AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
11 Aug 2026

Compiling and Benchmarking Task-State Horizons for Embodied Agents

Model ReleasesDGX agent

arXiv:2608.08036v1 Announce Type: new Abstract: Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long

Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge

Model ReleasesDGX agent

arXiv:2608.07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous structures. Howe

Decoupled Descent: Enforcing Exact Train-Test Error Tracking Via AMP Onsager Corrections [R]

Model Releases
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

Link: https://arxiv.org/pdf/2604.27883 Hi, Most of use are familiar with the headache of training a neural network using gradient descent where the training error may go to zero but the test error may

DeepSeek-V4-Flash-0731 (284B MoE) at 75 tok/s on 2× DGX Spark — full recipe, 11 gotchas, reboot-proof cluster, Codex CLI integration

Model ReleasesDGX agent

Spent two nights getting deepseek-ai/DeepSeek-V4-Flash-0731 (284B MoE, 13B active, native FP4/FP8, 1M context) running production-grade on two DGX Sparks connected by one QSFP DAC cable. Everything —

Full-bandwidth transformer

Model ReleasesDGX agent

arXiv:2608.08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each

GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction

Model ReleasesDGX agent

arXiv:2608.09493v1 Announce Type: new Abstract: Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains chal

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment

Model ReleasesDGX agent

arXiv:2608.08212v1 Announce Type: new Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Existing prompts

Hidden Language Consistency Phenomena in Reasoning LLMs

Model ReleasesDGX agent

arXiv:2608.08447v1 Announce Type: cross Abstract: Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended langu

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems

Model ReleasesDGX agent

arXiv:2608.07861v1 Announce Type: new Abstract: Vision-language models (VLMs) are becoming a practical backend for mobile visual question answering (VQA) systems, enabling smartphones and smart glasse

I ran Muse Glimmer @ 1M context - All tests passed.

Model ReleasesDGX agent

Heeeey all! I just completed some fun tests with Muse Glimmer, I thought I'd let you know. In fact, the summary below was written by Muse itself! I ran a 2× DGX Spark cluster and got Meta's day-old Mu

ICM Out! Better Tournament Strategy from Computed Continuations, vs. Solvers and LLMs

Model ReleasesDGX agent

arXiv:2608.09586v1 Announce Type: new Abstract: The Independent Chip Model (ICM) converts tournament chips into reference prize equity, and policies are routinely constructed against those values. Bec

IndexTTS 2.5 Technical Report

Model ReleasesDGX agent

arXiv:2601.03888v4 Announce Type: replace-cross Abstract: In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-base

Learning to Triage Vulnerability Reports from Program Analysis: An Empirical Study in Node.js

Model ReleasesDGX agent

arXiv:2510.20739v2 Announce Type: replace-cross Abstract: Program analysis tools often produce large volumes of candidate vulnerability reports that require costly manual review, creating a practical

LexKairos: Benchmarking Legal Temporal Capabilities in LLMs

Model ReleasesDGX agent

arXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that

Looker’s semantic layer governs Gemini Enterprise data for user trust

Model ReleasesDGX agent

For organizations deploying AI agents at scale, there’s often a critical divide between structured and unstructured data. While large language models (LLMs) excel at parsing text documents, emails, an

LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks

Model ReleasesDGX agent

arXiv:2608.07749v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resources, but a si

Math-Vision Diagrams: A Comprehensive Benchmark for Evaluating LLM Mathematical Diagram Generation Capabilities

Model ReleasesDGX agent

arXiv:2608.08964v1 Announce Type: new Abstract: The generation of mathematically precise diagrams from tex- tual prompts has emerged as a critical yet underexplored capability of Large Language Models

MemeMind: Reference-Guided Trace Construction for Offline Context Optimization

Model ReleasesDGX agent

arXiv:2608.09316v1 Announce Type: new Abstract: Offline context optimization improves an agent by revising its instructions and examples while keeping the model frozen. This approach learns from rollo

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

Model ReleasesDGX agent

arXiv:2608.09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information e

Persistent Semantic Entities in Tool-Augmented LLM Systems

Model ReleasesDGX agent

arXiv:2608.07952v1 Announce Type: cross Abstract: Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries---

Same Question, Different Answer? Measuring and Mitigating Prompt Privilege for Equitable AI Access

Model ReleasesDGX agent

arXiv:2608.08942v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should provide

Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic

Model ReleasesDGX agent

arXiv:2601.22510v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanis

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

Model ReleasesDGX agent

arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the app

SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation

Model ReleasesDGX agent

arXiv:2601.09974v2 Announce Type: replace Abstract: Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant over tim

Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification

Local AiDGX agent

arXiv:2608.08195v1 Announce Type: cross Abstract: Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because

The Authority Expectancy Effect in Multi-User Conflict

Model ReleasesDGX agent

arXiv:2608.08026v1 Announce Type: new Abstract: We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a m

The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation

Model ReleasesDGX agent

arXiv:2511.02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independentl

Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs

Model ReleasesDGX agent

arXiv:2608.08506v1 Announce Type: new Abstract: Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter cou

v0.32.8

Model ReleasesDGX agent

Muse Glimmer Muse Glimmer is now available on all platforms. Muse Glimmer can power coding agent applications such as Claude Code, Codex, Pi and more, as well as long-running personal assistants such

When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information

ResearchDGX agent

arXiv:2608.09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability u

10 Aug 2026

b10338

Model ReleasesDGX agent

model-saver : fix expert shared/chunk FFN length key clobber (#26693) The saver called add_kv with LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH twice, the second time passing n_ff_chexp. gguf_set_val_u32

Best open-source harness like Claude Code?

Model ReleasesDGX agent

Avid claude code user here looking to do equivalent things with local models. Just want to plug in something like Qwen and have the interface be 1:1 with claude code. Any suggestion? submitted by /u/N

CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment

SafetyDGX agent

arXiv:2604.00310v2 Announce Type: replace-cross Abstract: Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Mod

Corma launches with $60M in funding for defensive cybersecurity AI

Model ReleasesDGX agent

Defensive cybersecurity startup Corma Labs Ltd. today announced it has raised 60 million in seed funding to build a foundation model purpose-built for security defense. Founded in 2025, Corma runs off

Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation

Model ReleasesDGX agent

arXiv:2608.06631v1 Announce Type: cross Abstract: Cryptanalytic extraction has been demonstrated for ReLU networks, for networks using componentwise activations such as GELU or SiLU, and for a Transfo

How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality

ResearchDGX agent

arXiv:2604.06756v2 Announce Type: replace Abstract: Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible

Let's Unlearn Stereotypes Before Decision-Making: Assessing the Impact of Intrinsic Bias Mitigation on Downstream Fairness in LLMs

SafetyDGX agent

arXiv:2509.16462v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly used in high-stakes decision-making systems, where biased predictions can reinforce social and economi

Natural Language Processing Psychometrics

ResearchDGX agent

arXiv:2608.07316v1 Announce Type: cross Abstract: Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content,

Semantic Adapter Routing with Fine-Tuning Task Embeddings

Model ReleasesDGX agent

arXiv:2606.19079v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning (PEFT) has led to model ecosystems in which a single backbone is paired with many task-specialized adapters. Given s

Semi Edge Inference Idea [D]

TutorialsDGX agent

Today the most important factor in AI is cost. My idea is to split ML models inference (closed ones, proprietary) across server and edge computing on clients, and I would like to hear what do you thin

SLED: Scalable Location Encoding via Distillation

Model ReleasesDGX agent

arXiv:2608.06612v1 Announce Type: cross Abstract: The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer siz

Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs

Model ReleasesDGX agent

arXiv:2608.07014v1 Announce Type: new Abstract: Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing

Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control

Model ReleasesDGX agent

arXiv:2608.06758v1 Announce Type: new Abstract: We present Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built from Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16

// The Bitter Lesson of Tool Calling // Tool calling is a design choice, and the defaults are quietly costing accuracy. How so? New research…

Model ReleasesDGX agent

// The Bitter Lesson of Tool Calling // Tool calling is a design choice, and the defaults are quietly costing accuracy. How so? New research releases a generation-spanning comparison of programmatic t

The Sparsity Whisperer

Model ReleasesDGX agent

arXiv:2608.06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We

UAV3DCrop: Benchmarking 3D Reconstruction in Repeated Multi-Angle UAV Crop Surveys

Model ReleasesDGX agent

arXiv:2608.06404v1 Announce Type: new Abstract: Accurate 3D crop monitoring underpins data-driven precision agriculture by enabling field-scale analysis of plant structure, growth dynamics, and manage

We've used GPT-5.6-Cyber extensively in real-world vulnerability research, including work that uncovered previously unknown vulnerabilities …

Model ReleasesDGX agent

OpenAI announced the release of GPT‑5.6‑Cyber as part of its Cybersecurity Initiative, “Daybreak.” The model is aimed at advanced, authorized security research and testing, helping trusted defenders d

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

Model ReleasesDGX agent

arXiv:2608.07341v1 Announce Type: cross Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. extbf{Contamination mitigation

9 Aug 2026

[2606.05682] Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation

Model ReleasesDGX agent

Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost constrained production environments. Quantization

The Gemma team will host a special event on August 20

Model ReleasesDGX agent

Tweet by u/hackerllama Could be copium, but I would love to see Gemma 4.1 there with unified audio input for all model sizes perhaps even up to 120B, much improved tool calling (even with the latest t

8 Aug 2026

any reasonably fast public benchmarks I should run quants of deepseek flash 0731 on?

Model ReleasesDGX agent

I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 m

Anyone else amped up over Qwen 3.8?

Model ReleasesDGX agent

I’ve been using 3.6 27B Q4, and that quant is fast on an M5. The code has been average, but consistently “good enough.” And, after a year, I can see home LLMs being served at home much like streaming

Auto mode is now the default in Claude Code for Pro, Max, and Team plans

Model ReleasesDGX agent

Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode, to the point that they are making it the default setting for new ses

Qwen 35B-A3B MoE vs 27B dense in local coding tests: ~4× faster, much smaller quality gap than I expected

Model ReleasesDGX agent

I compared Qwen 35B-A3B MoE against Qwen 27B dense on a series of local coding-maintenance tasks. On my R9700/llama.cpp setup, the MoE model generated about 3.9× faster (~116 vs ~30 tok/s), but the co

Tesla V100 Qwen3.6 27B Performance

Model ReleasesDGX agent

Looking for V100 users to share your config and it's performance. GPU: Tesla V100 PCIE 32Gb Qwen3.6 27B Q4_K_M + Q8_0 MTP 128K context length Pi coding agent llama.cpp model preset: [*] spec-default =

7 Aug 2026

Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation

Model ReleasesDGX agent

arXiv:2503.03556v3 Announce Type: replace Abstract: Object affordance reasoning, the ability to infer object functionalities based on physical properties, is fundamental for task-oriented planning and

Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks

Model ReleasesDGX agent

arXiv:2608.05266v1 Announce Type: new Abstract: Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and sync

Anyone running DeepSeek-V4-Flash-0731 on MI325X with vLLM? Mine is behaving completely broken

Model ReleasesDGX agent

Is anyone here successfully running DeepSeek-V4-Flash-0731 locally with vLLM, especially on AMD MI325X? My setup: GPU: 1x AMD Instinct MI325X Model: deepseek-ai/DeepSeek-V4-Flash-0731 vLLM: 0.26.0 ROC

Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers

Model ReleasesDGX agent

arXiv:2608.06111v1 Announce Type: cross Abstract: Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to extit{syntactic structure}. We introduce extbf{

Continual Learning in Transition

Model ReleasesDGX agent

arXiv:2608.06216v1 Announce Type: cross Abstract: Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g.,

← Previous
1…295296297298299…1036
Next →