AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,106 results
11 Aug 2026

Shape Mutating Expert Compression:LorExperts and BTExperts

Model ReleasesDGX agent

arXiv:2608.07814v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many ex

Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic

Model ReleasesDGX agent

arXiv:2601.22510v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanis

SHE: Trajectory-driven Safety Harness Evolution for LLM Agents

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, per

SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision

Model ReleasesDGX agent

arXiv:2608.09097v1 Announce Type: new Abstract: Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly

SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model

Model ReleasesDGX agent

arXiv:2606.14574v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmarks

SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment

Model ReleasesDGX agent

arXiv:2608.07639v1 Announce Type: cross Abstract: Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill s

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

Model ReleasesDGX agent

arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the app

SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance

Model ReleasesDGX agent

arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable pr

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

Model ReleasesDGX agent

arXiv:2608.09771v1 Announce Type: new Abstract: Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every

Smart Compaction: Predicting Compaction Utility from Lakehouse Table Metadata

Model ReleasesDGX agent

arXiv:2608.08639v1 Announce Type: new Abstract: Open lakehouse table formats accumulate small data files over time, which degrades query performance. Deciding when compaction is worthwhile remains thr

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

Model ReleasesDGX agent

arXiv:2608.09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and imp

SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat

Sources: Nvidia is developing a Nemotron 4 model with 1T+ parameters, up from Nemotron 3 Ultra's 550B parameters but smaller than leading Chinese open models (The Information)

Model ReleasesDGX agent

The Information: Sources: Nvidia is developing a Nemotron 4 model with 1T+ parameters, up from Nemotron 3 Ultra's 550B parameters but smaller than leading Chinese open models — Nvidia is doubling down

Sparks of Cooperative Reasoning: LLMs as Strategic Hanabi Agents

Model ReleasesDGX agent

arXiv:2601.18077v3 Announce Type: replace Abstract: Cooperative reasoning under incomplete information remains challenging for both humans and multi-agent systems. The card game Hanabi embodies this c

Sparse corruption in low-rank matrix inference: the PCA benchmark

Model ReleasesDGX agent

arXiv:2511.11927v2 Announce Type: replace-cross Abstract: Principal Component Analysis (PCA) is a standard tool for extracting a low-rank signal from noisy observations. It is known that applying PCA

SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding

Model ReleasesDGX agent

arXiv:2608.07915v1 Announce Type: new Abstract: Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Th

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention

Model ReleasesDGX agent

arXiv:2608.07921v1 Announce Type: cross Abstract: We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bu

SpikeWorld: Fast-State Adaptation for Frozen Spiking World Models

Model ReleasesDGX agent

arXiv:2608.07712v1 Announce Type: cross Abstract: A predictive model receives a self-supervised signal whenever the consequence of an action is observed. Using that signal after deployment is difficul

SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation

Model ReleasesDGX agent

arXiv:2601.09974v2 Announce Type: replace Abstract: Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant over tim

Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents

Model ReleasesDGX agent

arXiv:2608.08282v1 Announce Type: new Abstract: Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model dist

Staying True to the Origin: Continuous Image Stylization with Smooth Transitions

Model ReleasesDGX agent

arXiv:2608.08125v1 Announce Type: new Abstract: Recent advances in generative models have achieved remarkable performance in text- and image-conditioned editing. However, preserving the content of a g

Stealing Reasoning Traces from Proprietary LLM APIs

Model ReleasesDGX agent

Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name (stolen-thoughts.com) for a neat paper: Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that

Subjective Multi-Bias Detection with Large Language Models

Model ReleasesDGX agent

arXiv:2608.09126v1 Announce Type: new Abstract: In this project, we delved into the pervasive challenge of bias detection within the text content. More specifically, our focus lies on the identificati

SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation

Model ReleasesDGX agent

arXiv:2510.14357v2 Announce Type: replace-cross Abstract: Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily relyin

SuperCoder: Assembly Program Superoptimization with Large Language Models

Model ReleasesDGX agent

arXiv:2505.11480v4 Announce Type: replace-cross Abstract: Superoptimization is the task of transforming a program into a faster one, and ideally the very fastest possible one, while preserving its inp

SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents

Model ReleasesDGX agent

arXiv:2608.08253v1 Announce Type: new Abstract: AI agents are becoming shared infrastructure, yet durable memory is commonly assembled from separate retrieval, governance, and operational components.

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

Model ReleasesDGX agent

arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguist

SurgWMBench: A Vision-Based Benchmark for World-Modeling Surgical Instrument Motion Planning

Model ReleasesDGX agent

arXiv:2608.08070v1 Announce Type: new Abstract: Reliable surgical planning requires models that move beyond recognizing the current surgical step or imitating expert demonstrations, and instead antici

SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

Model ReleasesDGX agent

arXiv:2608.07641v1 Announce Type: new Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

Model ReleasesDGX agent

arXiv:2608.09802v1 Announce Type: new Abstract: As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluati

SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping

Model ReleasesDGX agent

arXiv:2608.09497v1 Announce Type: new Abstract: Operational crop mapping requires models that generalise across years, resolve fine-grained crop taxonomies, and distinguish cropland from surrounding l

Tabular Numeric Stretch Transformation

Model ReleasesDGX agent

arXiv:2608.09162v1 Announce Type: cross Abstract: Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distributions, scale

take away the symbolic part of this and it just would not have worked. Claude on Riemann is yet another victory for hybrid, neurosymbolic AI…

Model ReleasesDGX agent

On August 11 2026, Gary Marcus commented that if the symbolic components were removed from Claude’s Riemann implementation it would fail to work, highlighting a concrete win for hybrid neurosymbolic A

Task-Oriented Formation Decision via Reinforcement Learning: Herding an Attacking Swarm

Model ReleasesDGX agent

arXiv:2608.09258v1 Announce Type: new Abstract: Multi-robot systems can accomplish tasks that are difficult for a single robot by organizing into task-specific formations. Different from existing stud

Task-to-Model Optimization for Enterprise LLM Coding Assistants: A Data-Driven Framework for Cost-Optimal Routing

Model ReleasesDGX agent

arXiv:2608.08528v1 Announce Type: new Abstract: Enterprise AI coding assistants incur substantial inference spend, and naive token-cost minimization often fails to reduce end-to-end cost once retries,

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

Model ReleasesDGX agent

arXiv:2608.09538v1 Announce Type: cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation.

TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?

Model ReleasesDGX agent

arXiv:2608.07899v1 Announce Type: new Abstract: Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure orig

TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing

Model ReleasesDGX agent

arXiv:2608.07851v1 Announce Type: cross Abstract: Residual connections rely on a static residual pathway, and are essential for training deep neural networks. Hyper-connections (HC) increase the expre

Temporal Generalization in fNIRS-Based Autism Classification: A Cross-Time-Window Transfer Benchmark

Model ReleasesDGX agent

arXiv:2608.07567v1 Announce Type: cross Abstract: Functional near-infrared spectroscopy (fNIRS) is a promising modality for autism spectrum disorder (ASD) classification, yet existing approaches assum

Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law

Model ReleasesDGX agent

arXiv:2608.09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the ap

Test-time Generalization for Physics through Neural Operator Splitting

Model ReleasesDGX agent

arXiv:2602.00884v2 Announce Type: replace Abstract: Neural operators have shown promise in learning solution maps of partial differential equations (PDEs), but they often struggle to generalize when t

Test-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation

Model ReleasesDGX agent

arXiv:2608.08290v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing

Tested in Coding: BF16 Muse Glimmer vs BF16 Qwen3.6 27B

Model ReleasesDGX agent

I'm guessing that many people have been waiting for this comparison. For clarity, both models are running at full FP16 KV-cache. Due to VRAM limitations, Muse Glimmer is running full 262,144 context,

Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers

Model ReleasesDGX agent

arXiv:2608.08809v1 Announce Type: new Abstract: A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the rig

TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

Model ReleasesDGX agent

arXiv:2608.07617v1 Announce Type: new Abstract: Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatche

⚙️That is the framework Mistral is building toward: one in which enterprises, governments, and startups can use the best AI available, shape…

Model ReleasesDGX agent

⚙️That is the framework Mistral is building toward: one in which enterprises, governments, and startups can use the best AI available, shape it around their own knowledge, and retain the value it crea

The Authority Expectancy Effect in Multi-User Conflict

Model ReleasesDGX agent

arXiv:2608.08026v1 Announce Type: new Abstract: We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a m

The Cell Must Go On: Agar.io for Continual Reinforcement Learning

Model ReleasesDGX agent

arXiv:2505.18347v3 Announce Type: replace-cross Abstract: Continual reinforcement learning (RL) concerns agents that are expected to learn continually, rather than converge to a policy that is then fi

The ChatGPT desktop app is now available in preview for desktop variants of these Linux distributions: • Ubuntu 24.04 LTS and 26.04 LTS • De…

Model ReleasesDGX agent

The ChatGPT desktop app is now available in preview for desktop variants of these Linux distributions: • Ubuntu 24.04 LTS and 26.04 LTS • Debian 13 • Fedora 43 and 44 Install with .deb or .rpm package

The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation

Model ReleasesDGX agent

arXiv:2511.02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independentl

The Cost of Adaptivity: Matching Lower Bounds Across Learning Problems

Model ReleasesDGX agent

arXiv:2608.08826v1 Announce Type: new Abstract: Adaptive procedures must work without nuisance information an oracle may use, such as a gradient scale or smoothness index, and robust procedures may ha

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

Model ReleasesDGX agent

arXiv:2608.08650v1 Announce Type: new Abstract: Mixture-of-Experts models increase parameter capacity while keeping the computation activated by each token bounded, but their architectural evolution c

The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction

Model ReleasesDGX agent

arXiv:2608.07517v1 Announce Type: cross Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aw

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

Model ReleasesDGX agent

arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The

The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games

Model ReleasesDGX agent

arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move fro

The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task

Model ReleasesDGX agent

arXiv:2608.08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it reaches its tool

The small open weight models are scarier in AI development

Model ReleasesDGX agent

Imagine if your everyday laptop could run an AI model smart enough to take care of 90% of your work—totally private, lightning fast, and completely free of monthly fees. That is the exact tipping poin

💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to c…

Model ReleasesDGX agent

💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to choose the right model for the right task. As part of this, we

Theory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration

Model ReleasesDGX agent

arXiv:2608.08881v1 Announce Type: new Abstract: The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments we

Think Deep, Speak Once: Relit, A Recursive Latent Implicit Transformer Framework

Model ReleasesDGX agent

arXiv:2608.08113v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has become the dominant paradigm for eliciting reasoning in Large Language Models (LLMs), yet it creates substantial co

← Previous
1…1011121314…369
Next →