AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
16,965 results
Model Releases

SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment

DGX agent

arXiv:2608.07639v1 Announce Type: cross Abstract: Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill s

model-releasesarxiv-cs-ai
11 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

DGX agent

arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the app

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance

DGX agent

arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable pr

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

DGX agent

arXiv:2608.09771v1 Announce Type: new Abstract: Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every

model-releasesarxiv-cs-ro
11 Aug 2026
Model Releases

Smart Compaction: Predicting Compaction Utility from Lakehouse Table Metadata

DGX agent

arXiv:2608.08639v1 Announce Type: new Abstract: Open lakehouse table formats accumulate small data files over time, which degrades query performance. Deciding when compaction is worthwhile remains thr

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

DGX agent

arXiv:2608.09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and imp

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

DGX agent

arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Sparks of Cooperative Reasoning: LLMs as Strategic Hanabi Agents

DGX agent

arXiv:2601.18077v3 Announce Type: replace Abstract: Cooperative reasoning under incomplete information remains challenging for both humans and multi-agent systems. The card game Hanabi embodies this c

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Sparse corruption in low-rank matrix inference: the PCA benchmark

DGX agent

arXiv:2511.11927v2 Announce Type: replace-cross Abstract: Principal Component Analysis (PCA) is a standard tool for extracting a low-rank signal from noisy observations. It is known that applying PCA

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding

DGX agent

arXiv:2608.07915v1 Announce Type: new Abstract: Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Th

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention

DGX agent

arXiv:2608.07921v1 Announce Type: cross Abstract: We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bu

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SpikeWorld: Fast-State Adaptation for Frozen Spiking World Models

DGX agent

arXiv:2608.07712v1 Announce Type: cross Abstract: A predictive model receives a self-supervised signal whenever the consequence of an action is observed. Using that signal after deployment is difficul

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation

DGX agent

arXiv:2601.09974v2 Announce Type: replace Abstract: Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant over tim

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Stateful CARS: Exact Cross-History Reuse for Policy-Constrained LLM Agents

DGX agent

arXiv:2608.08282v1 Announce Type: new Abstract: Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model dist

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Staying True to the Origin: Continuous Image Stylization with Smooth Transitions

DGX agent

arXiv:2608.08125v1 Announce Type: new Abstract: Recent advances in generative models have achieved remarkable performance in text- and image-conditioned editing. However, preserving the content of a g

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Subjective Multi-Bias Detection with Large Language Models

DGX agent

arXiv:2608.09126v1 Announce Type: new Abstract: In this project, we delved into the pervasive challenge of bias detection within the text content. More specifically, our focus lies on the identificati

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation

DGX agent

arXiv:2510.14357v2 Announce Type: replace-cross Abstract: Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily relyin

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SuperCoder: Assembly Program Superoptimization with Large Language Models

DGX agent

arXiv:2505.11480v4 Announce Type: replace-cross Abstract: Superoptimization is the task of transforming a program into a faster one, and ideally the very fastest possible one, while preserving its inp

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents

DGX agent

arXiv:2608.08253v1 Announce Type: new Abstract: AI agents are becoming shared infrastructure, yet durable memory is commonly assembled from separate retrieval, governance, and operational components.

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

DGX agent

arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguist

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

SurgWMBench: A Vision-Based Benchmark for World-Modeling Surgical Instrument Motion Planning

DGX agent

arXiv:2608.08070v1 Announce Type: new Abstract: Reliable surgical planning requires models that move beyond recognizing the current surgical step or imitating expert demonstrations, and instead antici

model-releasesarxiv-cs-ro
11 Aug 2026
Model Releases

SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

DGX agent

arXiv:2608.07641v1 Announce Type: new Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

DGX agent

arXiv:2608.09802v1 Announce Type: new Abstract: As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluati

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping

DGX agent

arXiv:2608.09497v1 Announce Type: new Abstract: Operational crop mapping requires models that generalise across years, resolve fine-grained crop taxonomies, and distinguish cropland from surrounding l

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Tabular Numeric Stretch Transformation

DGX agent

arXiv:2608.09162v1 Announce Type: cross Abstract: Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distributions, scale

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Task-Oriented Formation Decision via Reinforcement Learning: Herding an Attacking Swarm

DGX agent

arXiv:2608.09258v1 Announce Type: new Abstract: Multi-robot systems can accomplish tasks that are difficult for a single robot by organizing into task-specific formations. Different from existing stud

model-releasesarxiv-cs-ro
11 Aug 2026
Model Releases

Task-to-Model Optimization for Enterprise LLM Coding Assistants: A Data-Driven Framework for Cost-Optimal Routing

DGX agent

arXiv:2608.08528v1 Announce Type: new Abstract: Enterprise AI coding assistants incur substantial inference spend, and naive token-cost minimization often fails to reduce end-to-end cost once retries,

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

DGX agent

arXiv:2608.09538v1 Announce Type: cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation.

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?

DGX agent

arXiv:2608.07899v1 Announce Type: new Abstract: Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure orig

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing

DGX agent

arXiv:2608.07851v1 Announce Type: cross Abstract: Residual connections rely on a static residual pathway, and are essential for training deep neural networks. Hyper-connections (HC) increase the expre

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Temporal Generalization in fNIRS-Based Autism Classification: A Cross-Time-Window Transfer Benchmark

DGX agent

arXiv:2608.07567v1 Announce Type: cross Abstract: Functional near-infrared spectroscopy (fNIRS) is a promising modality for autism spectrum disorder (ASD) classification, yet existing approaches assum

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law

DGX agent

arXiv:2608.09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the ap

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Test-time Generalization for Physics through Neural Operator Splitting

DGX agent

arXiv:2602.00884v2 Announce Type: replace Abstract: Neural operators have shown promise in learning solution maps of partial differential equations (PDEs), but they often struggle to generalize when t

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Test-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation

DGX agent

arXiv:2608.08290v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers

DGX agent

arXiv:2608.08809v1 Announce Type: new Abstract: A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the rig

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

DGX agent

arXiv:2608.07617v1 Announce Type: new Abstract: Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatche

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

The Authority Expectancy Effect in Multi-User Conflict

DGX agent

arXiv:2608.08026v1 Announce Type: new Abstract: We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a m

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

The Cell Must Go On: Agar.io for Continual Reinforcement Learning

DGX agent

arXiv:2505.18347v3 Announce Type: replace-cross Abstract: Continual reinforcement learning (RL) concerns agents that are expected to learn continually, rather than converge to a policy that is then fi

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation

DGX agent

arXiv:2511.02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independentl

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

The Cost of Adaptivity: Matching Lower Bounds Across Learning Problems

DGX agent

arXiv:2608.08826v1 Announce Type: new Abstract: Adaptive procedures must work without nuisance information an oracle may use, such as a gradient scale or smoothness index, and robust procedures may ha

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

DGX agent

arXiv:2608.08650v1 Announce Type: new Abstract: Mixture-of-Experts models increase parameter capacity while keeping the computation activated by each token bounded, but their architectural evolution c

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction

DGX agent

arXiv:2608.07517v1 Announce Type: cross Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aw

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

The Knowing-Saying Gap: When Probes See Errors that Confidence Misses

DGX agent

arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games

DGX agent

arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move fro

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task

DGX agent

arXiv:2608.08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it reaches its tool

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Theory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration

DGX agent

arXiv:2608.08881v1 Announce Type: new Abstract: The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments we

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Think Deep, Speak Once: Relit, A Recursive Latent Implicit Transformer Framework

DGX agent

arXiv:2608.08113v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has become the dominant paradigm for eliciting reasoning in Large Language Models (LLMs), yet it creates substantial co

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders

DGX agent

arXiv:2608.08168v1 Announce Type: new Abstract: While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this e

model-releasesarxiv-cs-cl
11 Aug 2026
← Previous
1…910111213…354
Next →