AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
86,965Total entries
1Added by human
86,964Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,106 results
Model Releases

AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios

DGX agent

arXiv:2605.27995v1 Announce Type: new Abstract: Large language model (LLM)-based agents have shown strong capabilities in using external tools to solve complex tasks. However, existing evaluations oft

model-releasesarxiv-cs-ai
28 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Automating Formal Verification with Agent-Guided Tree Search

DGX agent

arXiv:2605.27485v1 Announce Type: cross Abstract: Formal verification offers a path to provably correct software, but writing verified code remains expensive enough that the technique is rarely used i

model-releasesarxiv-cs-lg
28 May 2026
Safety

Behavioural Analysis of Alignment Faking

DGX agent

arXiv:2605.27681v1 Announce Type: new Abstract: Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deploym

safetyarxiv-cs-ai
28 May 2026
Model Releases

Bounded-Compute Multimodal Regression for Product-Rating Prediction

DGX agent

arXiv:2605.27737v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly attractive for multimodal quality assessment, but their default reliance on autoregressive text generatio

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback

DGX agent

arXiv:2605.28010v1 Announce Type: new Abstract: Self-evolving large language models (LLMs) learn by generating their own training tasks and solutions, reducing reliance on human-curated supervision. H

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Data-Efficient On-Policy Distillation for Automatic Speech Recognition

DGX agent

arXiv:2605.28139v1 Announce Type: new Abstract: Building competitive automatic speech recognition (ASR) models usually requires large-scale au- dio supervision, which makes reproduction and specializa

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024

DGX agent

arXiv:2503.02857v5 Announce Type: replace-cross Abstract: In the age of increasingly realistic generative AI, robust deepfake detection is essential for mitigating fraud and disinformation. While many

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

DEPART: DEcomposing PARiTy across Multilingual LLMs

DGX agent

arXiv:2605.28163v1 Announce Type: cross Abstract: Multilingual Large Language Models (mLLMs) leaderboards report per-language accuracy but rarely explain why disparities emerge, leaving systemic biase

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Do LLMs Favor Their Providers? Measuring Vertical Integration Bias in Code Generation

DGX agent

arXiv:2605.28515v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become an integral part of software development, especially with the advent of agentic capabilities. Yet, many front

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter AdaptationTarget

DGX agent

arXiv:2605.27390v1 Announce Type: cross Abstract: Speculative decoding accelerates Large Language Model inference via a draft-then-verify paradigm, yet the output projection layer becomes a bottleneck

model-releasesarxiv-cs-ai
28 May 2026
Safety

FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning

DGX agent

arXiv:2605.28389v1 Announce Type: new Abstract: While large language models have made significant progress in mathematical reasoning, they remain unreliable at judging the correctness of their own sol

safetyarxiv-cs-cl
28 May 2026
Model Releases

Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment

DGX agent

arXiv:2605.28188v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making settings such as legal reasoning, where consistency under factuall

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification

DGX agent

arXiv:2605.28500v1 Announce Type: cross Abstract: Large language models have shown impressive capabilities in code generation, yet they often produce functionally incorrect code. Uncertainty quantific

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

DGX agent

arXiv:2605.28013v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision.

model-releasesarxiv-cs-cl
28 May 2026
Safety

LACUNA: Safe Agents as Recursive Program Holes

DGX agent

arXiv:2605.28617v1 Announce Type: new Abstract: LLM agents increasingly act by writing code, yet a split persists between the runtime that drives the agent and the code the model writes. The runtime o

safetyarxiv-cs-ai
28 May 2026
Tutorials

Learn from your own latents and not from tokens: A sample-complexity theory

DGX agent

arXiv:2605.27734v1 Announce Type: new Abstract: Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude lar

tutorialsarxiv-cs-lg
28 May 2026
Model Releases

ProgVLA: Progress-Aware Robot Manipulation Skill Learning

DGX agent

arXiv:2605.28231v1 Announce Type: cross Abstract: We present ProgVLA, a compact vision-language-action (VLA) model designed for reliable robot manipulation under tight compute and memory budgets. The

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents

DGX agent

arXiv:2605.28122v1 Announce Type: cross Abstract: A coding agent executes a benign task as a sequence of shell, file, and network actions, any of which can quietly exceed the authorized scope while th

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling

DGX agent

arXiv:2605.28179v1 Announce Type: new Abstract: Scaling laws guide large language model training by relating compute to cross-entropy loss, and recent work further extends them to predict downstream b

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and Language

DGX agent

arXiv:2605.27886v1 Announce Type: new Abstract: Tactile sensing is essential for robots to achieve human-like gentle manipulation. However, existing Vision-Language-Action (VLA) models struggle to exp

model-releasesarxiv-cs-ro
28 May 2026
Research

The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling

DGX agent

arXiv:2602.14862v2 Announce Type: replace-cross Abstract: Temperature scaling is a simple method that allows to control the uncertainty of probabilistic models. It is mostly used in two contexts: impr

researcharxiv-cs-ai
28 May 2026
Model Releases

UniMaia: Steering Chess Policies with Language for Human-like Play

DGX agent

arXiv:2605.27767v1 Announce Type: cross Abstract: Recent advances in large language models have enabled natural language to serve as a flexible interface for controlling complex systems, but often at

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

When do complex-valued neural networks help? A study of representation, geometry, and optimization

DGX agent

arXiv:2605.27673v1 Announce Type: new Abstract: Complex-valued Neural Networks (CVNNs) are often motivated by domains where information is naturally encoded in magnitude and phase. Yet complex-valued

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay

DGX agent

arXiv:2605.28069v1 Announce Type: new Abstract: Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression metho

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection

DGX agent

arXiv:2605.26533v1 Announce Type: cross Abstract: Automated industrial inspection requires both precise defect localization and structured maintenance report generation; in current practice these task

model-releasesarxiv-cs-ai
27 May 2026
Safety

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

DGX agent

arXiv:2605.26174v1 Announce Type: cross Abstract: Production language-model systems answer a request by partitioning it across an invisible orchestration of worker agents that recompose one integrated

safetyarxiv-cs-ai
27 May 2026
Model Releases

Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

DGX agent

arXiv:2605.26772v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may compl

model-releasesarxiv-cs-ai
27 May 2026
Research

Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens

DGX agent

arXiv:2505.13775v4 Announce Type: replace-cross Abstract: Recent impressive results from large reasoning models have been interpreted as a triumph of Chain of Thought (CoT), and especially of the proc

researcharxiv-cs-ai
27 May 2026
Model Releases

ChartAct: A Benchmark for Dynamic Chart Understanding

DGX agent

arXiv:2605.26994v1 Announce Type: new Abstract: Charts are widely used to present complex data for analysis and decision making. Existing chart understanding benchmarks mainly focus on static charts,

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

ConVer: Using Contracts and Loop Invariant Synthesis for Scalable Formal Software Verification

DGX agent

arXiv:2605.27051v1 Announce Type: cross Abstract: Formal verification of large C programs is impeded by state-space explosion: Bounded Model Checking (BMC) tools must encode the entire state space up

model-releasesarxiv-cs-ai
27 May 2026
Tutorials

Demystifying Video Reasoning

DGX agent

arXiv:2603.16870v2 Announce Type: replace-cross Abstract: Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabil

tutorialsarxiv-cs-ai
27 May 2026
Model Releases

Entropy Sentinel: Continuous LLM Accuracy Monitoring from Decoding Entropy Traces in STEM

DGX agent

arXiv:2601.09001v4 Announce Type: replace Abstract: Deploying LLMs raises two coupled challenges: (1) monitoring--estimating where a model underperforms as traffic and domains drift--and (2) improveme

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

'Give Me BF16 or Give Me Death'? Accuracy-Performance Trade-Offs in LLM Quantization

DGX agent

arXiv:2411.02355v4 Announce Type: replace-cross Abstract: Quantization is a powerful tool for accelerating large language model (LLM) inference, but the accuracy-performance trade-offs across differen

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training

DGX agent

arXiv:2602.02518v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly rely on external knowledge to improve factuality, yet many real-world knowledge sources are organize

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Memory Architectures for Multi-Turn Text-to-SQL: A Benchmark and Empirical Study

DGX agent

arXiv:2605.26394v1 Announce Type: new Abstract: Multi-turn Text-to-SQL is central to enterprise analytics yet remains predominantly evaluated in single-turn settings. We introduce EnterpriseMem-Bench,

model-releasesarxiv-cs-cl
27 May 2026
Research

Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning

DGX agent

arXiv:2605.26761v1 Announce Type: new Abstract: Multimodal instruction tuning is the de facto recipe for adapting vision language models (VLMs), yet instruction data are highly redundant, making data

researcharxiv-cs-cv
27 May 2026
Model Releases

PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions

DGX agent

arXiv:2605.27015v1 Announce Type: new Abstract: Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We in

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Qiskit QuantumKatas: Adapting Microsoft's Quantum Computing exercises for LLM evaluation

DGX agent

arXiv:2605.27210v1 Announce Type: cross Abstract: We adapt Microsoft's QuantumKatas -- a well-established quantum computing curriculum -- from Q# to Qiskit, the most widely-adopted quantum computing f

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference

DGX agent

arXiv:2605.27081v1 Announce Type: cross Abstract: Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining hi

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure

DGX agent

arXiv:2601.10566v5 Announce Type: replace Abstract: Entity-level unlearning is usually evaluated by what a model says: whether it stops naming the target, refuses a query, or shifts a Truth Ratio dist

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Rethinking the Multilingual Reasoning Gap with Layer Swap

DGX agent

arXiv:2605.26735v1 Announce Type: new Abstract: Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior wor

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Separating Semantic Competition from Context Length in RAG Reading

DGX agent

arXiv:2605.27294v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems can respond incorrectly even when the correct passage was retrieved. The model must still read the retrieve

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

DGX agent

arXiv:2510.09606v2 Announce Type: replace Abstract: With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still strug

model-releasesarxiv-cs-cv
27 May 2026
Research

SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking

DGX agent

arXiv:2511.04711v2 Announce Type: replace-cross Abstract: Large-scale vision-language models, especially CLIP, have demonstrated remarkable performance across diverse downstream tasks. Soft prompts, a

researcharxiv-cs-ai
27 May 2026
Model Releases

UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems

DGX agent

arXiv:2605.26646v1 Announce Type: new Abstract: LLM-based multi-agent systems decompose complex tasks into interacting roles, but most remain manually orchestrated by prompts, tools, and control rules

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization

DGX agent

arXiv:2605.26457v1 Announce Type: cross Abstract: AI coding agents are increasingly used to write real-world software, but ensuring that their outputs are correct remains a fundamental challenge. Form

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes

DGX agent

arXiv:2605.26380v1 Announce Type: cross Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios

DGX agent

arXiv:2508.19988v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved high accuracy on complex commonsense and mathematical problems that involve the composition of multiple r

model-releasesarxiv-cs-cl
26 May 2026
← Previous
1…319320321322323…1065
Next →