AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlog
88,419Total entries
1Added by human
88,418Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,638 results
Model Releases

llm-anthropic 0.25.1

DGX agent

Release: llm-anthropic 0.25.1 New model: Claude Opus 4.8 (claude-opus-4.8). New -o fast 1 option for fast mode, for organizations with that feature enabled on their account. Default max_tokens for eac

model-releasessimon-willison
28 May 2026
Model Releases

ProgVLA: Progress-Aware Robot Manipulation Skill Learning

DGX agent

arXiv:2605.28231v1 Announce Type: cross Abstract: We present ProgVLA, a compact vision-language-action (VLA) model designed for reliable robot manipulation under tight compute and memory budgets. The

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
model-releasesarxiv-cs-lg
28 May 2026
Model Releases

SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents

DGX agent

arXiv:2605.28122v1 Announce Type: cross Abstract: A coding agent executes a benign task as a sequence of shell, file, and network actions, any of which can quietly exceed the authorized scope while th

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling

DGX agent

arXiv:2605.28179v1 Announce Type: new Abstract: Scaling laws guide large language model training by relating compute to cross-entropy loss, and recent work further extends them to predict downstream b

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and Language

DGX agent

arXiv:2605.27886v1 Announce Type: new Abstract: Tactile sensing is essential for robots to achieve human-like gentle manipulation. However, existing Vision-Language-Action (VLA) models struggle to exp

model-releasesarxiv-cs-ro
28 May 2026
Research

The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling

DGX agent

arXiv:2602.14862v2 Announce Type: replace-cross Abstract: Temperature scaling is a simple method that allows to control the uncertainty of probabilistic models. It is mostly used in two contexts: impr

researcharxiv-cs-ai
28 May 2026
Model Releases

UniMaia: Steering Chess Policies with Language for Human-like Play

DGX agent

arXiv:2605.27767v1 Announce Type: cross Abstract: Recent advances in large language models have enabled natural language to serve as a flexible interface for controlling complex systems, but often at

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

When do complex-valued neural networks help? A study of representation, geometry, and optimization

DGX agent

arXiv:2605.27673v1 Announce Type: new Abstract: Complex-valued Neural Networks (CVNNs) are often motivated by domains where information is naturally encoded in magnitude and phase. Yet complex-valued

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay

DGX agent

arXiv:2605.28069v1 Announce Type: new Abstract: Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression metho

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection

DGX agent

arXiv:2605.26533v1 Announce Type: cross Abstract: Automated industrial inspection requires both precise defect localization and structured maintenance report generation; in current practice these task

model-releasesarxiv-cs-ai
27 May 2026
Safety

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

DGX agent

arXiv:2605.26174v1 Announce Type: cross Abstract: Production language-model systems answer a request by partitioning it across an invisible orchestration of worker agents that recompose one integrated

safetyarxiv-cs-ai
27 May 2026
Model Releases

Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

DGX agent

arXiv:2605.26772v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may compl

model-releasesarxiv-cs-ai
27 May 2026
Research

Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens

DGX agent

arXiv:2505.13775v4 Announce Type: replace-cross Abstract: Recent impressive results from large reasoning models have been interpreted as a triumph of Chain of Thought (CoT), and especially of the proc

researcharxiv-cs-ai
27 May 2026
Model Releases

ChartAct: A Benchmark for Dynamic Chart Understanding

DGX agent

arXiv:2605.26994v1 Announce Type: new Abstract: Charts are widely used to present complex data for analysis and decision making. Existing chart understanding benchmarks mainly focus on static charts,

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

ConVer: Using Contracts and Loop Invariant Synthesis for Scalable Formal Software Verification

DGX agent

arXiv:2605.27051v1 Announce Type: cross Abstract: Formal verification of large C programs is impeded by state-space explosion: Bounded Model Checking (BMC) tools must encode the entire state space up

model-releasesarxiv-cs-ai
27 May 2026
Tutorials

Demystifying Video Reasoning

DGX agent

arXiv:2603.16870v2 Announce Type: replace-cross Abstract: Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabil

tutorialsarxiv-cs-ai
27 May 2026
Model Releases

Entropy Sentinel: Continuous LLM Accuracy Monitoring from Decoding Entropy Traces in STEM

DGX agent

arXiv:2601.09001v4 Announce Type: replace Abstract: Deploying LLMs raises two coupled challenges: (1) monitoring--estimating where a model underperforms as traffic and domains drift--and (2) improveme

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

'Give Me BF16 or Give Me Death'? Accuracy-Performance Trade-Offs in LLM Quantization

DGX agent

arXiv:2411.02355v4 Announce Type: replace-cross Abstract: Quantization is a powerful tool for accelerating large language model (LLM) inference, but the accuracy-performance trade-offs across differen

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training

DGX agent

arXiv:2602.02518v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly rely on external knowledge to improve factuality, yet many real-world knowledge sources are organize

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Memory Architectures for Multi-Turn Text-to-SQL: A Benchmark and Empirical Study

DGX agent

arXiv:2605.26394v1 Announce Type: new Abstract: Multi-turn Text-to-SQL is central to enterprise analytics yet remains predominantly evaluated in single-turn settings. We introduce EnterpriseMem-Bench,

model-releasesarxiv-cs-cl
27 May 2026
Research

Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning

DGX agent

arXiv:2605.26761v1 Announce Type: new Abstract: Multimodal instruction tuning is the de facto recipe for adapting vision language models (VLMs), yet instruction data are highly redundant, making data

researcharxiv-cs-cv
27 May 2026
Model Releases

PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions

DGX agent

arXiv:2605.27015v1 Announce Type: new Abstract: Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We in

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Qiskit QuantumKatas: Adapting Microsoft's Quantum Computing exercises for LLM evaluation

DGX agent

arXiv:2605.27210v1 Announce Type: cross Abstract: We adapt Microsoft's QuantumKatas -- a well-established quantum computing curriculum -- from Q# to Qiskit, the most widely-adopted quantum computing f

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference

DGX agent

arXiv:2605.27081v1 Announce Type: cross Abstract: Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining hi

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure

DGX agent

arXiv:2601.10566v5 Announce Type: replace Abstract: Entity-level unlearning is usually evaluated by what a model says: whether it stops naming the target, refuses a query, or shifts a Truth Ratio dist

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Rethinking the Multilingual Reasoning Gap with Layer Swap

DGX agent

arXiv:2605.26735v1 Announce Type: new Abstract: Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior wor

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Separating Semantic Competition from Context Length in RAG Reading

DGX agent

arXiv:2605.27294v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems can respond incorrectly even when the correct passage was retrieved. The model must still read the retrieve

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

DGX agent

arXiv:2510.09606v2 Announce Type: replace Abstract: With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still strug

model-releasesarxiv-cs-cv
27 May 2026
Research

SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking

DGX agent

arXiv:2511.04711v2 Announce Type: replace-cross Abstract: Large-scale vision-language models, especially CLIP, have demonstrated remarkable performance across diverse downstream tasks. Soft prompts, a

researcharxiv-cs-ai
27 May 2026
Model Releases

UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems

DGX agent

arXiv:2605.26646v1 Announce Type: new Abstract: LLM-based multi-agent systems decompose complex tasks into interacting roles, but most remain manually orchestrated by prompts, tools, and control rules

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization

DGX agent

arXiv:2605.26457v1 Announce Type: cross Abstract: AI coding agents are increasingly used to write real-world software, but ensuring that their outputs are correct remains a fundamental challenge. Form

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes

DGX agent

arXiv:2605.26380v1 Announce Type: cross Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios

DGX agent

arXiv:2508.19988v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved high accuracy on complex commonsense and mathematical problems that involve the composition of multiple r

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data

DGX agent

arXiv:2605.25549v1 Announce Type: cross Abstract: High-quality expert chain-of-thought (CoT) data is one of the core bottlenecks in large language model (LLM) post-training. Existing data production m

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

BODHI: Precise OS Kernel Specification Inference

DGX agent

arXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express

DGX agent

arXiv:2605.25891v1 Announce Type: cross Abstract: We find a mismatch between what large language models encode about a causal question and what they answer. On anti-commonsense CLadder items, a fixed

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference

DGX agent

arXiv:2510.02361v2 Announce Type: replace-cross Abstract: Transformer-based large models excel in natural language processing and computer vision, but face severe computational inefficiencies due to t

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

CITYREP: A Unified Benchmark for Urban Representations Across Cities, Tasks, and Modalities

DGX agent

arXiv:2605.26036v1 Announce Type: new Abstract: Urban representation learning encodes complex urban environments into general-purpose embeddings for diverse downstream tasks and emerging urban foundat

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

CMAP: Cross-Modal Adaptive Prompting for Multi-Domain Task-Incremental Learning

DGX agent

arXiv:2605.25708v1 Announce Type: cross Abstract: Multi-domain task-incremental learning requires a model to sequentially acquire knowledge across visually diverse domains without forgetting prior tas

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

DRInQ: Evaluating Conversational Implicature with Controlled Context Variation

DGX agent

arXiv:2605.24267v1 Announce Type: new Abstract: Human conversation relies heavily on conversational implicature, in which speakers convey meanings that are suggested rather than explicitly stated. Alt

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

DropoutTS: Sample-Adaptive Dropout for Robust Time Series Forecasting

DGX agent

arXiv:2601.21726v2 Announce Type: replace Abstract: Deep time series models are vulnerable to noisy data ubiquitous in real-world applications. Existing robustness strategies either prune data or rely

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Efficient DP-SGD for LLMs with Randomized Clipping

DGX agent

arXiv:2605.24879v1 Announce Type: new Abstract: Large language models (LLMs) are trained on vast datasets that may contain sensitive information. Differential privacy (DP), the de facto standard for f

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

DGX agent

arXiv:2605.25052v1 Announce Type: new Abstract: Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these t

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations

DGX agent

arXiv:2507.07644v4 Announce Type: replace Abstract: We introduce FloorplanQA, a diagnostic benchmark for evaluating spatial reasoning in large language models (LLMs). FloorplanQA is grounded in struct

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap

DGX agent

arXiv:2605.24432v1 Announce Type: new Abstract: Large Language Model (LLM) interactions are typically underspecified, with users clarifying all necessary details across multiple conversational turns.

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

DGX agent

arXiv:2602.18527v2 Announce Type: replace-cross Abstract: Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. T

model-releasesarxiv-cs-ai
26 May 2026
Safety

Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation

DGX agent

arXiv:2605.25036v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) extend large language models with visual understanding, but remain vulnerable to hallucination, where outputs are

safetyarxiv-cs-ai
26 May 2026
Model Releases

Llamion Technical Report

DGX agent

arXiv:2605.25676v1 Announce Type: new Abstract: We release Llamion, a family of 14B-parameter open-weight language models obtained by transforming Orion-14B into the standardized Llama-family architec

model-releasesarxiv-cs-cl
26 May 2026
← Previous
1…397398399400401…1326
Next →