AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,020
  • Agents7,759
  • Applications5,540
  • Concepts5
  • Hardware1,925
  • Industry6,204
  • Local Ai5,102
  • Model Releases24,783
  • Research20,783
  • Safety13,742
  • Syntheses17
  • Tools1,680
  • Tutorials3,480

Source
HumanDGX agent

Content type
91,020Total entries
1Added by human
91,019Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,766 results
Model Releases

ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure

DGX agent

arXiv:2605.30284v1 Announce Type: new Abstract: Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks ha

model-releasesarxiv-cs-ai
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation

DGX agent

arXiv:2605.29579v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) have achieved rapid progress in vision-language understanding, they remain prone to multimodal hallucinat

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

DGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Structure-Aware Text Recognition for Ancient Greek Critical Editions

DGX agent

arXiv:2603.02803v2 Announce Type: replace Abstract: Recent advances in visual language models (VLMs) have transformed end-to-end document understanding. However, their ability to interpret the complex

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

DGX agent

arXiv:2602.15382v2 Announce Type: replace Abstract: Multi-Agent Systems (MAS) powered by Large Language Models have unlocked advanced collaborative reasoning, yet they remain bottlenecked by discrete

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection

DGX agent

arXiv:2605.30189v1 Announce Type: cross Abstract: We show that LoRA adapters, the dominant distribution format for fine-tuned LLMs, can be reliably backdoored through training data poisoning while pre

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Toward Ethical Facial Age Estimation: A Generalized Zero-Shot Benchmark Without Training on Children's Data

DGX agent

arXiv:2605.29230v1 Announce Type: cross Abstract: Age estimation from facial images typically relies on training data that includes images of minors, a practice that raises serious ethical, legal, and

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Training Deliberative Monitors for Black-Box Scheming Detection

DGX agent

arXiv:2605.29601v1 Announce Type: cross Abstract: As autonomous agents become more capable of performing real-world tasks, distinguishing scheming behavior from benign task pursuit may become a centra

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Who can we trust? LLM-as-a-jury for Comparative Assessment

DGX agent

arXiv:2602.16610v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly applied as automatic evaluators for natural language generation assessment often using pairwise

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

A Query Engine for the Agents

DGX agent

arXiv:2605.27785v1 Announce Type: new Abstract: The fastest-growing data in production today is unstructured text: agent traces, chat logs, reasoning chains, model outputs. People want to analyze it,

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios

DGX agent

arXiv:2605.27995v1 Announce Type: new Abstract: Large language model (LLM)-based agents have shown strong capabilities in using external tools to solve complex tasks. However, existing evaluations oft

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Automating Formal Verification with Agent-Guided Tree Search

DGX agent

arXiv:2605.27485v1 Announce Type: cross Abstract: Formal verification offers a path to provably correct software, but writing verified code remains expensive enough that the technique is rarely used i

model-releasesarxiv-cs-lg
28 May 2026
Safety

Behavioural Analysis of Alignment Faking

DGX agent

arXiv:2605.27681v1 Announce Type: new Abstract: Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deploym

safetyarxiv-cs-ai
28 May 2026
Model Releases

Bounded-Compute Multimodal Regression for Product-Rating Prediction

DGX agent

arXiv:2605.27737v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly attractive for multimodal quality assessment, but their default reliance on autoregressive text generatio

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback

DGX agent

arXiv:2605.28010v1 Announce Type: new Abstract: Self-evolving large language models (LLMs) learn by generating their own training tasks and solutions, reducing reliance on human-curated supervision. H

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Data-Efficient On-Policy Distillation for Automatic Speech Recognition

DGX agent

arXiv:2605.28139v1 Announce Type: new Abstract: Building competitive automatic speech recognition (ASR) models usually requires large-scale au- dio supervision, which makes reproduction and specializa

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024

DGX agent

arXiv:2503.02857v5 Announce Type: replace-cross Abstract: In the age of increasingly realistic generative AI, robust deepfake detection is essential for mitigating fraud and disinformation. While many

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

DEPART: DEcomposing PARiTy across Multilingual LLMs

DGX agent

arXiv:2605.28163v1 Announce Type: cross Abstract: Multilingual Large Language Models (mLLMs) leaderboards report per-language accuracy but rarely explain why disparities emerge, leaving systemic biase

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Do LLMs Favor Their Providers? Measuring Vertical Integration Bias in Code Generation

DGX agent

arXiv:2605.28515v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become an integral part of software development, especially with the advent of agentic capabilities. Yet, many front

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter AdaptationTarget

DGX agent

arXiv:2605.27390v1 Announce Type: cross Abstract: Speculative decoding accelerates Large Language Model inference via a draft-then-verify paradigm, yet the output projection layer becomes a bottleneck

model-releasesarxiv-cs-ai
28 May 2026
Safety

FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning

DGX agent

arXiv:2605.28389v1 Announce Type: new Abstract: While large language models have made significant progress in mathematical reasoning, they remain unreliable at judging the correctness of their own sol

safetyarxiv-cs-cl
28 May 2026
Model Releases

Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment

DGX agent

arXiv:2605.28188v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making settings such as legal reasoning, where consistency under factuall

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification

DGX agent

arXiv:2605.28500v1 Announce Type: cross Abstract: Large language models have shown impressive capabilities in code generation, yet they often produce functionally incorrect code. Uncertainty quantific

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

DGX agent

arXiv:2605.28013v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision.

model-releasesarxiv-cs-cl
28 May 2026
Safety

LACUNA: Safe Agents as Recursive Program Holes

DGX agent

arXiv:2605.28617v1 Announce Type: new Abstract: LLM agents increasingly act by writing code, yet a split persists between the runtime that drives the agent and the code the model writes. The runtime o

safetyarxiv-cs-ai
28 May 2026
Tutorials

Learn from your own latents and not from tokens: A sample-complexity theory

DGX agent

arXiv:2605.27734v1 Announce Type: new Abstract: Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude lar

tutorialsarxiv-cs-lg
28 May 2026
Model Releases

llm-anthropic 0.25.1

DGX agent

Release: llm-anthropic 0.25.1 New model: Claude Opus 4.8 (claude-opus-4.8). New -o fast 1 option for fast mode, for organizations with that feature enabled on their account. Default max_tokens for eac

model-releasessimon-willison
28 May 2026
Model Releases

ProgVLA: Progress-Aware Robot Manipulation Skill Learning

DGX agent

arXiv:2605.28231v1 Announce Type: cross Abstract: We present ProgVLA, a compact vision-language-action (VLA) model designed for reliable robot manipulation under tight compute and memory budgets. The

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents

DGX agent

arXiv:2605.28122v1 Announce Type: cross Abstract: A coding agent executes a benign task as a sequence of shell, file, and network actions, any of which can quietly exceed the authorized scope while th

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling

DGX agent

arXiv:2605.28179v1 Announce Type: new Abstract: Scaling laws guide large language model training by relating compute to cross-entropy loss, and recent work further extends them to predict downstream b

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and Language

DGX agent

arXiv:2605.27886v1 Announce Type: new Abstract: Tactile sensing is essential for robots to achieve human-like gentle manipulation. However, existing Vision-Language-Action (VLA) models struggle to exp

model-releasesarxiv-cs-ro
28 May 2026
Research

The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling

DGX agent

arXiv:2602.14862v2 Announce Type: replace-cross Abstract: Temperature scaling is a simple method that allows to control the uncertainty of probabilistic models. It is mostly used in two contexts: impr

researcharxiv-cs-ai
28 May 2026
Model Releases

UniMaia: Steering Chess Policies with Language for Human-like Play

DGX agent

arXiv:2605.27767v1 Announce Type: cross Abstract: Recent advances in large language models have enabled natural language to serve as a flexible interface for controlling complex systems, but often at

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

When do complex-valued neural networks help? A study of representation, geometry, and optimization

DGX agent

arXiv:2605.27673v1 Announce Type: new Abstract: Complex-valued Neural Networks (CVNNs) are often motivated by domains where information is naturally encoded in magnitude and phase. Yet complex-valued

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay

DGX agent

arXiv:2605.28069v1 Announce Type: new Abstract: Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression metho

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection

DGX agent

arXiv:2605.26533v1 Announce Type: cross Abstract: Automated industrial inspection requires both precise defect localization and structured maintenance report generation; in current practice these task

model-releasesarxiv-cs-ai
27 May 2026
Safety

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

DGX agent

arXiv:2605.26174v1 Announce Type: cross Abstract: Production language-model systems answer a request by partitioning it across an invisible orchestration of worker agents that recompose one integrated

safetyarxiv-cs-ai
27 May 2026
Model Releases

Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

DGX agent

arXiv:2605.26772v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may compl

model-releasesarxiv-cs-ai
27 May 2026
Research

Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens

DGX agent

arXiv:2505.13775v4 Announce Type: replace-cross Abstract: Recent impressive results from large reasoning models have been interpreted as a triumph of Chain of Thought (CoT), and especially of the proc

researcharxiv-cs-ai
27 May 2026
Model Releases

ChartAct: A Benchmark for Dynamic Chart Understanding

DGX agent

arXiv:2605.26994v1 Announce Type: new Abstract: Charts are widely used to present complex data for analysis and decision making. Existing chart understanding benchmarks mainly focus on static charts,

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

ConVer: Using Contracts and Loop Invariant Synthesis for Scalable Formal Software Verification

DGX agent

arXiv:2605.27051v1 Announce Type: cross Abstract: Formal verification of large C programs is impeded by state-space explosion: Bounded Model Checking (BMC) tools must encode the entire state space up

model-releasesarxiv-cs-ai
27 May 2026
Tutorials

Demystifying Video Reasoning

DGX agent

arXiv:2603.16870v2 Announce Type: replace-cross Abstract: Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabil

tutorialsarxiv-cs-ai
27 May 2026
Model Releases

Entropy Sentinel: Continuous LLM Accuracy Monitoring from Decoding Entropy Traces in STEM

DGX agent

arXiv:2601.09001v4 Announce Type: replace Abstract: Deploying LLMs raises two coupled challenges: (1) monitoring--estimating where a model underperforms as traffic and domains drift--and (2) improveme

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

'Give Me BF16 or Give Me Death'? Accuracy-Performance Trade-Offs in LLM Quantization

DGX agent

arXiv:2411.02355v4 Announce Type: replace-cross Abstract: Quantization is a powerful tool for accelerating large language model (LLM) inference, but the accuracy-performance trade-offs across differen

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training

DGX agent

arXiv:2602.02518v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly rely on external knowledge to improve factuality, yet many real-world knowledge sources are organize

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Memory Architectures for Multi-Turn Text-to-SQL: A Benchmark and Empirical Study

DGX agent

arXiv:2605.26394v1 Announce Type: new Abstract: Multi-turn Text-to-SQL is central to enterprise analytics yet remains predominantly evaluated in single-turn settings. We introduce EnterpriseMem-Bench,

model-releasesarxiv-cs-cl
27 May 2026
Research

Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning

DGX agent

arXiv:2605.26761v1 Announce Type: new Abstract: Multimodal instruction tuning is the de facto recipe for adapting vision language models (VLMs), yet instruction data are highly redundant, making data

researcharxiv-cs-cv
27 May 2026
Model Releases

PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions

DGX agent

arXiv:2605.27015v1 Announce Type: new Abstract: Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We in

model-releasesarxiv-cs-cl
27 May 2026
← Previous
1…411412413414415…1371
Next →