AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
29 May 2026

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

ResearchDGX agent

arXiv:2605.29800v1 Announce Type: new Abstract: LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a frame

On-Policy Replay for Continual Supervised Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.29495v1 Announce Type: new Abstract: Continual supervised fine-tuning (SFT) is the de facto recipe for adapting large language models (LLMs) to a stream of downstream tasks, but it suffers

OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation

Model ReleasesDGX agent

arXiv:2605.29829v1 Announce Type: new Abstract: Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient par

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Personalized Turn-Level User Conversation Satisfaction Benchmark

Model ReleasesDGX agent

arXiv:2605.29711v1 Announce Type: cross Abstract: User satisfaction with AI assistants is highly personalized: the same response may satisfy one user but disappoint another depending on what each user

PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

Model ReleasesDGX agent

arXiv:2605.29815v1 Announce Type: new Abstract: The growing number of submitted papers has motivated the exploration of Large Language Models (LLMs) as a means to support and augment the peer review p

ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure

Model ReleasesDGX agent

arXiv:2605.30284v1 Announce Type: new Abstract: Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks ha

ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation

Model ReleasesDGX agent

arXiv:2605.29579v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) have achieved rapid progress in vision-language understanding, they remain prone to multimodal hallucinat

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

Model ReleasesDGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

Structure-Aware Text Recognition for Ancient Greek Critical Editions

Model ReleasesDGX agent

arXiv:2603.02803v2 Announce Type: replace Abstract: Recent advances in visual language models (VLMs) have transformed end-to-end document understanding. However, their ability to interpret the complex

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2602.15382v2 Announce Type: replace Abstract: Multi-Agent Systems (MAS) powered by Large Language Models have unlocked advanced collaborative reasoning, yet they remain bottlenecked by discrete

Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection

Model ReleasesDGX agent

arXiv:2605.30189v1 Announce Type: cross Abstract: We show that LoRA adapters, the dominant distribution format for fine-tuned LLMs, can be reliably backdoored through training data poisoning while pre

Toward Ethical Facial Age Estimation: A Generalized Zero-Shot Benchmark Without Training on Children's Data

Model ReleasesDGX agent

arXiv:2605.29230v1 Announce Type: cross Abstract: Age estimation from facial images typically relies on training data that includes images of minors, a practice that raises serious ethical, legal, and

Training Deliberative Monitors for Black-Box Scheming Detection

Model ReleasesDGX agent

arXiv:2605.29601v1 Announce Type: cross Abstract: As autonomous agents become more capable of performing real-world tasks, distinguishing scheming behavior from benign task pursuit may become a centra

Who can we trust? LLM-as-a-jury for Comparative Assessment

Model ReleasesDGX agent

arXiv:2602.16610v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly applied as automatic evaluators for natural language generation assessment often using pairwise

28 May 2026

A Query Engine for the Agents

Model ReleasesDGX agent

arXiv:2605.27785v1 Announce Type: new Abstract: The fastest-growing data in production today is unstructured text: agent traces, chat logs, reasoning chains, model outputs. People want to analyze it,

AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios

Model ReleasesDGX agent

arXiv:2605.27995v1 Announce Type: new Abstract: Large language model (LLM)-based agents have shown strong capabilities in using external tools to solve complex tasks. However, existing evaluations oft

Automating Formal Verification with Agent-Guided Tree Search

Model ReleasesDGX agent

arXiv:2605.27485v1 Announce Type: cross Abstract: Formal verification offers a path to provably correct software, but writing verified code remains expensive enough that the technique is rarely used i

Behavioural Analysis of Alignment Faking

SafetyDGX agent

arXiv:2605.27681v1 Announce Type: new Abstract: Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deploym

Bounded-Compute Multimodal Regression for Product-Rating Prediction

Model ReleasesDGX agent

arXiv:2605.27737v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly attractive for multimodal quality assessment, but their default reliance on autoregressive text generatio

Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback

Model ReleasesDGX agent

arXiv:2605.28010v1 Announce Type: new Abstract: Self-evolving large language models (LLMs) learn by generating their own training tasks and solutions, reducing reliance on human-curated supervision. H

Data-Efficient On-Policy Distillation for Automatic Speech Recognition

Model ReleasesDGX agent

arXiv:2605.28139v1 Announce Type: new Abstract: Building competitive automatic speech recognition (ASR) models usually requires large-scale au- dio supervision, which makes reproduction and specializa

Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024

Model ReleasesDGX agent

arXiv:2503.02857v5 Announce Type: replace-cross Abstract: In the age of increasingly realistic generative AI, robust deepfake detection is essential for mitigating fraud and disinformation. While many

DEPART: DEcomposing PARiTy across Multilingual LLMs

Model ReleasesDGX agent

arXiv:2605.28163v1 Announce Type: cross Abstract: Multilingual Large Language Models (mLLMs) leaderboards report per-language accuracy but rarely explain why disparities emerge, leaving systemic biase

Do LLMs Favor Their Providers? Measuring Vertical Integration Bias in Code Generation

Model ReleasesDGX agent

arXiv:2605.28515v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become an integral part of software development, especially with the advent of agentic capabilities. Yet, many front

EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter AdaptationTarget

Model ReleasesDGX agent

arXiv:2605.27390v1 Announce Type: cross Abstract: Speculative decoding accelerates Large Language Model inference via a draft-then-verify paradigm, yet the output projection layer becomes a bottleneck

FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning

SafetyDGX agent

arXiv:2605.28389v1 Announce Type: new Abstract: While large language models have made significant progress in mathematical reasoning, they remain unreliable at judging the correctness of their own sol

Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment

Model ReleasesDGX agent

arXiv:2605.28188v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making settings such as legal reasoning, where consistency under factuall

Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification

Model ReleasesDGX agent

arXiv:2605.28500v1 Announce Type: cross Abstract: Large language models have shown impressive capabilities in code generation, yet they often produce functionally incorrect code. Uncertainty quantific

KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

Model ReleasesDGX agent

arXiv:2605.28013v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision.

LACUNA: Safe Agents as Recursive Program Holes

SafetyDGX agent

arXiv:2605.28617v1 Announce Type: new Abstract: LLM agents increasingly act by writing code, yet a split persists between the runtime that drives the agent and the code the model writes. The runtime o

Learn from your own latents and not from tokens: A sample-complexity theory

TutorialsDGX agent

arXiv:2605.27734v1 Announce Type: new Abstract: Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude lar

llm-anthropic 0.25.1

Model ReleasesDGX agent

Release: llm-anthropic 0.25.1 New model: Claude Opus 4.8 (claude-opus-4.8). New -o fast 1 option for fast mode, for organizations with that feature enabled on their account. Default max_tokens for eac

ProgVLA: Progress-Aware Robot Manipulation Skill Learning

Model ReleasesDGX agent

arXiv:2605.28231v1 Announce Type: cross Abstract: We present ProgVLA, a compact vision-language-action (VLA) model designed for reliable robot manipulation under tight compute and memory budgets. The

SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents

Model ReleasesDGX agent

arXiv:2605.28122v1 Announce Type: cross Abstract: A coding agent executes a benign task as a sequence of shell, file, and network actions, any of which can quietly exceed the authorized scope while th

SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling

Model ReleasesDGX agent

arXiv:2605.28179v1 Announce Type: new Abstract: Scaling laws guide large language model training by relating compute to cross-entropy loss, and recent work further extends them to predict downstream b

Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and Language

Model ReleasesDGX agent

arXiv:2605.27886v1 Announce Type: new Abstract: Tactile sensing is essential for robots to achieve human-like gentle manipulation. However, existing Vision-Language-Action (VLA) models struggle to exp

The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling

ResearchDGX agent

arXiv:2602.14862v2 Announce Type: replace-cross Abstract: Temperature scaling is a simple method that allows to control the uncertainty of probabilistic models. It is mostly used in two contexts: impr

UniMaia: Steering Chess Policies with Language for Human-like Play

Model ReleasesDGX agent

arXiv:2605.27767v1 Announce Type: cross Abstract: Recent advances in large language models have enabled natural language to serve as a flexible interface for controlling complex systems, but often at

When do complex-valued neural networks help? A study of representation, geometry, and optimization

Model ReleasesDGX agent

arXiv:2605.27673v1 Announce Type: new Abstract: Complex-valued Neural Networks (CVNNs) are often motivated by domains where information is naturally encoded in magnitude and phase. Yet complex-valued

ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay

Model ReleasesDGX agent

arXiv:2605.28069v1 Announce Type: new Abstract: Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression metho

27 May 2026

A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection

Model ReleasesDGX agent

arXiv:2605.26533v1 Announce Type: cross Abstract: Automated industrial inspection requires both precise defect localization and structured maintenance report generation; in current practice these task

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

SafetyDGX agent

arXiv:2605.26174v1 Announce Type: cross Abstract: Production language-model systems answer a request by partitioning it across an invisible orchestration of worker agents that recompose one integrated

Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

Model ReleasesDGX agent

arXiv:2605.26772v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may compl

Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens

ResearchDGX agent

arXiv:2505.13775v4 Announce Type: replace-cross Abstract: Recent impressive results from large reasoning models have been interpreted as a triumph of Chain of Thought (CoT), and especially of the proc

ChartAct: A Benchmark for Dynamic Chart Understanding

Model ReleasesDGX agent

arXiv:2605.26994v1 Announce Type: new Abstract: Charts are widely used to present complex data for analysis and decision making. Existing chart understanding benchmarks mainly focus on static charts,

ConVer: Using Contracts and Loop Invariant Synthesis for Scalable Formal Software Verification

Model ReleasesDGX agent

arXiv:2605.27051v1 Announce Type: cross Abstract: Formal verification of large C programs is impeded by state-space explosion: Bounded Model Checking (BMC) tools must encode the entire state space up

Demystifying Video Reasoning

TutorialsDGX agent

arXiv:2603.16870v2 Announce Type: replace-cross Abstract: Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabil

Entropy Sentinel: Continuous LLM Accuracy Monitoring from Decoding Entropy Traces in STEM

Model ReleasesDGX agent

arXiv:2601.09001v4 Announce Type: replace Abstract: Deploying LLMs raises two coupled challenges: (1) monitoring--estimating where a model underperforms as traffic and domains drift--and (2) improveme

'Give Me BF16 or Give Me Death'? Accuracy-Performance Trade-Offs in LLM Quantization

Model ReleasesDGX agent

arXiv:2411.02355v4 Announce Type: replace-cross Abstract: Quantization is a powerful tool for accelerating large language model (LLM) inference, but the accuracy-performance trade-offs across differen

GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training

Model ReleasesDGX agent

arXiv:2602.02518v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly rely on external knowledge to improve factuality, yet many real-world knowledge sources are organize

Memory Architectures for Multi-Turn Text-to-SQL: A Benchmark and Empirical Study

Model ReleasesDGX agent

arXiv:2605.26394v1 Announce Type: new Abstract: Multi-turn Text-to-SQL is central to enterprise analytics yet remains predominantly evaluated in single-turn settings. We introduce EnterpriseMem-Bench,

Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning

ResearchDGX agent

arXiv:2605.26761v1 Announce Type: new Abstract: Multimodal instruction tuning is the de facto recipe for adapting vision language models (VLMs), yet instruction data are highly redundant, making data

PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions

Model ReleasesDGX agent

arXiv:2605.27015v1 Announce Type: new Abstract: Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We in

Qiskit QuantumKatas: Adapting Microsoft's Quantum Computing exercises for LLM evaluation

Model ReleasesDGX agent

arXiv:2605.27210v1 Announce Type: cross Abstract: We adapt Microsoft's QuantumKatas -- a well-established quantum computing curriculum -- from Q# to Qiskit, the most widely-adopted quantum computing f

ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference

Model ReleasesDGX agent

arXiv:2605.27081v1 Announce Type: cross Abstract: Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining hi

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure

Model ReleasesDGX agent

arXiv:2601.10566v5 Announce Type: replace Abstract: Entity-level unlearning is usually evaluated by what a model says: whether it stops naming the target, refuses a query, or shifts a Truth Ratio dist

Rethinking the Multilingual Reasoning Gap with Layer Swap

Model ReleasesDGX agent

arXiv:2605.26735v1 Announce Type: new Abstract: Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior wor

Separating Semantic Competition from Context Length in RAG Reading

Model ReleasesDGX agent

arXiv:2605.27294v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems can respond incorrectly even when the correct passage was retrieved. The model must still read the retrieve

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

Model ReleasesDGX agent

arXiv:2510.09606v2 Announce Type: replace Abstract: With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still strug

SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking

ResearchDGX agent

arXiv:2511.04711v2 Announce Type: replace-cross Abstract: Large-scale vision-language models, especially CLIP, have demonstrated remarkable performance across diverse downstream tasks. Soft prompts, a

← Previous
1…308309310311312…1036
Next →