AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

86,993Total entries
1Added by human
86,992Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,477 results
11 Aug 2026

ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

Model ReleasesDGX agent

arXiv:2608.09476v1 Announce Type: cross Abstract: Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavio

ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection

Model ReleasesDGX agent

arXiv:2608.09789v1 Announce Type: new Abstract: Industrial anomaly detection (IAD) requires identifying fine-grained deviations from normal visual patterns. Multimodal large language models (MLLMs) ca

ASPaeroFlow: Decomposition Heuristics for Joint Air Traffic Flow & Capacity Management

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.09315v1 Announce Type: new Abstract: While mathematical models act as vital decision support systems for operational Air Traffic Flow and Capacity Management (ATFCM), existing approaches is

Backward Compatibility in Tree-Based Explanations and Enhanced CART Algorithm

ApplicationsDGX agent

arXiv:2608.08674v1 Announce Type: new Abstract: In the operation of machine learning models, model update is a fundamental process that requires careful consideration of its impact on downstream decis

Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs

ResearchDGX agent

arXiv:2608.09344v1 Announce Type: new Abstract: Recent advances in large vision-language models (LVLMs) have enabled powerful multimodal reasoning by integrating visual encoders with large language mo

BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

Model ReleasesDGX agent

arXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

Model ReleasesDGX agent

arXiv:2608.09555v1 Announce Type: new Abstract: External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Yet their effe

CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation

Model ReleasesDGX agent

arXiv:2608.09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, sup

Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents

ResearchDGX agent

arXiv:2608.09485v1 Announce Type: new Abstract: Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission,

CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits

Model ReleasesDGX agent

arXiv:2608.09374v1 Announce Type: new Abstract: Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, sel

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

Model ReleasesDGX agent

arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investiga

Closing the loop in learning with missing data

Model ReleasesDGX agent

arXiv:2608.09030v1 Announce Type: cross Abstract: What should a machine learning model learn when data is missing during training? We look at the learning process from a dynamical systems perspective,

Contamination Means Overestimation? A Fine-Grained Empirical Study in Code Intelligence

Model ReleasesDGX agent

arXiv:2506.02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning

Model ReleasesDGX agent

arXiv:2608.09324v1 Announce Type: new Abstract: On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding th

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Model ReleasesDGX agent

arXiv:2608.09766v1 Announce Type: cross Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evalu

Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE

Local AiDGX agent

arXiv:2608.08032v1 Announce Type: new Abstract: Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in

Design Space of Self--Consistent Electrostatic Machine Learning Interatomic Potentials

HardwareDGX agent

arXiv:2603.14700v2 Announce Type: replace-cross Abstract: Machine learning interatomic potentials (MLIPs) have become widely used tools in atomistic simulations. For much of the history of this field,

DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference

Model ReleasesDGX agent

arXiv:2608.08878v1 Announce Type: cross Abstract: Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequen

DocAtlas: Long-Document Understanding as Mutable-State Interaction

Model ReleasesDGX agent

arXiv:2608.07527v1 Announce Type: cross Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-a

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

ResearchDGX agent

arXiv:2608.08002v1 Announce Type: new Abstract: Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response q

FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search

Model ReleasesDGX agent

arXiv:2608.09474v1 Announce Type: new Abstract: Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained prima

Financial Numerical Prediction and Allocation as Token Generation

SafetyDGX agent

arXiv:2608.09880v1 Announce Type: new Abstract: Financial prediction typically relies on task-specific regression, ranking, or policy heads, separating the language model from the numerical object ult

ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration

Model ReleasesDGX agent

arXiv:2608.08605v1 Announce Type: new Abstract: Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common ba

From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings

Model ReleasesDGX agent

arXiv:2608.08896v1 Announce Type: new Abstract: Imaging device downtime is a major barrier to healthcare delivery in low- and middle-income countries (LMICs), often driven by limited access to special

GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis

Model ReleasesDGX agent

arXiv:2608.09921v1 Announce Type: new Abstract: Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from engineering domains such as power s

Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification

Model ReleasesDGX agent

arXiv:2603.19329v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can generate plausible code but offer limited guarantees of correctness. Formally verifying that implementations

Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses

Model ReleasesDGX agent

arXiv:2608.08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the harne

High-Layer Attention Pruning with Rescaling

Model ReleasesDGX agent

arXiv:2507.01900v3 Announce Type: replace Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional

Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation

Model ReleasesDGX agent

arXiv:2608.07498v1 Announce Type: cross Abstract: Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deploym

LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

Model ReleasesDGX agent

arXiv:2503.19990v4 Announce Type: replace Abstract: Many real-world applications of spatial intelligence, such as robotic control, autonomous driving, and automated assembly, require spatial reasoning

LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4

Model ReleasesDGX agent

arXiv:2607.15509v2 Announce Type: replace-cross Abstract: We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture desig

MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks

Model ReleasesDGX agent

arXiv:2507.03162v2 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These mod

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

Model ReleasesDGX agent

arXiv:2608.09046v1 Announce Type: new Abstract: Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure doe

MixFormer: Linear Transformer with Mixture of Memory Experts

ResearchDGX agent

arXiv:2608.09468v1 Announce Type: cross Abstract: State Space Models (SSMs), as a mainstream research direction of linear Transformers, aim to achieve higher efficiency than standard Transformers in l

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

Model ReleasesDGX agent

arXiv:2608.07993v1 Announce Type: new Abstract: Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor moti

NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2510.18940v2 Announce Type: replace-cross Abstract: Existing parameter-efficient fine-tuning (PEFT) methods primarily fall into two categories: addition-based and selective in-situ adaptation. T

Neurosymbolic Discovery of Algebraic Graph Constructions

Model ReleasesDGX agent

arXiv:2608.08118v1 Announce Type: new Abstract: There are several methods for searching for graphs with prescribed properties, such as SAT solvers and specialized generators. These methods return the

OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents

Model ReleasesDGX agent

arXiv:2608.08264v1 Announce Type: new Abstract: Large language model agents are becoming operational interfaces to files, memories, registries, and external tools. This deployment shift creates a new

On the Robustness of LLMs' Internal Representation of Code Correctness

ResearchDGX agent

arXiv:2608.08266v1 Announce Type: cross Abstract: Code generated by modern language models often reads naturally. Yet, it also often fails to implement what was asked. This should be no surprise, as r

Optimal Learning Under Tsybakov Noise

ResearchDGX agent

arXiv:2608.08416v1 Announce Type: new Abstract: Probably Approximately Correct (PAC) learning [Val84] is a fundamental learning model that has been extensively investigated. In this model, H subseteq

P^{3}: Joint Program-and-Proof Planning for Verified Code Generation

Model ReleasesDGX agent

arXiv:2608.09277v1 Announce Type: new Abstract: Verified code generation asks a large language model (LLM) to generate both an executable program and a machine-checkable proof that the program meets a

PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

Model ReleasesDGX agent

arXiv:2608.07631v1 Announce Type: cross Abstract: LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue

Preserving Item Semantics for Free: Rethinking Token Initialization in LLM-Based Generative Recommendation

Model ReleasesDGX agent

arXiv:2608.07816v1 Announce Type: cross Abstract: Recent advances in generative recommendation (GR) leverage large language models (LLMs) as recommender backbones, enabling LLMs to directly generate r

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

Model ReleasesDGX agent

arXiv:2608.07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on e

REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation

Model ReleasesDGX agent

arXiv:2503.22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require

Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction

Model ReleasesDGX agent

arXiv:2608.09182v1 Announce Type: cross Abstract: Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Existing localiz

RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning

Model ReleasesDGX agent

arXiv:2608.09123v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a s

RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not Enough

Model ReleasesDGX agent

arXiv:2608.07583v1 Announce Type: cross Abstract: Multi-agent LLM systems route among model-backed advisors, yet a deployer rarely knows before shipping whether routing will help at all. Prevailing ro

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

Model ReleasesDGX agent

arXiv:2608.09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance,

SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision

Model ReleasesDGX agent

arXiv:2608.09097v1 Announce Type: new Abstract: Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

Model ReleasesDGX agent

arXiv:2608.09771v1 Announce Type: new Abstract: Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every

SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat

SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding

Model ReleasesDGX agent

arXiv:2608.07915v1 Announce Type: new Abstract: Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Th

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

Model ReleasesDGX agent

arXiv:2608.09538v1 Announce Type: cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation.

Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law

Model ReleasesDGX agent

arXiv:2608.09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the ap

TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

Model ReleasesDGX agent

arXiv:2608.07617v1 Announce Type: new Abstract: Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatche

The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction

Model ReleasesDGX agent

arXiv:2608.07517v1 Announce Type: cross Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aw

The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games

Model ReleasesDGX agent

arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move fro

Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

Model ReleasesDGX agent

arXiv:2608.08795v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external o

UNSPECIFIC: General Constraint Synthesis for Breaking Copy-and-Paste Shortcut in LLM Instruction Following

Model ReleasesDGX agent

arXiv:2608.09154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly expected to follow long lists of constraints in complex instructions, and synthesizing instructions from a

← Previous
1…327328329330331…1042
Next →