AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
28 Jul 2026

SpecAHD: Localize to Specialize for Automated Heuristic Design in Large-Scale Routing Problems

Local AiDGX agent

arXiv:2607.23676v1 Announce Type: new Abstract: LLM-based automated heuristic design (AHD) typically scores executable programs on complete instances or within fixed solver components. In large-scale

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving

AgentsDGX agent

arXiv:2607.23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces

Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks

SafetyDGX agent

arXiv:2607.22758v1 Announce Type: cross Abstract: The integration of iterative LLMs within multi-agent diagnostic frameworks requires a rigorous quantitative reevaluation of underlying communication t


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Speed Reading Tool Powered by Artificial Intelligence for Students with ADHD, Dyslexia, and Short Attention Span

Model ReleasesDGX agent

arXiv:2307.14544v2 Announce Type: replace-cross Abstract: This paper presents an artificial intelligence tool designed to assist students with dyslexia, ADHD, and short attention spans in processing t

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

Model ReleasesDGX agent

arXiv:2607.23123v1 Announce Type: new Abstract: Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced

Stability of AI Governance Systems: A Coupled Dynamics Model of Public Trust and Social Disruptions

Model ReleasesDGX agent

arXiv:2603.20248v2 Announce Type: replace-cross Abstract: AI systems are increasingly entrenched in public governance, yet scholarship lacks formal tools to determine when deviations of public trust i

STAIF: A Stage-wise Optimization for Complex Instruction Following

SafetyDGX agent

arXiv:2607.22649v1 Announce Type: new Abstract: Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment m

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

Model ReleasesDGX agent

arXiv:2607.22658v1 Announce Type: new Abstract: Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain l

StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting

Model ReleasesDGX agent

arXiv:2607.24191v1 Announce Type: cross Abstract: Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key l

Statistically Supported LLM Ingredient and Recipe Data Collection in Computational Nutrition

ResearchDGX agent

arXiv:2607.23273v1 Announce Type: cross Abstract: Computational nutrition needs precise ingredient data, but current databases are incomplete, inconsistent, and built for human reference rather than a

Steerable Chatbots: Exploring Personalization Control Interfaces via LLM Activation Steering

ResearchDGX agent

arXiv:2505.04260v3 Announce Type: replace-cross Abstract: Personalizing LLM responses typically requires users to articulate their preferences through prompting, which can be burdensome at cold start

Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls

Model ReleasesDGX agent

arXiv:2607.24519v1 Announce Type: cross Abstract: Pretrained EEG foundation models are increasingly proposed for clinical decoding, but their transfer across populations and robustness to negative con

Stress-testing large language model agents in a robotic chemistry laboratory

AgentsDGX agent

arXiv:2607.23045v1 Announce Type: new Abstract: AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. He

Structural Preservation Governs Data Augmentation in Deep Learning-Based Laser Speckle Material Classification

ResearchDGX agent

arXiv:2607.22725v1 Announce Type: cross Abstract: Data augmentation is routinely used to improve generalization in image classification, but the assumptions underlying standard policies are poorly mat

Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning

TutorialsDGX agent

arXiv:2607.23077v1 Announce Type: new Abstract: Modeling multi-entity temporal data requires capturing dependencies across entities, time, and their interactions. Transformer-based approaches perform

Structure Over Scale: Schema-Constrained Causal Graphs for RAG

ResearchDGX agent

arXiv:2607.22592v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (GraphRAG) grounds answers in structured knowledge, but current systems extract entities and relationships ex

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.24054v1 Announce Type: new Abstract: A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes i

SwitchBraidNet: Quantisation-Aware Lightweight Architecture for Hybrid Brain-Computer Interface

ResearchDGX agent

arXiv:2606.18816v2 Announce Type: replace-cross Abstract: Hybrid brain-computer interfaces (BCIs) that integrate motor imagery (MI) and steady-state visual evoked potentials (SSVEP) provide high-dimen

SymStep: Symbolic Step Verification for Logical Reasoning

Model ReleasesDGX agent

arXiv:2607.23055v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

SafetyDGX agent

arXiv:2607.22563v1 Announce Type: new Abstract: Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. H

SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding

SafetyDGX agent

arXiv:2607.23991v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirements. However,

TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning

AgentsDGX agent

arXiv:2509.06278v4 Announce Type: replace Abstract: Table reasoning requires models to jointly perform comprehensive semantic understanding and precise numerical operations. Although recent large lang

Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models

Model ReleasesDGX agent

arXiv:2607.23976v1 Announce Type: cross Abstract: Appending a two-word confirmation tag to a decision question -- 'Is X the better choice?' versus 'X is the better choice, right?' -- changes whether a

Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis

ResearchDGX agent

arXiv:2607.24539v1 Announce Type: new Abstract: Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish

Teacher Knows It Best: Spontaneous Symmetry Breaking and Tipping Points in Networked Langevin Dynamics AI Sycophancy

ResearchDGX agent

arXiv:2607.24304v1 Announce Type: cross Abstract: We formulate a statistical physics framework to model a networked stochastic dynamical system exhibiting bistability, driven by additive noise and soc

Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models

ResearchDGX agent

arXiv:2607.22575v1 Announce Type: new Abstract: Human episodic memory supports the retrieval of experiences that unfold over extended timescales, yet the computational mechanisms underlying this abili

Test-Time Coverage: Test-Conditioned Data Curation for Deployment-Aware Learning

AgentsDGX agent

arXiv:2607.22697v1 Announce Type: new Abstract: Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution. However, st

TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2

Model ReleasesDGX agent

arXiv:2606.19259v2 Announce Type: replace-cross Abstract: Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation model

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

Model ReleasesDGX agent

arXiv:2607.24063v1 Announce Type: new Abstract: On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is tow

The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs

ResearchDGX agent

arXiv:2602.15491v2 Announce Type: replace-cross Abstract: Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly with

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

Model ReleasesDGX agent

arXiv:2607.24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked

The Illusion of Secure LLM Code: Closing the Security Gap via Iterative Reprompting

ApplicationsDGX agent

arXiv:2607.23710v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly integrated into software development workflows, yet their ability to autonomously generate secure authen

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

SafetyDGX agent

arXiv:2607.24720v1 Announce Type: cross Abstract: Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are tra

The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.22585v1 Announce Type: new Abstract: Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools,

The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing

ApplicationsDGX agent

arXiv:2607.24396v1 Announce Type: cross Abstract: In deep learning, efficiency gets more and more important to compensate for the ongoing growth in model sizes and applications. Neuromorphic hardware

The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages

Model ReleasesDGX agent

arXiv:2607.24276v1 Announce Type: cross Abstract: Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are tr

The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding

Local AiDGX agent

arXiv:2607.24570v1 Announce Type: cross Abstract: Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violations occur

Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models

Model ReleasesDGX agent

arXiv:2607.23054v1 Announce Type: cross Abstract: Multi-head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key-value pairs through a shared low-rank bottleneck (cKV), achieving 81% KV-

TLA^{+}-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA+ Specification Generation

Model ReleasesDGX agent

arXiv:2607.23425v1 Announce Type: cross Abstract: Large language models increasingly write TLA^{+} formal specifications from natural-language descriptions, but progress is hard to measure: existing r

Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

ResearchDGX agent

arXiv:2607.23493v1 Announce Type: cross Abstract: Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in

Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations

Model ReleasesDGX agent

arXiv:2607.22610v1 Announce Type: new Abstract: When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those dependencies

TokenMem: Faithful Knowledge Injection for Frozen LLMs

Model ReleasesDGX agent

arXiv:2607.22625v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge, but suffers from knowledge conflicts: when retrieved

Too much evidence, too little time: From text to actionable recommendations through multi-objective evidence reasoning

ResearchDGX agent

arXiv:2607.22574v1 Announce Type: new Abstract: Evidence-based clinical decision making requires specialists to identify, evaluate and synthesize relevant scientific literature. However, PubMed search

TopoFE: topology-aware LLM-guided Automated Feature Engineering

ResearchDGX agent

arXiv:2607.23286v1 Announce Type: new Abstract: Automatic feature engineering (AutoFE) for tabular learning can be naturally formulated as a program synthesis problem, where the objective is to discov

Towards High-Level Semantic Intelligence

ResearchDGX agent

arXiv:2607.24082v1 Announce Type: new Abstract: Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development

Towards Nexus-Score: Metadata Gaps Limit Scholarly AI Attribution

TutorialsDGX agent

arXiv:2607.22684v1 Announce Type: cross Abstract: Artificial intelligence systems increasingly mediate how science is found and credited. We asked whether missing metadata prevents AI systems from cre

Towards simultaneous decoding of kinetic and kinematic movement parameters during grasp and lift task by noninvasive brain imaging

Model ReleasesDGX agent

arXiv:2607.24081v1 Announce Type: cross Abstract: Brain-machine interfaces (BMIs) can assist individuals with limited mobility, such as stroke survivors or amputees. One of the key challenges in devel

TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs

ApplicationsDGX agent

arXiv:2607.22639v1 Announce Type: new Abstract: Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via co

TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs

SafetyDGX agent

arXiv:2607.24563v1 Announce Type: new Abstract: Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outputs remain fa

Traceable LLM Reasoning for Fake-Order Fraud Detection

ApplicationsDGX agent

arXiv:2607.23075v1 Announce Type: cross Abstract: Detecting fake-order fraud at scale remains a critical challenge for large online-to-offline (O2O) service platforms, as existing approaches often rel

Training Language Models to Cooperate with Inference-Time Controllers

Local AiDGX agent

arXiv:2607.23771v1 Announce Type: new Abstract: Large language model (LLM) performance increasingly depends not only on the base model, but also on the inference-time controller used to organize reaso

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

SafetyDGX agent

arXiv:2607.23333v1 Announce Type: cross Abstract: We revisit the regret loss framework introduced in Park et al. (2025), which uses decision-theoretic regret as a direct loss function for training mod

TRE: Training-Free Hallucination Detection for Diffusion Language Models

Model ReleasesDGX agent

arXiv:2607.22661v1 Announce Type: new Abstract: Diffusion large language models (D-LLMs) have recently gained increasing attention, yet their reliability is significantly hindered by the hallucination

TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.23838v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base at query t

TriSP: Tri-Signal Structured Pruning for Large Language Models

Model ReleasesDGX agent

arXiv:2607.22587v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance across diverse tasks but their deployment is constrained by the memory and compute cost of their

Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search

ResearchDGX agent

arXiv:2607.23854v1 Announce Type: new Abstract: Humans often find good solutions to combinatorial optimization problems that are computationally hard even for advanced computer algorithms. In the Eucl

Understanding Machine Unlearning Through the Lens of Mode Connectivity

Model ReleasesDGX agent

arXiv:2607.23970v1 Announce Type: cross Abstract: Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress, the loss la

Understanding Tone-Dependent Inference Cost in Large Language Models

Model ReleasesDGX agent

arXiv:2607.23915v1 Announce Type: cross Abstract: We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were perf

Unequal Trips, Unequal Places: Diagnosing and Mitigating Delay Inequity in Autonomous Vehicle Fleet Coordination

SafetyDGX agent

arXiv:2607.24336v1 Announce Type: new Abstract: City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is distributed

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

Model ReleasesDGX agent

arXiv:2512.24125v3 Announce Type: replace-cross Abstract: General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action executio

← Previous
1…5253545556…354
Next →