AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
20,981 results
12 Aug 2026

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

Model ReleasesDGX agent

arXiv:2608.10042v1 Announce Type: cross Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool u

V-FiLLM: Verified Financial LLM Reasoning Benchmark

Model ReleasesDGX agent

arXiv:2608.11047v1 Announce Type: new Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains compar

VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection

AgentsDGX agent

arXiv:2511.19436v2 Announce Type: replace-cross Abstract: Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary mode


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus

ResearchDGX agent

arXiv:2608.10665v1 Announce Type: new Abstract: Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approache

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

Model ReleasesDGX agent

arXiv:2608.10875v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained re

What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research

SafetyDGX agent

arXiv:2608.10431v1 Announce Type: cross Abstract: Responsible AI (RAI) has become a central concern for technology companies, regulators, and the public. How industry practitioners interpret, implemen

When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting

AgentsDGX agent

arXiv:2606.16465v2 Announce Type: replace Abstract: AI agents can now take irreversible actions in operational systems, but agent-caused losses are still not clearly assigned, priced, or transferred.

When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

Model ReleasesDGX agent

arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of

Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition

ResearchDGX agent

arXiv:2608.10836v1 Announce Type: cross Abstract: The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory tr

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

Model ReleasesDGX agent

arXiv:2608.11095v1 Announce Type: new Abstract: Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wh

Withholding the Completing Chunk: Deterministic Pair-Completion Guardrails for Streaming LLM Output

Model ReleasesDGX agent

arXiv:2608.10279v1 Announce Type: cross Abstract: Streaming language-model output creates a release-timing problem: complete-response moderation acts after streamed text has escaped, whereas repeated

Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data

Model ReleasesDGX agent

arXiv:2608.11022v1 Announce Type: cross Abstract: Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their con

XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving

SafetyDGX agent

arXiv:2608.10976v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verb

'YES! YES! I absolutely love this insight!' Affirmative Narration as Interactional Strategy in Dialogues with LLM Chatbots

ResearchDGX agent

arXiv:2607.28646v2 Announce Type: cross Abstract: This article analyses narrative mechanisms that are common in dialogues with LLM chatbots. In combination, these mechanisms produce an interactional s

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

SafetyDGX agent

arXiv:2608.10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream dec

11 Aug 2026

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

Model ReleasesDGX agent

arXiv:2608.08814v1 Announce Type: cross Abstract: We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment construc

A Combined Feature-Based Framework for Disguise and Spoofing Detection in Face Recognition Systems

ResearchDGX agent

arXiv:2608.08521v1 Announce Type: cross Abstract: Face recognition systems face two distinct, commonly-separated failure modes: spoofing, where an impostor presents a photograph or video of an authori

A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences

Model ReleasesDGX agent

arXiv:2608.08240v1 Announce Type: new Abstract: This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage

A foundation model of numerical intelligence with cross-disciplinary generalization

ResearchDGX agent

arXiv:2607.28432v2 Announce Type: replace Abstract: Intelligence is commonly understood as the ability to acquire and apply knowledge, adapt to unfamiliar situations and solve new problems. Large lang

A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization

AgentsDGX agent

arXiv:2608.08180v1 Announce Type: cross Abstract: Abstractive text summarization systems frequently generate fluent yet unfaithful summaries by fabricating or distorting relationships between entities

A Minimal kappa--au Logic for Risk-Sensitive Abduction

ResearchDGX agent

arXiv:2608.08192v1 Announce Type: new Abstract: Standard approaches to abductive reasoning can retain multiple candidate explanations, but they do not generally combine explicit compositional cross-hy

A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition

ResearchDGX agent

arXiv:2608.09088v1 Announce Type: new Abstract: Mixed emotions represent a clinically relevant but still underexplored target for automatic emotion recognition. EEG provides millisecond-level access t

A New Approach to Characterising Optimisation Problems Using Programmatic Representation and Complexity Measures

ResearchDGX agent

arXiv:2608.08898v1 Announce Type: cross Abstract: Characterising optimisation problem instances is a fundamental part of understanding the behaviour and performance of different algorithms as well as

A QUBO-Inspired Computational Framework for Airport Landside Bottleneck Diagnosis and Dynamic Dispatch Optimization

ResearchDGX agent

arXiv:2608.08632v1 Announce Type: new Abstract: Airport landside traffic centers connect terminal arrivals with taxis, ride-hailing vehicles, private cars, buses, metro services, parking facilities, a

A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence

Model ReleasesDGX agent

arXiv:2501.17629v2 Announce Type: replace-cross Abstract: Several studies claim that large language models have passed the Turing Test and hence can 'think', yet none follow Turing's original instruct

A Sobering Look at Tabular Data Generation via Probabilistic Circuits

ResearchDGX agent

arXiv:2603.23016v2 Announce Type: replace-cross Abstract: Tabular data is more challenging to generate than text and images, due to its heterogeneous features and much lower sample sizes. On this task

A Structural Dynamics Graph World Model: Unified Modeling, Constrained Rollout, and Interpretable Calibration

SafetyDGX agent

arXiv:2608.08689v1 Announce Type: new Abstract: The state evolution of a complex system arises jointly from object laws, relational propagation, domain conservation, and unmodeled error. Forcing all s

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning

SafetyDGX agent

arXiv:2608.08158v1 Announce Type: new Abstract: Sparse, delayed, and weakly informative rewards remain central obstacles to efficient reinforcement learning. Reward shaping addresses these limitations

A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents

Model ReleasesDGX agent

arXiv:2608.09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.

Abstracted Away: Resisting Alienation and Ungrounded Abstraction in AI Research Communities

ResearchDGX agent

arXiv:2608.08408v1 Announce Type: cross Abstract: Logics of abstraction in computational AI research often push important forms of knowledge and reflection aside: dominant standards of legitimacy sepa

ACEvo: Adversarial Co-Evolution of Problem Distributions and Solvers for Combinatorial Optimization

Model ReleasesDGX agent

arXiv:2506.02594v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to synthesize heuristic programs, yet most existing pipelines optimize solvers against fixed benc

ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

Model ReleasesDGX agent

arXiv:2608.09476v1 Announce Type: cross Abstract: Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavio

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

Model ReleasesDGX agent

arXiv:2607.10180v2 Announce Type: replace-cross Abstract: We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. T

Adaptive Semantic Capacity Allocation for Parallel Generative Recommendation

ResearchDGX agent

arXiv:2608.09685v1 Announce Type: new Abstract: Autoregressive semantic ID recommenders are constrained by expensive beam-search decoding, which limits the practical length of item identifiers. Parall

Adaptive Sequential Test Planning for Multi-Mechanism Reliability Qualification via Bayesian Monte Carlo Tree Search

SafetyDGX agent

arXiv:2608.09622v1 Announce Type: new Abstract: Reliability qualification of advanced semiconductor devices requires sequential stress decisions that balance characterization objectives against multip

Adaptive Symmetry Discovery for Dynamical System Identification

TutorialsDGX agent

arXiv:2608.08091v1 Announce Type: cross Abstract: Dynamical systems model trajectory data generated by fixed underlying dynamics, with applications ranging from biology to physics. Especially in scien

Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service Classes

TutorialsDGX agent

arXiv:2608.07747v1 Announce Type: new Abstract: We study how to share a single conserved capacity budget across many locations and two service classes when demand is uneven, time-varying, and can exce

Adversarial Attacks on Deep OCR Systems

Model ReleasesDGX agent

arXiv:2608.07636v1 Announce Type: cross Abstract: Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at l

Adversarial Latent-State Training for Robust Policies in Partially Observable Domains

Model ReleasesDGX agent

arXiv:2603.07313v4 Announce Type: replace-cross Abstract: Robustness under latent distribution shift remains challenging in partially observable reinforcement learning. We formalize a focused setting

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation

HardwareDGX agent

arXiv:2608.08469v1 Announce Type: new Abstract: Existing streaming multimodal models process observations incrementally but still follow a turn-based prefill-then-decode pattern, making them non-duple

AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization

Model ReleasesDGX agent

arXiv:2608.07557v1 Announce Type: cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist en

Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD Campaigns

AgentsDGX agent

arXiv:2608.07637v1 Announce Type: new Abstract: Long-running molecular simulation campaigns require repeated continuation from saved states, provenance-aware progression, adaptive assessment, and occa

Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) Brachytherapy

Local AiDGX agent

arXiv:2608.08163v1 Announce Type: new Abstract: The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized peda

Agentic AI for Clustering, Relationship Discovery, and Semantic Trading in Prediction Markets

Model ReleasesDGX agent

arXiv:2512.02436v2 Announce Type: replace Abstract: Prediction markets allow users to trade on outcomes of real-world events, but are prone to fragmentation with overlapping questions, implicit equiva

Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data

Model ReleasesDGX agent

arXiv:2608.08859v1 Announce Type: cross Abstract: Wireless Body Area Networks (WBANs) generate multivariate physiological time series that are highly nonstationary and must often be processed under st

Agentic Auto-Research is Fuzz Testing

AgentsDGX agent

arXiv:2608.09855v1 Announce Type: new Abstract: Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the proposer and ra

Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy

SafetyDGX agent

arXiv:2608.09857v1 Announce Type: cross Abstract: Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on

Agentic Router: An Execution-Grounded Continual Learning Approach With Memory

AgentsDGX agent

arXiv:2608.09184v1 Announce Type: new Abstract: Large language model (LLM) agents provide a promising interface for command-line-based network operations, but a plausible command may still fail or int

Agentic Stage-One Stellarator Optimization: Autonomous Multi-Objective Search for Finite-Beta Equilibria

Local AiDGX agent

arXiv:2608.01344v2 Announce Type: replace Abstract: Stage-one stellarator design searches a high-dimensional family of three-dimensional plasma boundaries and fixed-boundary MHD equilibria for configu

AI Evaluation Should Measure Verification Cost, Not Correctness Alone

Model ReleasesDGX agent

arXiv:2608.08709v1 Announce Type: new Abstract: The reliability of AI generative models is typically measured by output correctness, yet in practice it depends on the effort required to verify those o

AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting

ApplicationsDGX agent

arXiv:2608.09775v1 Announce Type: new Abstract: Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels d

AIVV: Neuro-Symbolic LLM Agent-Integrated Verification and Validation for Trustworthy Autonomous Systems

AgentsDGX agent

arXiv:2604.02478v2 Announce Type: replace Abstract: Deep learning models excel at detecting anomaly patterns in normal data. However, they do not provide a direct solution for anomaly classification a

AkasicDB: Demonstrating Omni RAG with a Unified Vector-Graph-Relational DBMS

ResearchDGX agent

arXiv:2608.09214v1 Announce Type: cross Abstract: Recent Retrieval-Augmented Generation (RAG) systems increasingly combine vector retrieval with structured knowledge, such as Graph RAG and Filtered ve

An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography

Model ReleasesDGX agent

arXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency.

An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop

Local AiDGX agent

arXiv:2608.07542v1 Announce Type: new Abstract: Autonomous research loops driven by large language models can run machine-learning experiments at scale but tend to drift toward local refinements of wh

An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism

ResearchDGX agent

arXiv:2607.15511v2 Announce Type: replace-cross Abstract: Serverless computing provides automatic resource management and pay-per-use execution, but effective autoscaling remains challenging because o

An evolutionary model of animats with VLM-based subjective evaluation

ResearchDGX agent

arXiv:2608.07537v1 Announce Type: cross Abstract: In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language Model (VLM) into the fitness evaluation a

An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning

Model ReleasesDGX agent

arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated b

An Explainable GNN Framework for Component-Level Anomaly Diagnosis

SafetyDGX agent

arXiv:2608.09246v1 Announce Type: new Abstract: Industrial processes are complex systems composed of multiple interacting sensors that generate multivariate time series (MTS). Detecting anomalies in s

AndroidReality: How Far Are Mobile Agents from the Real World?

Model ReleasesDGX agent

arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-worl

← Previous
1…34567…350
Next →