AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
24 Jul 2026

Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry

Local AiDGX agent

arXiv:2607.21495v1 Announce Type: new Abstract: AI agents are increasingly created inside organizations by non-engineering users through low-code, no-code, and conversational development environments.

Toward cryptographically verifiable authorization for autonomous AI agents: A security hypothesis, preliminary formal model, and proof-of-concept implementation

SafetyDGX agent

arXiv:2607.21325v1 Announce Type: cross Abstract: Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limited human oversight. Existing authenticat

Towards a Certifying Grounder

ResearchDGX agent

arXiv:2607.21199v1 Announce Type: cross Abstract: Grounding, the translation of high-level theories into equivalent quantifier-free formulas, is a crucial step in declarative solving, yet it has so fa


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Towards an Automated Test of LLM Security Knowledge

Model ReleasesDGX agent

arXiv:2607.18496v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM perf

Towards Faithful Graph Explanations with Synergistic Edge Effects via Granular Balls

Model ReleasesDGX agent

arXiv:2607.21381v1 Announce Type: new Abstract: Instance-level explanations aim to reveal the rationale behind a model's decisions for a specific graph. Previous methods explain graph neural networks

Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI

AgentsDGX agent

arXiv:2607.20916v1 Announce Type: new Abstract: Generative AI lets large language models produce scholarly-looking text within seconds, yet fluency does not equal valid explanation. The deepest risk i

Tractable Hierarchical Control of Autoregressive Language Models

ResearchDGX agent

arXiv:2607.20483v1 Announce Type: new Abstract: Constraining the generation of autoregressive large language models (LLMs) is an important component of integrating language models into formal systems.

Training Large Language Models for Self-Explanation Faithfulness

Model ReleasesDGX agent

arXiv:2607.21090v1 Announce Type: cross Abstract: We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated r

Transition-Related Potentials as Markers of Narrative Comprehension in Continuous EEG

ResearchDGX agent

arXiv:2607.20720v1 Announce Type: cross Abstract: Harnessing the potential of electroencephalography (EEG) for brain research is fundamentally limited by intrinsic noise and the diffuse projection of

TwistedMerge: Certified Higher-Order Diagnostics and Abstention for Model Merging

SafetyDGX agent

arXiv:2607.20887v1 Announce Type: cross Abstract: Model merging combines independently trained or fine-tuned models, but pairwise alignability does not imply globally consistent alignment. We formulat

U-CFR: Uncertainty-Guided Cascade Forward Refinement for Interactive Segmentation

Model ReleasesDGX agent

arXiv:2607.20705v1 Announce Type: cross Abstract: Interactive image segmentation is critical for efficient image annotation; however, existing methods often require many corrective clicks or rely on p

Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement

ResearchDGX agent

arXiv:2607.20529v1 Announce Type: cross Abstract: Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs. However, existing aggre

Understanding Critical Thinking in Generative Artificial Intelligence Use: Development, Validation, and Correlates of the Critical Thinking in AI Use Scale

SafetyDGX agent

arXiv:2512.12413v2 Announce Type: replace Abstract: Generative AI tools are increasingly embedded in everyday work and learning, yet their fluency, opacity, and propensity to hallucinate mean that use

Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning

Model ReleasesDGX agent

arXiv:2607.21300v1 Announce Type: cross Abstract: Machine unlearning has emerged as a tool for removing personal data from trained models to comply with recent AI regulations. To evaluate unlearning e

Unsupervised Consensus-Based Anomaly Detection for Spatiotemporal Malaria Incidence in Ghana

ResearchDGX agent

arXiv:2607.21559v1 Announce Type: new Abstract: A consensus anomaly detection framework was applied to monthly malaria surveillance data from Ghana (2014-2023) to identify atypical transmission patter

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

SafetyDGX agent

arXiv:2607.21151v1 Announce Type: new Abstract: As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuiti

Verifier-First Evaluation of Agentic LLMs for Infrastructure-as-Code Generation

Model ReleasesDGX agent

arXiv:2607.20478v1 Announce Type: cross Abstract: Infrastructure-as-Code (IaC) generation from natural language requires satisfying provider schemas, dependency planning, and organizational policy con

VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification

ResearchDGX agent

arXiv:2607.20474v1 Announce Type: new Abstract: Natural language interfaces can greatly benefit the accessibility and usability of optimization modeling, and recent advances in large language models (

Vision-Language-Policy Model for Dynamic Robot Task Planning

SafetyDGX agent

arXiv:2512.19178v2 Announce Type: replace-cross Abstract: Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robotic

Visual Contrastive Self-Distillation

Model ReleasesDGX agent

arXiv:2607.21556v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymme

VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method

Model ReleasesDGX agent

arXiv:2607.21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions commonly encod

VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory

Model ReleasesDGX agent

arXiv:2603.04910v2 Announce Type: replace-cross Abstract: Imitation learning from human demonstrations has achieved significant success in robotic control, yet most visuomotor policies still condition

WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms

Model ReleasesDGX agent

arXiv:2607.20638v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in code generation and reasoning, yet their ability to perform temporal reasoning ove

WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance

Model ReleasesDGX agent

arXiv:2511.12997v2 Announce Type: replace Abstract: Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex browsing tas

What Matters for Simulation to Online Reinforcement Learning on Real Robots

ApplicationsDGX agent

arXiv:2602.20220v2 Announce Type: replace-cross Abstract: We investigate what specific design choices enable successful online reinforcement learning (RL) on physical robots. Across 100 real-world tra

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation

Model ReleasesDGX agent

arXiv:2607.21401v1 Announce Type: cross Abstract: A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard that watches that answer has to keep up w

When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers

Model ReleasesDGX agent

arXiv:2607.20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm? We answer with four findings from controlled

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion

Model ReleasesDGX agent

arXiv:2607.20543v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. We study thi

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning

Model ReleasesDGX agent

arXiv:2607.09328v2 Announce Type: replace-cross Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across dis

Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles

Local AiDGX agent

arXiv:2607.18269v2 Announce Type: replace Abstract: The wisdom of crowds -- the finding that aggregating judgments across individuals often outperforms the best individual -- has been extensively stud

Workflow-Localized Mechanism Learning: Attribution-Guided Repair and Knowledge Reuse for Structured Agent Skills

Model ReleasesDGX agent

arXiv:2607.20999v1 Announce Type: new Abstract: Agent Skills package reusable procedural knowledge as external artifacts for frozen language-model agents, yet existing optimizers do not jointly resolv

Workload-Aware Caching for Multi-Agent Systems

SafetyDGX agent

arXiv:2607.20495v1 Announce Type: new Abstract: Multi-agent systems decompose complex tasks into directed acyclic graphs (DAGs) of specialized agent executions, creating natural opportunities for cach

23 Jul 2026

A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace

AgentsDGX agent

arXiv:2607.19941v1 Announce Type: cross Abstract: As AI agents become integral to business workflows, establishing guiding user experience (UX) principles is crucial for ensuring user trust and succes

A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering

Model ReleasesDGX agent

arXiv:2507.07046v3 Announce Type: replace-cross Abstract: Nowadays, speech emotion recognition (SER) plays a vital role in the field of human-computer interaction (HCI) and the evolution of artificial

A Systematic Benchmark of Intensity Normalisation Methods for 3D Knee MRI Segmentation and Cross-Domain Generalisability

Model ReleasesDGX agent

arXiv:2607.20028v1 Announce Type: cross Abstract: Robust out-of-the-box performance is essential for the clinical deployment of deep learning models in medical imaging. An important but underexplored

Active Inference as a Convex Markov Decision Process

Local AiDGX agent

arXiv:2607.20152v1 Announce Type: cross Abstract: Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic objectives within

AdaHome: An Adaptive Smart Home Assistant using Local Small Language Models

Local AiDGX agent

arXiv:2607.18034v2 Announce Type: replace Abstract: Smart home assistants interpret a wide range of user commands, from explicit device control to underspecified and preference dependent requests. Whi

Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts

AgentsDGX agent

arXiv:2607.19629v1 Announce Type: cross Abstract: Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vulnerable states request information that

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

ResearchDGX agent

arXiv:2607.19363v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform freq

AgentCgroup: Understanding and Controlling OS Resources of AI Agents

Model ReleasesDGX agent

arXiv:2602.09345v3 Announce Type: replace-cross Abstract: AI agents are increasingly deployed in multi-tenant cloud environments, where they execute diverse tool calls within sandboxed containers, eac

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

SafetyDGX agent

arXiv:2607.19190v2 Announce Type: replace-cross Abstract: Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a str

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

Model ReleasesDGX agent

arXiv:2607.14573v3 Announce Type: replace Abstract: Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows,

An Automated Framework for Extracting Reachable Attack Chains from Cyber Threat Intelligence Reports

ApplicationsDGX agent

arXiv:2607.19742v1 Announce Type: cross Abstract: Cyber Threat Intelligence (CTI) reports richly describe real-world attack processes, but their unstructured narratives cannot be directly used for aut

An Intelligent-Cloud Edge Multimodal Interaction System for Robots

ResearchDGX agent

arXiv:2607.14675v2 Announce Type: replace-cross Abstract: Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task p

An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies

ResearchDGX agent

arXiv:2607.19771v1 Announce Type: cross Abstract: Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individ

An LLM-powered Agentic Recommendation System for Connected TV Content Discovery

AgentsDGX agent

arXiv:2607.09988v3 Announce Type: replace-cross Abstract: Recommendation systems, from traditional multi-stage to recent unified generative architectures, face challenges in incorporating diverse cont

Anatomy of a Sound Neural Reasoner: One-Shot Amortization, First-Pass Poisoning, and Search Inertness in Clue-Rich Completion

ResearchDGX agent

arXiv:2607.19635v1 Announce Type: cross Abstract: Neural solvers are built to deduce, branch, and revise intermediate states. The Lattice Deduction Transformer (LDT) appears to do exactly that. In clu

Are Attributions of Consciousness to AI Chatbots Epistemically Innocent?

ResearchDGX agent

arXiv:2607.20001v1 Announce Type: cross Abstract: Artificial intelligence (AI) chatbots (e.g., ChatGPT) can communicate in strikingly humanlike ways. This has prompted many chatbot users to attribute

ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking

AgentsDGX agent

arXiv:2601.06487v3 Announce Type: replace-cross Abstract: Reinforcement learning has substantially improved the performance of LLM agents on tasks with verifiable outcomes, but it still struggles on o

Associative Emotional Learning in Convolutional Neural Networks

ResearchDGX agent

arXiv:2607.19327v2 Announce Type: replace Abstract: Associative emotional learning enables organisms to adaptively link pleasant or unpleasant outcomes to the presence of predictive stimuli. Whereas c

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

ResearchDGX agent

arXiv:2607.20166v1 Announce Type: cross Abstract: Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g

Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology

TutorialsDGX agent

arXiv:2607.19415v1 Announce Type: cross Abstract: High-content morphological profiling (Cell Painting) yields sensitive, high-dimensional signatures of cellular state, but translating longitudinal mor

AuditVotes: Elevating Provable Defense for GNNs with Efficient Augmentation and Conditional Smoothing

ResearchDGX agent

arXiv:2503.22998v3 Announce Type: replace-cross Abstract: Despite advancements in Graph Neural Networks (GNNs), adaptive attacks continue to challenge their robustness. Certified robustness via random

Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models

Model ReleasesDGX agent

arXiv:2607.19847v1 Announce Type: cross Abstract: Predicting missing cell values in tabular data is a fundamental problem in data cleaning. While state-of-the-art reasoning models show great promise i

AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic

AgentsDGX agent

arXiv:2607.11338v2 Announce Type: replace Abstract: Symbolic expressions can effectively characterize and predict circuit behavior, but deriving them directly from circuit schematics is challenging. T

Avoiding Obfuscation with Prover-Estimator Debate

ResearchDGX agent

arXiv:2506.13609v2 Announce Type: replace Abstract: Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks.

BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators

Model ReleasesDGX agent

arXiv:2607.19438v1 Announce Type: cross Abstract: Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed

Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX

Model ReleasesDGX agent

arXiv:2607.19353v1 Announce Type: new Abstract: Confidential computing is becoming a practical deployment requirement for AI inference workloads that process sensitive inputs or protect proprietary mo

Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes

TutorialsDGX agent

arXiv:2607.19843v1 Announce Type: cross Abstract: Large language models (LLMs) have made automated program repair (APR) increasingly practical for real-world bugs, but repairing directly from bug repo

Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models

ApplicationsDGX agent

arXiv:2607.19369v1 Announce Type: new Abstract: Hidden-state probes often recover latent labels in imperfect-information sequence models, but this alone does not establish that a model maintains a pos

← Previous
1…5859606162…354
Next →