AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,688 results
Applications

Solipsistic Superintelligence is Unlikely to be Cooperative

DGX agent

arXiv:2606.03237v1 Announce Type: new Abstract: AI's central challenge is shifting from capability to coexistence. The dominant paradigm in AI research focuses on developing powerful agents that treat

applicationsarxiv-cs-ai
3 Jun 2026
Agents
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SPADE: Sketch-guided Path Planning Augmented with Diffusion Experts

DGX agent

arXiv:2606.03512v1 Announce Type: cross Abstract: Path planning is essential for Autonomous Mobile Robots (AMRs). Conventional methods for incorporating human preferences into planning typically rely

agentsarxiv-cs-ai
3 Jun 2026
Safety

Sparse-View Lung Nodule Volumetry from Digitally Reconstructed Radiographs via AReT: Anatomy-Regularized TensoRF

DGX agent

arXiv:2606.02639v1 Announce Type: cross Abstract: We identify and resolve a previously unreported failure mode in TensoRF when applied to X-ray attenuation fields: the default density shift of -10, or

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

Spike-Aware C++ INT8 Inference for Sparse Spiking Language Models on Commodity CPUs

DGX agent

arXiv:2606.03026v1 Announce Type: cross Abstract: Spiking language models expose activation sparsity that dense Transformer runtimes do not directly exploit. This paper studies that property from a sy

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Staying Alive: Uncensored Survival Analysis with Tabular Foundation Models

DGX agent

arXiv:2606.03689v1 Announce Type: cross Abstract: Survival Analysis (SA) is a statistical framework that models the time span until some event of interest occurs. Widely used in several domains, inclu

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems

DGX agent

arXiv:2606.03467v1 Announce Type: new Abstract: LLM-based multi-agent systems exhibit remarkable collaborative capabilities in complex multi-step tasks. However, these systems are highly sensitive to

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Strongly Polynomial Time Complexity of Policy Iteration for L_infty Robust MDPs

DGX agent

arXiv:2601.23229v2 Announce Type: replace Abstract: Markov decision processes (MDPs) are a fundamental model in sequential decision making. Robust MDPs (RMDPs) extend this framework by allowing uncert

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

DGX agent

arXiv:2606.02642v1 Announce Type: cross Abstract: Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing be

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation

DGX agent

arXiv:2606.03348v1 Announce Type: cross Abstract: Recent generative models can now produce visual artifacts with realistic embedded text and layouts, creating a new misinformation threat: synthetic cr

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments

DGX agent

arXiv:2606.03892v1 Announce Type: cross Abstract: Training LLMs to orchestrate multi-step tool calls is held back by three coupled obstacles: realistic stateful execution environments are costly to bu

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering

DGX agent

arXiv:2606.02624v1 Announce Type: cross Abstract: AI for scientific discovery is entering an agentic era, where protein-engineering systems are expected to prioritize future wet-lab experiments rather

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Taiji: Pareto Optimal Policy Optimization with Semantics-IDs Trade-off for Industrial LLM-Enhanced Recommendation

DGX agent

arXiv:2606.03866v1 Announce Type: cross Abstract: Scaling recommender systems via large language models (LLMs) has become a prominent trend in the industry. However, aligning the LLM's semantic space

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation

DGX agent

arXiv:2509.09685v5 Announce Type: replace-cross Abstract: We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline. In th

model-releasesarxiv-cs-ai
3 Jun 2026
Research

Target Updates May Stabilize Linear Q-Learning: Periodic and Soft Dynamics

DGX agent

arXiv:2606.02645v1 Announce Type: cross Abstract: Periodic target updates in Q-learning and soft target updates in actor-critic methods are empirically well established stabilization mechanisms, but t

researcharxiv-cs-ai
3 Jun 2026
Research

Test-Time Optimization of Physical Query Plans with LLMs

DGX agent

arXiv:2602.10387v2 Announce Type: replace-cross Abstract: Traditional query optimization relies on cost-based optimizers that estimate execution cost (e.g., runtime, memory, and I/O) using predefined

researcharxiv-cs-ai
3 Jun 2026
Model Releases

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks

DGX agent

arXiv:2606.03606v1 Announce Type: cross Abstract: Large language models achieve strong performance on arithmetic reasoning benchmarks, and one common response to arithmetic brittleness is to delegate

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios

DGX agent

arXiv:2601.08173v2 Announce Type: replace Abstract: The rapid evolution of Multi-modal Large Language Models (MLLMs) has advanced workflow automation; however, existing research mainly targets perform

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

The DeepSpeak-Agentic Dataset

DGX agent

arXiv:2606.03686v1 Announce Type: new Abstract: We present DeepSpeak-Agentic, a dataset of videos comprising over 37 hours of semi-structured conversations between a human and an embodied AI agent. We

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

The Epi-LLM Framework: probing LLM behavioral priors through epidemiological agent-based models

DGX agent

arXiv:2606.02867v1 Announce Type: cross Abstract: Human behaviour during epidemics affects infectious disease dynamics, but quantifying this remains deeply challenging. Here we introduce the Epi-LLM f

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

The Impact of Configuring Agentic AI Coding Tools on Build-vs-Buy Decisions: A Study Protocol

DGX agent

arXiv:2606.03907v1 Announce Type: cross Abstract: Agentic AI coding tools write code with increasing autonomy and in doing so decide when to import a library and when to implement functionality from s

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

DGX agent

arXiv:2606.03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for de

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size

DGX agent

arXiv:2606.02646v1 Announce Type: cross Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence. We derive a two-paramete

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs

DGX agent

arXiv:2606.03092v1 Announce Type: new Abstract: Inference-time scaling has emerged as a critical avenue for enhancing Large Language Models' performance, yet real-world deployment is constrained by st

safetyarxiv-cs-ai
3 Jun 2026
Research

The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models

DGX agent

arXiv:2606.03645v1 Announce Type: cross Abstract: Large Language Models exhibit paradoxical fragility in fundamental arithmetic, implying a disconnect between internal computation and discrete output.

researcharxiv-cs-ai
3 Jun 2026
Research

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs

DGX agent

arXiv:2606.03357v1 Announce Type: cross Abstract: When prompting SLMs for psychometric assessments, researchers assume the outputs reflect semantic reasoning. We evaluate this premise across 13 open-w

researcharxiv-cs-ai
3 Jun 2026
Applications

The Violation Situation Pattern: A Knowledge-Graph Pattern for Compliance Violations

DGX agent

arXiv:2606.03326v1 Announce Type: new Abstract: Compliance pipelines detect violations as transient query results and do not keep the violation itself as a persistent graph object with review state, a

applicationsarxiv-cs-ai
3 Jun 2026
Safety

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

DGX agent

arXiv:2606.03137v1 Announce Type: new Abstract: LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existi

safetyarxiv-cs-ai
3 Jun 2026
Research

Thinking Past the Answer: Evaluating Harmful Overthinking in Large Reasoning Models

DGX agent

arXiv:2606.02835v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by generating explicit intermediate reasoning traces through increased test-time compute, yet the assu

researcharxiv-cs-ai
3 Jun 2026
Model Releases

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning

DGX agent

arXiv:2606.03503v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards (RLVR) on Chain-of-Thoughts (Co

model-releasesarxiv-cs-ai
3 Jun 2026
Research

TimeOmni-VL: Unified Models for Time Series Understanding and Generation

DGX agent

arXiv:2602.17149v2 Announce Type: replace-cross Abstract: Recent time series modeling faces a sharp divide between numerical generation and semantic understanding, with research showing that generatio

researcharxiv-cs-ai
3 Jun 2026
Research

Tonal parsimony in chord-sequence analysis: combining modulation cost and tonal vocabulary

DGX agent

arXiv:2606.03459v1 Announce Type: cross Abstract: We study the assignment of local tonalities to chord sequences, a task useful for harmonic analysis, composition, and jazz-oriented improvisation. Sta

researcharxiv-cs-ai
3 Jun 2026
Safety

Too Much of a Good Thing: When sim2real Efforts Impede Policy Learning (And What to Do About It)

DGX agent

arXiv:2606.02636v1 Announce Type: cross Abstract: While sim2real efforts are necessary for effective policy transfer to hardware, there is such a thing as too much of a good thing. We argue that sim2r

safetyarxiv-cs-ai
3 Jun 2026
Safety

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning

DGX agent

arXiv:2606.03762v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) equips large language models (LLMs) with tool-use capabilities that substantially improve reasoning on complex tas

safetyarxiv-cs-ai
3 Jun 2026
Agents

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents

DGX agent

arXiv:2606.03054v1 Announce Type: new Abstract: Tool-augmented vision-language agents can acquire external perceptual evidence through OCR, detection, segmentation, and other tools, but executing ever

agentsarxiv-cs-ai
3 Jun 2026
Local Ai

Toward a Modular Architecture for Embedded AI Agent Systems at the Edge

DGX agent

arXiv:2606.02862v1 Announce Type: new Abstract: The rise of Large Language Models (LLMs) has enabled agentic AI capable of complex reasoning and tool use; however, deploying such autonomy in pervasive

local-aiarxiv-cs-ai
3 Jun 2026
Safety

Towards a Science of AI Agent Reliability

DGX agent

arXiv:2602.16666v3 Announce Type: replace Abstract: AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many age

safetyarxiv-cs-ai
3 Jun 2026
Hardware

Towards Compact Autonomous Driving Perception with Balanced Learning and Multi-sensor Fusion

DGX agent

arXiv:2606.02979v1 Announce Type: cross Abstract: We present a novel compact deep multi-task learning model to handle various autonomous driving perception tasks in one forward pass. The model perform

hardwarearxiv-cs-ai
3 Jun 2026
Research

Towards Non-Monotonic Entailment in Propositional Defeasible Standpoint Logic

DGX agent

arXiv:2606.03655v1 Announce Type: new Abstract: Recent work in defeasible reasoning has seen notions of preferential semantics and entailment in the style of Kraus et al. applied to modal logics. Howe

researcharxiv-cs-ai
3 Jun 2026
Research

Tracking Urban Atmospheric Pollutants using Sentinel-5P Satellite Data

DGX agent

arXiv:2606.02592v1 Announce Type: cross Abstract: Urban nitrogen dioxide (NO_2) is a key indicator of combustion-related air pollution and exhibits strong spatial and temporal variability in cities. T

researcharxiv-cs-ai
3 Jun 2026
Model Releases

Trading Human Curation for Synthetic Augmentation in RLVR

DGX agent

arXiv:2606.03800v1 Announce Type: cross Abstract: The supply of high-quality training tasks is a central bottleneck for reinforcement learning from verifiable rewards (RLVR) on agentic language models

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection

DGX agent

arXiv:2606.02812v1 Announce Type: new Abstract: Modeling patient trajectories from longitudinal electronic health records (EHRs) requires reasoning over sparse, noisy, and long-context multimodal sequ

agentsarxiv-cs-ai
3 Jun 2026
Applications

TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches

DGX agent

arXiv:2603.23117v2 Announce Type: cross Abstract: By integrating Chain-of-Thought (CoT) reasoning, Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, pa

applicationsarxiv-cs-ai
3 Jun 2026
Model Releases

TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment

DGX agent

arXiv:2606.03036v1 Announce Type: new Abstract: LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services. The domain-w

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning

DGX agent

arXiv:2606.03629v1 Announce Type: new Abstract: Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions. Recently,

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics

DGX agent

arXiv:2606.03626v1 Announce Type: cross Abstract: Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks. However, most prior work focu

model-releasesarxiv-cs-ai
3 Jun 2026
Research

Typhoon: Towards an Effective Task-Specific Masking Strategy for Pre-trained Language Models

DGX agent

arXiv:2303.15619v2 Announce Type: replace-cross Abstract: The choice of which tokens to mask is a central, under-examined design decision in masked language modeling (MLM). Standard pretraining masks

researcharxiv-cs-ai
3 Jun 2026
Applications

Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models

DGX agent

arXiv:2606.03748v1 Announce Type: cross Abstract: Real-time vision demands models that are accurate, efficient, and simple to deploy across diverse hardware. The YOLO family has become widely deployed

applicationsarxiv-cs-ai
3 Jun 2026
Agents

Uncertainty-Aware Clarification in LLM Agents with Information Gain

DGX agent

arXiv:2606.03135v1 Announce Type: new Abstract: Large Language Model (LLM) agents often operate under underspecified user instructions, where latent uncertainty over user intent leads to erroneous too

agentsarxiv-cs-ai
3 Jun 2026
← Previous
1…214215216217218…452
Next →