AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
Safety

Beyond Policy Optimization: A Data Curation Flywheel for Sparse-Reward Long-Horizon Planning

DGX agent

arXiv:2508.03018v2 Announce Type: replace Abstract: Large Language Reasoning Models have demonstrated remarkable success on static tasks, yet their application to multi-round agentic planning in inter

safetyarxiv-cs-ai
19 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs

DGX agent

arXiv:2601.16527v2 Announce Type: replace-cross Abstract: Multimodal LLMs are powerful but prone to object hallucinations, which describe non-existent entities and harm reliability. While recent unlea

model-releasesarxiv-cs-ai
19 May 2026
Local Ai

Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks

DGX agent

arXiv:2605.18194v1 Announce Type: new Abstract: While Multi-Modal Large Language Models (MLLMs) demonstrate impressive capabilities in general reasoning, their embodied spatial intelligence remains ha

local-aiarxiv-cs-ai
19 May 2026
Model Releases

BioProAgent: Neuro-Symbolic Grounding for Constrained Scientific Planning

DGX agent

arXiv:2603.00876v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant reasoning capabilities in scientific discovery but struggle to bridge the gap to physical

model-releasesarxiv-cs-ai
19 May 2026
Local Ai

BLAgent: Agentic RAG for File-Level Bug Localization

DGX agent

arXiv:2605.17965v1 Announce Type: cross Abstract: Bug localization remains a key bottleneck in downstream software maintenance tasks, including root cause analysis, triage, and automated program repai

local-aiarxiv-cs-ai
19 May 2026
Model Releases

BlendedNet++: A dataset and benchmark for field-resolved aerodynamics and inverse design of blended wing body aircraft

DGX agent

arXiv:2512.03280v2 Announce Type: replace-cross Abstract: The conceptual design of Blended Wing Body (BWB) aircraft is often constrained by the high computational cost of resolving complex aerodynamic

model-releasesarxiv-cs-ai
19 May 2026
Safety

Body-Grounded Perspective Formation and Conative Attunement in Artificial Agents

DGX agent

arXiv:2605.16728v1 Announce Type: new Abstract: This paper proposes a minimal architecture for body-grounded perspective formation in artificial agents. Extending prior work, the model introduces an i

safetyarxiv-cs-ai
19 May 2026
Model Releases

BoLT: A Benchmark to Democratize Black-box Optimization Research for Expensive LLM Tasks

DGX agent

arXiv:2605.17000v1 Announce Type: cross Abstract: Optimization of LLM training and inference configurations, such as hyperparameters, data mixtures, and prompts, is critical to performance, but it is

model-releasesarxiv-cs-ai
19 May 2026
Research

Brain Vascular Age Prediction Using Cerebral Blood Flow Velocity and Machine Learning Algorithms

DGX agent

arXiv:2605.16969v1 Announce Type: new Abstract: Defining vascular age in terms of physiological function has become one focal point of the extensive studies to categorize and track chronological age.

researcharxiv-cs-ai
19 May 2026
Research

Breaking the accuracy-resource dilemma: a lightweight adaptive video inference enhancement

DGX agent

arXiv:2601.14568v2 Announce Type: replace-cross Abstract: Existing video inference (VI) enhancement methods typically aim to improve performance by scaling up model sizes and employing sophisticated n

researcharxiv-cs-ai
19 May 2026
Research

Bridging the Version Gap: Multi-version Training Improves ICD Code Prediction, Especially for Rare Codes

DGX agent

arXiv:2605.17755v1 Announce Type: cross Abstract: Clinical coding maps clinical documentation to standardized medical codes, an essential yet time-consuming administrative task that could benefit from

researcharxiv-cs-ai
19 May 2026
Safety

Building Reliable Arithmetic Multipliers Under NBTI Aging and Process Variations

DGX agent

arXiv:2605.18444v1 Announce Type: cross Abstract: Hardware aging poses a significant challenge for integrated circuits (ICs), leading to performance degradation and eventual failure. In this work, we

safetyarxiv-cs-ai
19 May 2026
Research

Byzantine-Resilient Federated Learning via QUBO-Based Client Selection on Quantum Annealers

DGX agent

arXiv:2605.16438v1 Announce Type: cross Abstract: Federated Learning (FL) trains a global model across decentralized clients while preserving data privacy, but at scale it is vulnerable to malicious u

researcharxiv-cs-ai
19 May 2026
Agents

Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents

DGX agent

arXiv:2602.16699v3 Announce Type: replace-cross Abstract: LLM agents are deployed in environments where they must interact to acquire information. In these scenarios, the agent must reason about inher

agentsarxiv-cs-ai
19 May 2026
Model Releases

CAM-Bench: A Benchmark for Computational and Applied Mathematics in Lean

DGX agent

arXiv:2605.17255v1 Announce Type: new Abstract: Formal theorem-proving benchmarks enable mechanically verifiable evaluation of mathematical reasoning in large language models. However, existing benchm

model-releasesarxiv-cs-ai
19 May 2026
Research

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection

DGX agent

arXiv:2605.17133v1 Announce Type: cross Abstract: The rapid advancement of Deepfake technologies and video manipulation tools poses a critical challenge to multimedia forensics, judicial evidence inte

researcharxiv-cs-ai
19 May 2026
Model Releases

Can Heterogeneous Language Models Be Fused?

DGX agent

arXiv:2604.01674v2 Announce Type: replace Abstract: Model merging aims to integrate multiple expert models into a single model that inherits their complementary strengths without incurring the inferen

model-releasesarxiv-cs-ai
19 May 2026
Agents

Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment

DGX agent

arXiv:2603.23638v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly tested on complex tasks, but their ability to allocate scarce resources over long horizons remain

agentsarxiv-cs-ai
19 May 2026
Research

Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks

DGX agent

arXiv:2510.01782v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) should refuse to answer questions beyond their knowledge. This capability, which we term knowledge-aware refusal,

researcharxiv-cs-ai
19 May 2026
Model Releases

Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench

DGX agent

arXiv:2605.17079v1 Announce Type: cross Abstract: LLMs are increasingly used as ``digital consumers'' to simulate public opinion, pre-test marketing decisions, and anticipate audience response. Howeve

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

CANSURF: An ASV-View Can Dataset and Benchmark for Detection and Tracking of Surface-Level Debris

DGX agent

arXiv:2605.16774v1 Announce Type: cross Abstract: Surface-level marine debris remains a practical bottleneck for autonomous clean-up, where small, reflective targets (e.g., aluminum cans) must be dete

model-releasesarxiv-cs-ai
19 May 2026
Research

Capturing LLM Capabilities via Evidence-Calibrated Query Clustering

DGX agent

arXiv:2605.17110v1 Announce Type: new Abstract: Query clustering organizes queries into groups that reflect shared latent capability demands, enabling capability-aware LLM evaluation. Existing cluster

researcharxiv-cs-ai
19 May 2026
Model Releases

CarbonScaling: Extending Neural Scaling Laws for Carbon Footprint in Large Language Models

DGX agent

arXiv:2508.06524v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly follow neural scaling laws that tie performance gains to rapidly expanding computational budgets, ra

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning

DGX agent

arXiv:2605.17176v1 Announce Type: new Abstract: Emotion understanding is a core capability for LLMs to interact effectively with humans, yet existing evaluation paradigms rely on discrete emotion labe

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

CasualSynth: Generating Structurally Sound Synthetic Data

DGX agent

arXiv:2605.17528v1 Announce Type: cross Abstract: Large Language Models (LLMs) generate realistic synthetic data but offer no guarantee that their outputs respect the causal mechanisms governing the t

model-releasesarxiv-cs-ai
19 May 2026
Research

CATA: Continual Machine Unlearning via Conflict-Averse Task Arithmetic

DGX agent

arXiv:2605.18610v1 Announce Type: cross Abstract: Vision-language models (VLMs) have shown remarkable ability in aligning visual and textual representations, enabling a wide range of multimodal applic

researcharxiv-cs-ai
19 May 2026
Safety

CatalyticMLLM: A Graph-Text Multimodal Large Language Model for Catalytic Materials

DGX agent

arXiv:2605.17254v1 Announce Type: new Abstract: Property prediction and inverse structural design of catalytic materials are typically modeled as two independent tasks: the former predicts target prop

safetyarxiv-cs-ai
19 May 2026
Research

Catastrophic Overfitting, Entropy Gap and Participation Ratio: A Noiseless l^p Norm Solution for Fast Adversarial Training

DGX agent

arXiv:2505.02360v2 Announce Type: replace-cross Abstract: Adversarial training is a cornerstone of robust deep learning, but fast methods like the Fast Gradient Sign Method (FGSM) often suffer from Ca

researcharxiv-cs-ai
19 May 2026
Model Releases

Causal Intervention-Based Memory Selection for Long-Horizon LLM Agents

DGX agent

arXiv:2605.17641v1 Announce Type: new Abstract: Long-horizon LLM agents rely on persistent memory to support interactions across sessions, yet existing memory systems often retrieve context using sema

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Causely: A Causal Intelligence Layer for Enterprise AI A Benchmark Study on SRE and Reliability Workflows

DGX agent

arXiv:2605.18327v1 Announce Type: new Abstract: AI agents deployed into SRE workflows currently derive their understanding of environment state from raw observability telemetry at query time, paying a

model-releasesarxiv-cs-ai
19 May 2026
Tutorials

CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning

DGX agent

arXiv:2605.16416v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have achieved strong performance on general multimodal reasoning, yet remain challenged in integrating nonlocal visual i

tutorialsarxiv-cs-ai
19 May 2026
Tutorials

CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings

DGX agent

arXiv:2605.17370v1 Announce Type: new Abstract: Cognitive behavioural therapy is widely used to help patients understand and manage psychological distress. It is often delivered through spoken convers

tutorialsarxiv-cs-ai
19 May 2026
Hardware

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference

DGX agent

arXiv:2605.17164v1 Announce Type: cross Abstract: Deploying large-scale LLM training and inference with optimal performance is exceptionally challenging due to a complex design space of parallelism st

hardwarearxiv-cs-ai
19 May 2026
Safety

ChartDesign: Towards LLM Designer of Data Visualization

DGX agent

arXiv:2605.16274v1 Announce Type: cross Abstract: Charts are the dominant medium for visualizing data, discovering patterns and trends, and communicating data driven insights, yet designing them still

safetyarxiv-cs-ai
19 May 2026
Local Ai

CheckSupport: A Local LLM-Powered Tool for Automated Manuscript Submission Checklist Selection and Completion

DGX agent

arXiv:2605.16377v1 Announce Type: cross Abstract: Transparent and standardized reporting is essential for reproducible scientific research, yet adherence to reporting guidelines remains inconsistent b

local-aiarxiv-cs-ai
19 May 2026
Safety

ChemVA: Advancing Large Language Models on Chemical Reaction Diagrams Understanding

DGX agent

arXiv:2605.17214v1 Announce Type: new Abstract: While Large Language Models (LLMs) have revolutionized scientific text processing, they exhibit a significant capability gap when interpreting chemical

safetyarxiv-cs-ai
19 May 2026
Model Releases

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

DGX agent

arXiv:2605.16679v1 Announce Type: cross Abstract: End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving

DGX agent

arXiv:2605.17284v1 Announce Type: cross Abstract: End-to-end autonomous driving systems powered by Vision-Language-Action (VLA) models achieve strong performance on common driving scenarios, yet remai

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

ClawArena: Benchmarking AI Agents in Evolving Information Environments

DGX agent

arXiv:2604.04202v2 Announce Type: replace-cross Abstract: AI agents deployed as persistent assistants must maintain correct beliefs as their information environment evolves. In practice, evidence is s

model-releasesarxiv-cs-ai
19 May 2026
Safety

Code as Agent Harness

DGX agent

arXiv:2605.18747v1 Announce Type: cross Abstract: Recent large language models (LLMs) have demonstrated strong capabilities in understanding and generating code, from competitive programming to reposi

safetyarxiv-cs-ai
19 May 2026
Safety

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook

DGX agent

arXiv:2605.18257v1 Announce Type: cross Abstract: Multimodal representation alignment is pivotal for large language models and robotics. Traditional methods are often hindered by cross-modal informati

safetyarxiv-cs-ai
19 May 2026
Research

CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models

DGX agent

arXiv:2602.17684v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based f

researcharxiv-cs-ai
19 May 2026
Tutorials

CoLLM-NAS: Collaborative Large Language Models for Efficient Knowledge-Guided Neural Architecture Search

DGX agent

arXiv:2509.26037v2 Announce Type: replace Abstract: The integration of Large Language Models (LLMs) with Neural Architecture Search (NAS) has introduced new possibilities for automating the design of

tutorialsarxiv-cs-ai
19 May 2026
Safety

COLSON: Controllable Learning-Based Social Navigation via Diffusion-Based Reinforcement Learning

DGX agent

arXiv:2503.13934v2 Announce Type: replace-cross Abstract: Mobile robot navigation in dynamic environments with pedestrian traffic is a key challenge in the development of autonomous mobile service rob

safetyarxiv-cs-ai
19 May 2026
Model Releases

CommitDistill: A Lightweight Knowledge-Centric Memory Layer for Software Repositories

DGX agent

arXiv:2605.18284v1 Announce Type: cross Abstract: Software repositories accumulate large amounts of unstructured knowledge in commit messages, pull-request discussions, and issue threads, but develope

model-releasesarxiv-cs-ai
19 May 2026
Research

Computational Challenges in Token Economics: Bridging Economic Theory and AI System Design

DGX agent

arXiv:2605.17410v1 Announce Type: new Abstract: Token economics has emerged as a useful lens for understanding resource allocation, value creation, and pricing in large language model systems. While r

researcharxiv-cs-ai
19 May 2026
Research

Concise and Logically Consistent Conformal Sets for Neuro-Symbolic Concept-Based Models

DGX agent

arXiv:2605.18202v1 Announce Type: cross Abstract: Neuro-Symbolic Concept-based Models (NeSy-CBMs) are a family of architectures that integrate neural networks with symbolic reasoning for enhanced reli

researcharxiv-cs-ai
19 May 2026
Safety

Confidence-Gated Robot Autonomy: When Does Uncertainty Actually Help?

DGX agent

arXiv:2605.18045v1 Announce Type: cross Abstract: Robotic systems often use predictive uncertainty to decide whether to act autonomously or defer to a fallback policy. In threshold-gated autonomy, unc

safetyarxiv-cs-ai
19 May 2026
← Previous
1…292293294295296…448
Next →