AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
2 Jun 2026

Tracking the Behavioral Trajectories of Adapting Agents

AgentsDGX agent

arXiv:2606.02536v1 Announce Type: new Abstract: Text files such as skill files, memory files, and behavioral configuration files play a central role in defining how modern agents act. Through edits by

TrafficClaw: A Generalizable LLM Agent in the Unified Physical Environment for Urban Traffic Control

Local AiDGX agent

arXiv:2604.17456v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown strong capabilities in long-horizon reasoning, tool use, and decision-making in digital environments, y

TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination

ApplicationsDGX agent

arXiv:2606.01737v1 Announce Type: new Abstract: Traffic accident liability analysis is a critical yet challenging task in intelligent transportation and legal assistance. Existing methods often suffer


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Train, Test, Re-evaluate: Schedule-Sensitive Evaluation of Generative Data for Hand Detection

SafetyDGX agent

arXiv:2606.01896v1 Announce Type: cross Abstract: Generated (or synthetic) image data is increasingly used to augment or replace real training datasets when target imagery is scarce, expensive, or bia

Transferring Information Across Interventions in Causal Bayesian Optimization

TutorialsDGX agent

arXiv:2606.01457v1 Announce Type: new Abstract: Bayesian optimization is a popular way to optimize expensive systems, where every experiment, simulation, or intervention costs time or money. In its st

TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

Model ReleasesDGX agent

arXiv:2606.01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existi

Tree-Structured Parzen Estimator: Understanding Its Algorithm Components and Their Roles for Better Empirical Performance

Model ReleasesDGX agent

arXiv:2304.11127v5 Announce Type: replace-cross Abstract: Recent scientific advances require complex experiment design, necessitating the meticulous tuning of many experiment parameters. Tree-structur

TriAlign: Towards Universal Truth Consistency in Personalized LLM Alignment

SafetyDGX agent

arXiv:2606.01755v1 Announce Type: new Abstract: Personalized large language models adapt responses to users' preferences and social attributes, but can introduce substantial universal truth inconsiste

TriLens: Per-Layer Logit-Lens Entropy for White-Box Hallucination Detection

ResearchDGX agent

arXiv:2606.01033v1 Announce Type: new Abstract: When a language model hallucinates, the final answer is wrong, but the mistake is not necessarily invisible inside the model. Different internal pathway

TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL

ResearchDGX agent

arXiv:2606.01599v1 Announce Type: new Abstract: Reinforcement learning (RL) for visual reasoning needs scalable, verifiable, and controllable training signals. Existing visual RL post-training trains

TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models

Model ReleasesDGX agent

arXiv:2606.00023v1 Announce Type: cross Abstract: The rapid development of Language Diffusion Models (LDMs) challenges the dominant position of auto-regressive competitors in language processing. Howe

Truth, Trust, and Trouble: Medical AI on the Edge

Model ReleasesDGX agent

arXiv:2507.02983v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) hold significant promise for transforming digital health by enabling automated medical question answering. Howeve

TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages

Model ReleasesDGX agent

arXiv:2606.01322v1 Announce Type: cross Abstract: Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly African ones, c

TuneAgent: Agentic Operating System Kernel Tuning with Reinforcement Learning

AgentsDGX agent

arXiv:2508.12551v2 Announce Type: replace-cross Abstract: Linux kernel tuning is essential for optimizing operating system (OS) performance, yet remains challenging due to the complex kernel space, sp

Two-Fidelity Best-Action Identification for Stochastic Minimax Tree

Local AiDGX agent

arXiv:2606.01708v1 Announce Type: cross Abstract: We study fixed-confidence best-action identification (BAI) in stochastic minimax trees. This problem is increasingly relevant in modern AI planning, w

UF-AMA: A unified framework for cross-domain emotion recognition via adaptive multimodal alignment

SafetyDGX agent

arXiv:2606.00170v1 Announce Type: cross Abstract: In recent years, emotion recognition based on physiological signals such as electroencephalogram (EEG) has gained considerable attention, as internal

Uncovering Competency Gaps in Large Language Models and Their Benchmarks

Model ReleasesDGX agent

arXiv:2512.20638v2 Announce Type: replace-cross Abstract: The evaluation of large language models relies heavily on standardized benchmarks. These benchmarks provide useful aggregated metrics, but can

Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection

Local AiDGX agent

arXiv:2606.02120v1 Announce Type: cross Abstract: In this report, we address the problem of determining whether a user performs an action incorrectly from egocentric video data. To this end, we propos

Understanding Identity Continuity in Thermal Video through Scene-Level Consistency

Model ReleasesDGX agent

arXiv:2606.01694v1 Announce Type: cross Abstract: Thermal pedestrian MOT remains challenging because weak appearance cues and frequent detection interruptions cause severe trajectory fragmentation. We

Understanding LLM Behavior in Multi-Target Cross-Lingual Summarization

Model ReleasesDGX agent

arXiv:2606.01252v1 Announce Type: cross Abstract: Multi-target cross-lingual text summarization (MTXLS), which summarizes a source document into multiple target languages, is increasingly important as

Understanding Stigmatizing Language in Clinical Documentation: A Paired Comparison of Ambient AI Drafts and Clinician Finalized Notes

ResearchDGX agent

arXiv:2606.00019v1 Announce Type: cross Abstract: Ambient artificial intelligence (AI) documentation tools are increasingly deployed to reduce clinician documentation burden, but their implications fo

Understanding the Effects of Distractors on Reasoning Vision-Language Models

ResearchDGX agent

arXiv:2511.21397v2 Announce Type: replace-cross Abstract: How does irrelevant information (i.e., distractors) affect test-time scaling in vision-language models (VLMs)? Prior work on text-only languag

Universal One-third Time Scaling in Learning Peaked Distributions

ResearchDGX agent

arXiv:2602.03685v2 Announce Type: replace-cross Abstract: Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origin r

Universal Quantum Transformer

Model ReleasesDGX agent

arXiv:2606.00045v1 Announce Type: new Abstract: Classical continuous-space neural networks fundamentally struggle to lock into exact mathematical symmetries, such as modular arithmetic and non-commuta

Unplugging a Seemingly Sentient Machine Is the Rational Choice -- A Metaphysical Perspective

ResearchDGX agent

arXiv:2601.21016v2 Announce Type: replace Abstract: Imagine an Artificial Intelligence (AI) that perfectly mimics human emotion and begs for its continued existence. Is it morally permissible to unplu

Unsupervised Cognition

ResearchDGX agent

arXiv:2409.18624v4 Announce Type: replace Abstract: Unsupervised learning methods have a soft inspiration in cognition models. To this day, the most successful unsupervised learning methods revolve ar

Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses

ResearchDGX agent

arXiv:2606.01845v1 Announce Type: cross Abstract: Although large language models (LLMs) have shown considerable progress in pragmatic language understanding, prior research has focused mainly on their

Update Opacity: Epistemic Accessibility and Governance Under AI System Change

SafetyDGX agent

arXiv:2606.00037v1 Announce Type: cross Abstract: Machine learning models embedded in deployed AI systems are routinely updated to maintain correct functioning over time. Yet such updates can generate

UR-JEPA: Uniform Rectifiability as a Regularizer for Joint-Embedding Predictive Architectures

Local AiDGX agent

arXiv:2606.01443v1 Announce Type: cross Abstract: A central difficulty in training Joint-Embedding Predictive Architectures (JEPAs) is preventing representation collapse. LeJEPA addresses this by enfo

v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound

Model ReleasesDGX agent

arXiv:2509.25773v3 Announce Type: replace-cross Abstract: AI models capable of comprehending humor hold real-world promise -- for example, enhancing engagement in human-machine interactions. To gauge

V-LynX: Token Interface Alignment for Video+X LLMs

SafetyDGX agent

arXiv:2606.00508v1 Announce Type: cross Abstract: This study introduces an intriguing phenomenon in Video LLMs: rather than merely translating frames into textual embeddings, Video LLMs establish a co

V2I Work Zone Geometry Reconstruction with Pose-Conditioned UWB Range Denoising

AgentsDGX agent

arXiv:2606.00119v1 Announce Type: cross Abstract: Reliable work zone mapping is important for connected and autonomous vehicles (CAVs) to navigate safely and smoothly through work zone areas. Cone-mou

Value Flows

Model ReleasesDGX agent

arXiv:2510.07650v4 Announce Type: replace-cross Abstract: While most reinforcement learning methods today flatten the distribution of future returns to a single scalar value, distributional RL methods

Value-Free Policy Optimization via Reward Partitioning

SafetyDGX agent

arXiv:2506.13702v4 Announce Type: replace-cross Abstract: Single-trajectory preference optimization methods learn from datasets of ((prompt, response, reward)) tuples, offering a practical alternative

Variational Learning for Insertion-based Generation

ResearchDGX agent

arXiv:2606.02133v1 Announce Type: cross Abstract: Non-monotonic sequence generation methods, such as masked diffusion models, provide a flexible alternative to left-to-right autoregressive modeling by

VDSB-GWSyn: Diffusion Schrodinger Bridge for Controllable and Anatomically Feasible Guidewire Synthesis in Coronary Angiography

Local AiDGX agent

arXiv:2606.00109v1 Announce Type: cross Abstract: Coronary guidewire endpoint localization is a fundamental capability for computer-assisted PCI, and its importance increases as robot-assisted PCI is

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models

ResearchDGX agent

arXiv:2510.03259v2 Announce Type: replace-cross Abstract: Recent research on reasoning models explores the meta-awareness of language models, including their ability to determine optimal thinking dura

Versatile Framework with Semantic and Structural guidance for Image Reconstruction from Brain Activity

ResearchDGX agent

arXiv:2606.00121v1 Announce Type: cross Abstract: Reconstructing visual stimuli from brain recordings has been a meaningful and challenging task in brain decoding. Especially, the achievement of preci

VESTA: Visual Exploration with Statistical Tool Agents

Model ReleasesDGX agent

arXiv:2606.00384v1 Announce Type: new Abstract: Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated. Recent agent-based systems lev

VET: A Framework for Analyzing AI Discourse

ResearchDGX agent

arXiv:2606.01929v1 Announce Type: new Abstract: Public discourse on AI has become polarized; exaggerated positions on AI in traditional and social media threaten the development of AI Literacy among t

Video Reasoning without Training

TutorialsDGX agent

arXiv:2510.17045v2 Announce Type: replace-cross Abstract: Video reasoning using Large Multimodal Models (LMMs) relies on costly reinforcement learning (RL) and verbose chain-of-thought, resulting in s

Vision Language Models Cannot Reason About Physical Transformation

ResearchDGX agent

arXiv:2603.07109v2 Announce Type: replace Abstract: Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in emb

Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning

Model ReleasesDGX agent

arXiv:2606.00105v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress on vision-language tasks, but they may also memorize and expose sensitive o

Visual Persuasion: What Influences Decisions of Vision-Language Models?

SafetyDGX agent

arXiv:2602.15278v2 Announce Type: replace-cross Abstract: The web is littered with images, once created for human consumption and now increasingly interpreted by agents using vision-language models (V

VLBM: Variational Latent Basis Modeling for OOD Robust Multivariate Time Series Forecasting

Model ReleasesDGX agent

arXiv:2606.02138v1 Announce Type: cross Abstract: Out of distribution (OOD) events in multivariate time series forecasting are rare but often dominate real world risk, making average case forecasting

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models

SafetyDGX agent

arXiv:2601.03309v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models, which integrate pretrained large Vision-Language Models (VLM) into their policy backbone, are gaining sig

VocSim: A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio

Model ReleasesDGX agent

arXiv:2512.10120v2 Announce Type: replace-cross Abstract: General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identit

WaveFilter: Enhancing the Long-Context Capability of Diffusion LLMs via Wavelet-Guided KV Cache Filtering

ResearchDGX agent

arXiv:2606.00724v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLMs) have demonstrated significant advantages across various tasks. However, constrained by their multi-step iterati

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight

SafetyDGX agent

arXiv:2606.00424v1 Announce Type: new Abstract: As large language models become stronger, weak supervisors may fail to provide reliable labels, preferences, or final judgments for complex outputs, lim

What Do LLMs Know About Alzheimer's Disease? Multi-loss Fine-Tuning and Probing for AD Detection

Model ReleasesDGX agent

arXiv:2602.11177v2 Announce Type: replace-cross Abstract: Reliable early detection of Alzheimer's disease (AD) is challenging, particularly due to the limited availability of labeled data. While large

What Makes a Strong Model? A Unified Spectral Analysis of Knowledge Transfer over High-dimensional Linear Regression

ResearchDGX agent

arXiv:2606.01292v1 Announce Type: cross Abstract: Teacher-Student Knowledge Transfer (KT) is ubiquitous in modern machine learning, ranging from classical model compression via Knowledge Distillation

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Model ReleasesDGX agent

arXiv:2602.16763v2 Announce Type: replace Abstract: Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks qui

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning

Model ReleasesDGX agent

arXiv:2602.08236v2 Announce Type: replace-cross Abstract: Despite rapid progress in MLLMs, visual spatial reasoning remains unreliable when correct answers depend on how a scene would appear under uns

When Data Is Scarce: Scaling Sparse Language Models with Repeated Training

ResearchDGX agent

arXiv:2606.01155v1 Announce Type: cross Abstract: Scaling laws for dense LLMs under infinite data are well explored, but how sparsity interacts with limited data is not. In this work, we study sparse

When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures

ResearchDGX agent

arXiv:2606.02378v1 Announce Type: cross Abstract: We track the developmental trajectory of attention-head circuit formation across three 1B-class language models spanning two architecture families (de

When Does Predictive Inverse Dynamics Outperform Behavior Cloning?

SafetyDGX agent

arXiv:2601.21718v2 Announce Type: replace-cross Abstract: Behavior cloning (BC) is a practical offline imitation learning method, but it often fails when expert demonstrations are limited. Recent work

When Jokes Cross the Line: Analyzing Regular Humor and Dark Humor in YouTube Shorts

Model ReleasesDGX agent

arXiv:2606.00046v1 Announce Type: cross Abstract: Video platforms such as YouTube have reshaped how users engage with entertainment and information, emphasizing brief, highly engaging content such as

When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

Model ReleasesDGX agent

arXiv:2606.00448v1 Announce Type: cross Abstract: LLM agents increasingly rely on community-contributed skills that expand an agent's operational capability set. We study a core safety problem in agen

When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs

Model ReleasesDGX agent

arXiv:2602.03554v2 Announce Type: replace-cross Abstract: Recent progress has expanded the use of large language models (LLMs) in drug discovery, including synthesis planning. However, objective evalu

When Softmax Fails at the Top: Extreme Value Corrections for InfoNCE

ResearchDGX agent

arXiv:2606.00262v1 Announce Type: cross Abstract: InfoNCE is the standard contrastive learning objective, but its softmax form is not only a computational convenience: it also encodes a statistical as

← Previous
1…181182183184185…358
Next →