AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
AllBlog
86,510Total entries
1Added by human
86,509Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
22,241 results
Research

The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models

DGX agent

arXiv:2606.03645v1 Announce Type: cross Abstract: Large Language Models exhibit paradoxical fragility in fundamental arithmetic, implying a disconnect between internal computation and discrete output.

researcharxiv-cs-ai
3 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs

DGX agent

arXiv:2606.03357v1 Announce Type: cross Abstract: When prompting SLMs for psychometric assessments, researchers assume the outputs reflect semantic reasoning. We evaluate this premise across 13 open-w

researcharxiv-cs-ai
3 Jun 2026
Applications

The Violation Situation Pattern: A Knowledge-Graph Pattern for Compliance Violations

DGX agent

arXiv:2606.03326v1 Announce Type: new Abstract: Compliance pipelines detect violations as transient query results and do not keep the violation itself as a persistent graph object with review state, a

applicationsarxiv-cs-ai
3 Jun 2026
Safety

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

DGX agent

arXiv:2606.03137v1 Announce Type: new Abstract: LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existi

safetyarxiv-cs-ai
3 Jun 2026
Research

Thinking Past the Answer: Evaluating Harmful Overthinking in Large Reasoning Models

DGX agent

arXiv:2606.02835v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by generating explicit intermediate reasoning traces through increased test-time compute, yet the assu

researcharxiv-cs-ai
3 Jun 2026
Model Releases

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning

DGX agent

arXiv:2606.03503v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards (RLVR) on Chain-of-Thoughts (Co

model-releasesarxiv-cs-ai
3 Jun 2026
Research

TimeOmni-VL: Unified Models for Time Series Understanding and Generation

DGX agent

arXiv:2602.17149v2 Announce Type: replace-cross Abstract: Recent time series modeling faces a sharp divide between numerical generation and semantic understanding, with research showing that generatio

researcharxiv-cs-ai
3 Jun 2026
Research

Tonal parsimony in chord-sequence analysis: combining modulation cost and tonal vocabulary

DGX agent

arXiv:2606.03459v1 Announce Type: cross Abstract: We study the assignment of local tonalities to chord sequences, a task useful for harmonic analysis, composition, and jazz-oriented improvisation. Sta

researcharxiv-cs-ai
3 Jun 2026
Safety

Too Much of a Good Thing: When sim2real Efforts Impede Policy Learning (And What to Do About It)

DGX agent

arXiv:2606.02636v1 Announce Type: cross Abstract: While sim2real efforts are necessary for effective policy transfer to hardware, there is such a thing as too much of a good thing. We argue that sim2r

safetyarxiv-cs-ai
3 Jun 2026
Safety

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning

DGX agent

arXiv:2606.03762v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) equips large language models (LLMs) with tool-use capabilities that substantially improve reasoning on complex tas

safetyarxiv-cs-ai
3 Jun 2026
Agents

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents

DGX agent

arXiv:2606.03054v1 Announce Type: new Abstract: Tool-augmented vision-language agents can acquire external perceptual evidence through OCR, detection, segmentation, and other tools, but executing ever

agentsarxiv-cs-ai
3 Jun 2026
Local Ai

Toward a Modular Architecture for Embedded AI Agent Systems at the Edge

DGX agent

arXiv:2606.02862v1 Announce Type: new Abstract: The rise of Large Language Models (LLMs) has enabled agentic AI capable of complex reasoning and tool use; however, deploying such autonomy in pervasive

local-aiarxiv-cs-ai
3 Jun 2026
Safety

Towards a Science of AI Agent Reliability

DGX agent

arXiv:2602.16666v3 Announce Type: replace Abstract: AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many age

safetyarxiv-cs-ai
3 Jun 2026
Hardware

Towards Compact Autonomous Driving Perception with Balanced Learning and Multi-sensor Fusion

DGX agent

arXiv:2606.02979v1 Announce Type: cross Abstract: We present a novel compact deep multi-task learning model to handle various autonomous driving perception tasks in one forward pass. The model perform

hardwarearxiv-cs-ai
3 Jun 2026
Research

Towards Non-Monotonic Entailment in Propositional Defeasible Standpoint Logic

DGX agent

arXiv:2606.03655v1 Announce Type: new Abstract: Recent work in defeasible reasoning has seen notions of preferential semantics and entailment in the style of Kraus et al. applied to modal logics. Howe

researcharxiv-cs-ai
3 Jun 2026
Research

Tracking Urban Atmospheric Pollutants using Sentinel-5P Satellite Data

DGX agent

arXiv:2606.02592v1 Announce Type: cross Abstract: Urban nitrogen dioxide (NO_2) is a key indicator of combustion-related air pollution and exhibits strong spatial and temporal variability in cities. T

researcharxiv-cs-ai
3 Jun 2026
Model Releases

Trading Human Curation for Synthetic Augmentation in RLVR

DGX agent

arXiv:2606.03800v1 Announce Type: cross Abstract: The supply of high-quality training tasks is a central bottleneck for reinforcement learning from verifiable rewards (RLVR) on agentic language models

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection

DGX agent

arXiv:2606.02812v1 Announce Type: new Abstract: Modeling patient trajectories from longitudinal electronic health records (EHRs) requires reasoning over sparse, noisy, and long-context multimodal sequ

agentsarxiv-cs-ai
3 Jun 2026
Applications

TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches

DGX agent

arXiv:2603.23117v2 Announce Type: cross Abstract: By integrating Chain-of-Thought (CoT) reasoning, Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, pa

applicationsarxiv-cs-ai
3 Jun 2026
Model Releases

TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment

DGX agent

arXiv:2606.03036v1 Announce Type: new Abstract: LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services. The domain-w

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning

DGX agent

arXiv:2606.03629v1 Announce Type: new Abstract: Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions. Recently,

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics

DGX agent

arXiv:2606.03626v1 Announce Type: cross Abstract: Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks. However, most prior work focu

model-releasesarxiv-cs-ai
3 Jun 2026
Research

Typhoon: Towards an Effective Task-Specific Masking Strategy for Pre-trained Language Models

DGX agent

arXiv:2303.15619v2 Announce Type: replace-cross Abstract: The choice of which tokens to mask is a central, under-examined design decision in masked language modeling (MLM). Standard pretraining masks

researcharxiv-cs-ai
3 Jun 2026
Applications

Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models

DGX agent

arXiv:2606.03748v1 Announce Type: cross Abstract: Real-time vision demands models that are accurate, efficient, and simple to deploy across diverse hardware. The YOLO family has become widely deployed

applicationsarxiv-cs-ai
3 Jun 2026
Agents

Uncertainty-Aware Clarification in LLM Agents with Information Gain

DGX agent

arXiv:2606.03135v1 Announce Type: new Abstract: Large Language Model (LLM) agents often operate under underspecified user instructions, where latent uncertainty over user intent leads to erroneous too

agentsarxiv-cs-ai
3 Jun 2026
Research

Unveiling the Structure of Do-Calculus Reasoning via Derivation Graphs

DGX agent

arXiv:2606.03719v1 Announce Type: new Abstract: The do-calculus defines a general system of inference for interventional queries, allowing causal quantities to be transformed through successive applic

researcharxiv-cs-ai
3 Jun 2026
Safety

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning

DGX agent

arXiv:2606.03962v1 Announce Type: cross Abstract: Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward. Yet, modern applicati

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

VidMsg: A Benchmark for Implicit Message Inference in Short Videos

DGX agent

arXiv:2606.03635v1 Announce Type: cross Abstract: Understanding short online videos involves more than identifying visible objects and actions; video makers often include an underlying message or purp

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

VistaHop: Benchmarking Multi-hop Visual Reasoning for Visual DeepSearch

DGX agent

arXiv:2606.03273v1 Announce Type: cross Abstract: Visual DeepSearch requires multimodal large reasoning model (MLRM) agents to answer complex visual queries by repeatedly inspecting image regions, gro

model-releasesarxiv-cs-ai
3 Jun 2026
Tutorials

Visual Graph Scaffolds for Structural Reasoning in Large Language Models

DGX agent

arXiv:2606.02673v1 Announce Type: new Abstract: Graphs have been used to enhance large language models (LLMs) for structured reasoning, mostly as external knowledge sources are provided to models at t

tutorialsarxiv-cs-ai
3 Jun 2026
Model Releases

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

DGX agent

arXiv:2603.04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- se

model-releasesarxiv-cs-ai
3 Jun 2026
Local Ai

VulnAgent-R2: Evidence-Calibrated Multi-Agent Auditing for Repository-Level Vulnerability Detection

DGX agent

arXiv:2603.13384v2 Announce Type: replace-cross Abstract: Software vulnerabilities often depend on cross-file data flow, build options, framework conventions, and runtime guards, so isolated function

local-aiarxiv-cs-ai
3 Jun 2026
Research

Wavelet as Tokenizer: Preliminary Results on a Shared Wavelet Token Schema for Natural Signals

DGX agent

arXiv:2606.02631v1 Announce Type: cross Abstract: This paper studies whether audio, images, and video can share a common wavelet token schema rather than relying on separate modality-specific latent g

researcharxiv-cs-ai
3 Jun 2026
Model Releases

Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning

DGX agent

arXiv:2509.19305v2 Announce Type: replace-cross Abstract: Diffusion probability models have shown significant promise in offline reinforcement learning by directly modeling trajectory sequences. Howev

model-releasesarxiv-cs-ai
3 Jun 2026
Local Ai

WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts

DGX agent

arXiv:2606.03220v1 Announce Type: cross Abstract: Existing benchmarks for MLLM-generated web artifacts assess interaction through local evidence and miss the requirement-induced states and transitions

local-aiarxiv-cs-ai
3 Jun 2026
Model Releases

What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents

DGX agent

arXiv:2606.02965v1 Announce Type: new Abstract: Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceed

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

What Makes Interaction Trajectories Effective for Training Terminal Agents?

DGX agent

arXiv:2606.03461v1 Announce Type: new Abstract: Stronger code agents are commonly assumed to be superior teachers for post-training, yet this assumption remains poorly disentangled from task difficult

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics

DGX agent

arXiv:2606.03569v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated remarkable capabilities but suffer from significant computational overhead during inference. While vis

safetyarxiv-cs-ai
3 Jun 2026
Agents

When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning

DGX agent

arXiv:2606.02866v1 Announce Type: new Abstract: When does multi-agent debate help data cleaning, and when does it hurt? Across three benchmarks, four model families, and over 6,000 task-condition pair

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

When Model Merging Breaks Routing: Training-Free Calibration for MoE

DGX agent

arXiv:2606.03391v1 Announce Type: cross Abstract: Model merging has emerged as a cost-effective approach for consolidating the capabilities of multiple LLMs without retraining. However, existing mergi

model-releasesarxiv-cs-ai
3 Jun 2026
Local Ai

When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming

DGX agent

arXiv:2606.03238v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) makes large-scale post-training possible by replacing an underspecified human objective with learned

local-aiarxiv-cs-ai
3 Jun 2026
Research

When Should LLMs Be Less Specific? Selective Abstraction for Reliable Long-Form Text Generation

DGX agent

arXiv:2602.11908v3 Announce Type: replace Abstract: LLMs are widely used, yet they remain prone to factual errors that erode user trust and limit adoption in high-risk settings. One approach to mitiga

researcharxiv-cs-ai
3 Jun 2026
Model Releases

When Should the Teacher Move? Temporal Coupling and Stability in Self On-Policy Distillation

DGX agent

arXiv:2606.03532v1 Announce Type: cross Abstract: Self on-policy distillation trains a student policy against a teacher derived from its own parameter history, yet the teacher's update schedule -- whi

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning

DGX agent

arXiv:2606.03741v1 Announce Type: new Abstract: Long-horizon reasoning requires a system to commit to medium-horizon intent without becoming rigid: re-plan too often and computation never coheres into

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

DGX agent

arXiv:2606.02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-regis

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System

DGX agent

arXiv:2602.08335v2 Announce Type: replace Abstract: Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving com

safetyarxiv-cs-ai
3 Jun 2026
Agents

Whom to Query for What: Adaptive Group Elicitation via Multi-Turn LLM Interactions

DGX agent

arXiv:2602.14279v2 Announce Type: replace-cross Abstract: Eliciting information to reduce uncertainty about latent group-level properties from surveys and other collective assessments requires allocat

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

Whose Name Comes Up? II: Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation

DGX agent

arXiv:2602.08873v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are now used for academic expert recommendation. Existing audits typically evaluate such recommendations in isola

model-releasesarxiv-cs-ai
3 Jun 2026
← Previous
1…226227228229230…464
Next →