AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
14 May 2026

ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding

Model ReleasesDGX agent

arXiv:2605.13228v1 Announce Type: cross Abstract: Video understanding requires active evidence seeking, motivating tool-augmented video agents for temporal reasoning, cross-modal understanding, and co

Retrieval-Augmented Tutoring for Algorithm Tracing and Problem-Solving in AI Education

TutorialsDGX agent

arXiv:2605.12988v1 Announce Type: new Abstract: Students learning algorithms often need support as they interpret traces, debug reasoning errors, and apply procedures across unfamiliar problem instanc

Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation

TutorialsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.12975v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become a standard approach for knowledge-intensive question answering, but existing systems remain brittle on m

Revealing Interpretable Failure Modes of VLMs

SafetyDGX agent

arXiv:2605.12674v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly used in safety-critical applications because of their broad reasoning capabilities and ability to general

Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency

SafetyDGX agent

arXiv:2605.13047v1 Announce Type: cross Abstract: Evaluating whether large vision-language models (VLMs) align with human perception for high-level semantic scene comprehension remains a challenge. Tr

Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding

TutorialsDGX agent

arXiv:2506.09522v3 Announce Type: replace-cross Abstract: Large Vision Language Models (LVLMs) achieve strong performance across multimodal tasks by integrating visual perception with language underst

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective

Model ReleasesDGX agent

arXiv:2605.12969v1 Announce Type: cross Abstract: RLVR has become a widely adopted paradigm for improving LLMs' reasoning capabilities, and GRPO is one of its most representative algorithms. In this p

RISED: A Pre-Deployment Safety Evaluation Framework for Clinical AI Decision-Support Systems

Model ReleasesDGX agent

arXiv:2605.12895v1 Announce Type: cross Abstract: Aggregate accuracy metrics dominate the evaluation of clinical AI decision-support systems but do not detect deployment-phase failures of input reliab

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations

ApplicationsDGX agent

arXiv:2507.00990v3 Announce Type: replace-cross Abstract: This work introduces Robots Imitating Generated Videos (RIGVid), a system that enables robots to perform complex manipulation tasks--such as p

Robust and Explainable Bicuspid Aortic Valve Diagnosis Using Stacked Ensembles on Echocardiography

ResearchDGX agent

arXiv:2605.13730v1 Announce Type: cross Abstract: Transthoracic echocardiography (TTE) is the first-line imaging modality for diagnosing bicuspid aortic valve (BAV), yet diagnostic performance varies

RS-Claw: Progressive Active Tool Exploration via Hierarchical Skill Trees for Remote Sensing Agents

Model ReleasesDGX agent

arXiv:2605.13391v1 Announce Type: new Abstract: The rise of multi-modal large language models (MLLMs) is shifting remote sensing (RS) intelligence from 'see' to 'action', as OpenClaw-style frameworks

RTLC -- Research, Teach-to-Learn, Critique: A three-stage prompting paradigm inspired by the Feynman Learning Technique that lifts LLM-as-judge accuracy on JudgeBench with no fine-tuning

Model ReleasesDGX agent

arXiv:2605.13695v1 Announce Type: cross Abstract: LLM-as-a-judge is now the default measurement instrument for open-ended generation, but on the public JudgeBench benchmark even strong instruction-tun

Scale-Gest: Scalable Model-Space Synthesis and Runtime Selection for On-Device Gesture Detection

Local AiDGX agent

arXiv:2605.12506v1 Announce Type: cross Abstract: Realizing on-device ML-based gesture detection under tight real-time performance, energy and memory constraints is challenging, especially when consid

Scale over Preference: The Impact of AI-Generated Content on Online Content Ecology

ApplicationsDGX agent

arXiv:2604.01690v2 Announce Type: replace Abstract: The rapid proliferation of Artificial Intelligence-Generated Content (AIGC) is fundamentally restructuring online content ecologies, necessitating a

Scaling few-shot spoken word classification with generative meta-continual learning

TutorialsDGX agent

arXiv:2605.13075v1 Announce Type: cross Abstract: Few-shot spoken word classification has largely been developed for applications where a small number of classes is considered, and so the potential of

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs

Model ReleasesDGX agent

arXiv:2510.18245v3 Announce Type: replace-cross Abstract: Scaling the number of parameters and the size of training data has proven to be an effective strategy for improving large language model (LLM)

Scaling Retrieval-Augmented Reasoning with Parallel Search and Explicit Merging

AgentsDGX agent

arXiv:2605.13534v1 Announce Type: new Abstract: Deep search agents have proven effective in enhancing LLMs by retrieving external knowledge during multi-step reasoning. However, existing methods often

ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles

SafetyDGX agent

arXiv:2605.13725v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent simulation offers a powerful testbed for studying social opinion dynamics. Yet current approaches often ado

SECOND-Grasp: Semantic Contact-guided Dexterous Grasping

SafetyDGX agent

arXiv:2605.13117v1 Announce Type: cross Abstract: Achieving reliable robotic manipulation, such as dexterous grasping, requires a synergy between physically stable interactions and semantic task guida

Seg-Agent: Test-Time Multimodal Reasoning for Training-Free Language-Guided Segmentation

Model ReleasesDGX agent

arXiv:2605.12953v1 Announce Type: cross Abstract: Language-guided segmentation transcends the scope limitations of traditional semantic segmentation, enabling models to segment arbitrary target region

Selective Off-Policy Reference Tuning with Plan Guidance

SafetyDGX agent

arXiv:2605.11505v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT a

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation

SafetyDGX agent

arXiv:2605.13554v1 Announce Type: cross Abstract: Contrastive reinforcement learning (CRL) learns goal-conditioned Q-values through a contrastive objective over state-action and goal representations,

Semantic knowledge guides innovation and drives cultural evolution

AgentsDGX agent

arXiv:2510.12837v3 Announce Type: replace-cross Abstract: Cultural evolution allows ideas and technologies to accumulate across generations, reaching their most complex and open-ended form in humans.

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs

Model ReleasesDGX agent

arXiv:2605.13737v1 Announce Type: new Abstract: When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perc

Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators

ResearchDGX agent

arXiv:2605.12748v1 Announce Type: cross Abstract: Large language models (LLMs) can fluently generate student-like responses, making them attractive as simulated students for training and evaluating AI

SP-GCRL: Influence Maximization on Incomplete Social Graphs

SafetyDGX agent

arXiv:2605.12513v1 Announce Type: cross Abstract: Influence maximization (IM) in real platforms is challenged by incomplete, noisy social graphs and non-stationary diffusion dynamics. We propose SP-GC

Spectral Flattening Is All Muon Needs: How Orthogonalization Controls Learning Rate and Convergence

ResearchDGX agent

arXiv:2605.13079v1 Announce Type: cross Abstract: Muon orthogonalizes the momentum buffer before each update, replacing its singular values with ones via Newton-Schulz iterations. This simple change l

SpectralTrain: A Universal Framework for Hyperspectral Image Classification

Model ReleasesDGX agent

arXiv:2511.16084v2 Announce Type: replace-cross Abstract: Hyperspectral image (HSI) classification typically involves large-scale data and computationally intensive training, which limits the practica

SPOT: Selective Prompt Projection via Total Variation for Inference-Only Safe Text-to-Image Generation

SafetyDGX agent

arXiv:2602.00616v3 Announce Type: replace Abstract: Text-to-Image (T2I) diffusion models enable high quality open ended synthesis, but practical use requires suppressing unsafe generations while prese

SSDA: Bridging Spectral and Structural Gaps via Dual Adaptation for Vision-Based Time Series Forecasting

ApplicationsDGX agent

arXiv:2605.12550v1 Announce Type: cross Abstract: Large vision models (LVMs) have recently proven to be surprisingly effective time series forecasters, simply by rendering temporal data as images. Thi

Stable Attention Response for Reliable Precipitation Nowcasting

ResearchDGX agent

arXiv:2605.13181v1 Announce Type: cross Abstract: Precipitation nowcasting remains challenging due to the highly localized, rapidly evolving, and heterogeneous nature of atmospheric dynamics. Although

STAR: Semantic-Temporal Adaptive Representation Learning for Few-Shot Action Recognition

SafetyDGX agent

arXiv:2605.13202v1 Announce Type: cross Abstract: Few-shot action recognition (FSAR) requires models to generalize to novel action categories from only a handful of annotated samples. Despite progress

State-Centric Decision Process

AgentsDGX agent

arXiv:2605.12755v1 Announce Type: new Abstract: Language environments such as web browsers, code terminals, and interactive simulations emit raw text rather than states, and provide none of the runtim

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos

Model ReleasesDGX agent

arXiv:2512.01707v3 Announce Type: replace-cross Abstract: Streaming video understanding requires models not only to process temporally incoming frames, but also to anticipate user intention for realis

Stress-Testing the Reasoning Competence of LLMs With Proofs Under Minimal Formalism

Model ReleasesDGX agent

arXiv:2605.12524v1 Announce Type: cross Abstract: We introduce ProofGrid, a benchmark suite for evaluating LLM reasoning through machine-checkable proofs rather than final answers alone. ProofGrid con

Strikingness-Aware Evaluation for Temporal Knowledge Graph Reasoning

ResearchDGX agent

arXiv:2605.13153v1 Announce Type: new Abstract: Temporal Knowledge Graph Reasoning (TKGR) aims at inferring missing (especially future) events from historical data. Current evaluation in TKGR uniforml

Stylized Text-to-Motion Generation via Hypernetwork-Driven Low-Rank Adaptation

ResearchDGX agent

arXiv:2605.13333v1 Announce Type: cross Abstract: Text-driven motion diffusion models are capable of generating realistic human motions, but text alone often struggles to express fine-level nuances of

SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management

Model ReleasesDGX agent

arXiv:2602.07342v2 Announce Type: replace Abstract: Large language models (LLMs) have shown promise in complex reasoning and tool-based decision making, motivating their application to real-world supp

Sustaining AI safety: Control-theoretic external impossibility, intrinsic necessity, and structural requirements

SafetyDGX agent

arXiv:2605.12963v1 Announce Type: new Abstract: As AI systems become increasingly capable, safety strategies must be evaluated not only by how much they reduce present risk, but by whether they could

SynCABEL: Synthetic Contextualized Augmentation for Biomedical Entity Linking

Model ReleasesDGX agent

arXiv:2601.19667v2 Announce Type: replace-cross Abstract: We present SynCABEL (Synthetic Contextualized Augmentation for Biomedical Entity Linking), a framework that addresses a central bottleneck in

Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs

Model ReleasesDGX agent

arXiv:2505.11556v4 Announce Type: replace-cross Abstract: Multi-agent systems built on large language models (LLMs) are expected to enhance decision-making by pooling distributed information, yet syst

T-TExTS (Teaching Text Expansion for Teacher Scaffolding): Enhancing Text Selection in High School Literature through Knowledge Graph-Based Recommendation

ResearchDGX agent

arXiv:2506.12075v2 Announce Type: replace-cross Abstract: High school English Literature teachers often encounter barriers to assembling diverse, thematically aligned text sets due to limited planning

Table-R1: Region-based Reinforcement Learning for Table Understanding

Model ReleasesDGX agent

arXiv:2505.12415v3 Announce Type: replace-cross Abstract: Tables present unique challenges for language models due to their structured row-column interactions, necessitating specialized approaches for

Teacher-Guided Policy Optimization for LLM Distillation

SafetyDGX agent

arXiv:2605.13230v1 Announce Type: cross Abstract: The convergence of reinforcement learning and imitation learning has positioned Reverse KL (RKL) as a promising paradigm for on-policy LLM distillatio

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment

SafetyDGX agent

arXiv:2605.13537v1 Announce Type: cross Abstract: Inference-time alignment techniques offer a lightweight alternative or complement to costly reinforcement learning, while enabling continual adaptatio

The Bayesian Origin of the Probability Weighting Function in Human Representation of Probabilities

ResearchDGX agent

arXiv:2510.04698v3 Announce Type: replace-cross Abstract: Humans systematically misrepresent probability in a stereotyped inverse-S pattern. It has been documented for decades, but its origin remains

The critical slowing down in diffusion models

Model ReleasesDGX agent

arXiv:2605.12597v1 Announce Type: cross Abstract: Computational sampling has been central to the sciences since the mid-20th century. While machine-learning-based approaches have recently enabled majo

The End Justifies the Mean: A Linear Ranking Rule for Proportional Sequential Decisions

SafetyDGX agent

arXiv:2605.12717v1 Announce Type: cross Abstract: AI alignment and participatory design motivate a new democratic design problem: how to collectively choose a decision rule to use repeatedly. We study

The Expressivity Boundary of Probabilistic Circuits: A Comparison with Large Language Models

ResearchDGX agent

arXiv:2605.12940v1 Announce Type: cross Abstract: Probabilistic Circuits (PCs) are deep generative models that support exact and efficient probabilistic inference. Yet in autoregressive language model

The Readability Spectrum: Patterns, Issues, and Prompt Effects in LLM-Generated Code

ResearchDGX agent

arXiv:2605.13280v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are transforming software development, the functional quality of generated code has become a central focus, leaving re

The WidthWall: A Strict Expressivity Hierarchy for Hypergraph Neural Networks

TutorialsDGX agent

arXiv:2605.13690v1 Announce Type: cross Abstract: Hypergraphs provide a natural framework to model higher-order interactions in scientific, social, and biological systems. Hypergraph neural networks (

Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents

SafetyDGX agent

arXiv:2605.12620v1 Announce Type: new Abstract: Building generalist embodied agents capable of solving complex real-world tasks remains a fundamental challenge in AI. Multimodal Large Language Models

TiCo: Time-Controllable Spoken Dialogue Model

Model ReleasesDGX agent

arXiv:2603.22267v2 Announce Type: replace-cross Abstract: We introduce TiCo, a time-controllable spoken dialogue model (SDM) that follows time-constrained instructions (e.g., 'Please generate a respon

TimelineReasoner: Advancing Timeline Summarization with Large Reasoning Models

ResearchDGX agent

arXiv:2605.12518v1 Announce Type: cross Abstract: The proliferation of online news poses a challenge to extracting structured timelines from unstructured content. While recent studies have shown that

TokaMind for Power Grid: Cross-Domain Transfer from Fusion Plasma

Model ReleasesDGX agent

arXiv:2605.11033v1 Announce Type: cross Abstract: TokaMind is a multi-modal transformer (MMT) foundation model pre-trained on tokamak plasma diagnostics data from MAST, where it was shown to outperfor

ToolWeave: Structured Synthesis of Complex Multi-Turn Tool-Calling Dialogues

Model ReleasesDGX agent

arXiv:2605.12521v1 Announce Type: cross Abstract: Multi-turn tool calling is essential for LLMs to function as autonomous agents, yet synthesizing the training data required for these capabilities rem

Topology-Preserving Neural Operator Learning via Hodge Decomposition

SafetyDGX agent

arXiv:2605.13834v1 Announce Type: cross Abstract: In this paper, we study solution operators of physical field equations on geometric meshes from a function-space perspective. We reveal that Hodge ort

Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents

Model ReleasesDGX agent

arXiv:2602.16246v3 Announce Type: replace Abstract: Interactive large language model (LLM) agents operating via multi-turn dialogue and multi-step tool calling are increasingly used in production. Ben

Towards a holistic understanding of Selection Bias for Causal Effect Identification

SafetyDGX agent

arXiv:2605.13430v1 Announce Type: cross Abstract: Selection bias is pervasive in observational studies. For example, large scale biobanks data can exhibit ``healthy volunteer bias'' when respondents a

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.13119v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are effective robot action executors, but they remain limited on long-horizon tasks due to the dual burden of exte

← Previous
1…260261262263264…358
Next →