AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
1 Jul 2026

Robustness of Robotic Manipulation: Foundations and Frontiers

SafetyDGX agent

arXiv:2606.31494v1 Announce Type: cross Abstract: Humans and animals exhibit remarkable robustness in physical manipulation, yet robots remain far behind. Progress toward human-level manipulation robu

RoPoLL: Robust Panel of LLM Judges

Model ReleasesDGX agent

arXiv:2606.30931v1 Announce Type: new Abstract: The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its st

SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA

AgentsDGX agent

arXiv:2504.07385v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evaluati


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Same-Origin Policy for Agentic Browsers

Model ReleasesDGX agent

arXiv:2606.14027v3 Announce Type: replace-cross Abstract: Agentic browsers integrate autonomous AI agents into web browsers, enabling users to accomplish web tasks through natural-language instruction

Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records

AgentsDGX agent

arXiv:2606.31131v1 Announce Type: new Abstract: To ensure safe on-road behavior, pre-deployment testing and failure discovery of Autonomous Driving Systems (ADS) is crucial. Present day simulation bas

Scientific Explanations in Health Sciences: Causality, Trust, and Epistemic Adequacy

ResearchDGX agent

arXiv:2606.31616v1 Announce Type: new Abstract: Medical Artificial Intelligence (AI) is widely expected to transform clinical practice, yet the decision-making processes of many Machine Learning (ML)

Security--Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense

Model ReleasesDGX agent

arXiv:2606.30783v1 Announce Type: cross Abstract: We identify a security-fidelity tradeoff in defending LLMs against indirect prompt injection: defenses resist injected instructions largely by suppres

Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

SafetyDGX agent

arXiv:2606.31719v1 Announce Type: cross Abstract: In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through interaction. W

Seeing Through Multiple Views: Parameter-Efficient Fine-Tuning via Selective Neurons for Consistent Radiology Report Generation

Model ReleasesDGX agent

arXiv:2606.31099v1 Announce Type: cross Abstract: Recent years have seen substantial advances in radiology report generation (RRG), yet existing approaches predominantly adopt direct feature fusion wh

Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA

Local AiDGX agent

arXiv:2606.32002v1 Announce Type: new Abstract: Language models are increasingly taught from synthetic question--answer (QA) supervision: a model generates questions about a document, answers them fro

ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping

AgentsDGX agent

arXiv:2606.31693v1 Announce Type: cross Abstract: The wave of AI-native applications is moving shopping beyond page- and feed-based browsing toward intent-driven experiences orchestrated by LLM agents

SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos

Model ReleasesDGX agent

arXiv:2606.31127v1 Announce Type: cross Abstract: To enable personalized, real-time coaching using Augmented Reality glasses or fixed camera setups in domains such as sports, cooking, or music, a syst

Smart charging of large fleets of Electric Vehicles: Independent Multi-Agent Reinforcement Learning approaches

SafetyDGX agent

arXiv:2606.31347v1 Announce Type: new Abstract: The electrification of transportation through electric vehicles introduces new challenges for power grid management, such as increased peak demand, volt

Sparsity-Inducing Divergence Losses for Biometric Verification

ResearchDGX agent

arXiv:2606.31664v1 Announce Type: cross Abstract: Performance in face and speaker verification is largely driven by margin-penalty softmax losses such as CosFace and ArcFace. Recently introduced alpha

Spatial Reasoning via Modality Switching Between Language and Symbolic Representation

ResearchDGX agent

arXiv:2606.31285v1 Announce Type: new Abstract: Human reasoning is inherently multimodal: when problems become difficult, we rarely think in words alone. We often externalize our reasoning by sketchin

Stage-Transition Dense Reward Modeling for Reinforcement Learning

ResearchDGX agent

arXiv:2606.31377v1 Announce Type: cross Abstract: Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping si

STEB: Style Text Embedding Benchmark

Model ReleasesDGX agent

arXiv:2606.31741v1 Announce Type: cross Abstract: While semantic embeddings are rigorously evaluated on the Massive Text Embedding Benchmark, the evaluation of style embeddings remains fragmented, wit

Structural Preservation and the Logical Expressiveness of Graph Neural Networks

SafetyDGX agent

arXiv:2606.17882v2 Announce Type: replace Abstract: Bridges between graph neural networks (GNNs) and logical formalisms have been established by fixing architectural choices, such as the types of aggr

Surprise as a Signal for Plasticity and Metacognition

ResearchDGX agent

arXiv:2606.31495v1 Announce Type: new Abstract: We study a single idea across two settings: that a prediction-error signal, computed by a small predictor over the latent space of a frozen encoder, can

Surrogate-Gated Generation and Foundation-Model Embeddings for Bayesian Materials Design

Model ReleasesDGX agent

arXiv:2606.28578v1 Announce Type: cross Abstract: Closed-loop materials discovery iterates between proposing candidate structures and evaluating their properties, and property evaluation dominates the

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

ResearchDGX agent

arXiv:2606.31259v1 Announce Type: cross Abstract: Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative multi-step den

TDGT: A Tabular Data Generation Toolkit supporting adaptive GPU-accelerated Bayesian mixture models, diffusion-based models, and latent-space generative modeling

SafetyDGX agent

arXiv:2606.31268v1 Announce Type: cross Abstract: The growing demand for privacy-preserving data sharing has positioned synthetic data generation as a critical component of responsible AI workflows. D

Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles

Model ReleasesDGX agent

arXiv:2606.23672v2 Announce Type: replace Abstract: This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Manipulation Puzzles. In this tas

Team MKC at CLPsych 2026: Capturing and Characterizing Mental Health Changes through Social Media Timeline Dynamics

ResearchDGX agent

arXiv:2606.31464v1 Announce Type: cross Abstract: Recent advances in Large Language Models (LLMs) have motivated their adoption across a wide range of domains, including Artificial Intelligence (AI) f

Temperature Field Reconstruction of Tungsten Monoblock Divertor on EAST using Physics-aware Neural Operator Transformer

Model ReleasesDGX agent

arXiv:2606.31574v1 Announce Type: cross Abstract: Accurate modeling of the divertor temperature field is essential for preventing material melting and damage and for extending the service life of fusi

Temporal Preservation over Processing: Diagnosing and Designing Spatiotemporal Single-Stage Video Detectors

ResearchDGX agent

arXiv:2606.31421v1 Announce Type: cross Abstract: Single-stage video object detectors are increasingly deployed in time-critical applications, yet it remains unclear whether these models genuinely rea

Test-Time Verification for Text-to-SQL via Outcome Reward Models

SafetyDGX agent

arXiv:2606.30851v1 Announce Type: cross Abstract: Improving the reliability of large language models (LLMs) at inference time is a central challenge in structured reasoning tasks such as Text-to-SQL.

The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes

AgentsDGX agent

arXiv:2606.30653v1 Announce Type: cross Abstract: Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification

The HydroGym Reinforcement Learning Platform for Fluid Dynamics

ResearchDGX agent

arXiv:2512.17534v2 Announce Type: replace-cross Abstract: Modeling and controlling fluids is critical across science and engineering. Effective flow control can increase lift, reduce drag, enhance mix

The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning

ResearchDGX agent

arXiv:2606.30875v1 Announce Type: cross Abstract: Foundation model pseudo-labeling - labeling data strictly via zero-shot inference - enables massive scale, but performance is undermined by hallucinat

The Past Is Prologue: A Plug-in Controller for Selective Updates in Sequentially Evolving LLM Memory

SafetyDGX agent

arXiv:2606.31121v1 Announce Type: new Abstract: Sequentially evolving LLM memory enables agents to reuse past experience, but existing systems usually deploy each locally generated memory update witho

Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents

Model ReleasesDGX agent

arXiv:2606.31648v1 Announce Type: new Abstract: We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-English enterp

Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism

ResearchDGX agent

arXiv:2606.31222v1 Announce Type: new Abstract: Composed image retrieval requires identifying a target image from a gallery by integrating a reference image with a textual modification instruction. In

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning

SafetyDGX agent

arXiv:2606.31599v1 Announce Type: cross Abstract: Vision-language models (VLMs) combining reinforcement learning (RL) ignite remarkable progress in multimodal reasoning, yet still struggle with medica

Toward AI-Resilient Assessment in Computer Science Courses in an AI-Native World

TutorialsDGX agent

arXiv:2606.30655v1 Announce Type: cross Abstract: AI-native course assessments in senior computer science courses and related fields should grade students by AI-resilient skill: the ability to achieve

Towards Inclusive Mobility Modeling: Characterizing and Evaluating Elderly Trajectory Patterns in Urban Systems

Local AiDGX agent

arXiv:2606.31207v1 Announce Type: new Abstract: The rapid advance of smart cities increasingly depends on trajectory data mining, yet underrepresented demographic groups, particularly the elderly, are

Toxicity Assessment in Preclinical Histopathology via Class-Aware Mahalanobis Distance for Known and Novel Anomalies

SafetyDGX agent

arXiv:2602.02124v2 Announce Type: replace-cross Abstract: Drug-induced toxicity is a leading cause of preclinical and early-clinical failure, making early detection critical. Histopathology is the gol

TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels

SafetyDGX agent

arXiv:2504.12557v3 Announce Type: replace-cross Abstract: Ensuring safe behavior in reinforcement learning (RL) is challenging when safety constraints are implicit and cannot be densely measured. In m

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

SafetyDGX agent

arXiv:2606.30887v1 Announce Type: cross Abstract: Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control

Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation

SafetyDGX agent

arXiv:2606.31184v1 Announce Type: cross Abstract: Adaptive experiments for average treatment effects (ATE) require randomized allocations balancing valid inference with statistical efficiency. The ora

TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models

SafetyDGX agent

arXiv:2606.31976v1 Announce Type: new Abstract: Human-labeled data are widely used as reference annotations in ML, despite known variability across annotators in many expert-driven domains. In additio

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2606.32017v1 Announce Type: cross Abstract: Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and objec

Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks

ResearchDGX agent

arXiv:2606.31074v1 Announce Type: cross Abstract: Existing AI-generated text detectors are vulnerable to attacks that manipulate textual characteristics. In this study, we propose a novel Triospect De

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

Model ReleasesDGX agent

arXiv:2603.29759v2 Announce Type: replace-cross Abstract: Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing ben

Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens

Model ReleasesDGX agent

arXiv:2606.30755v1 Announce Type: cross Abstract: Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. They take on

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

ResearchDGX agent

arXiv:2606.31128v1 Announce Type: cross Abstract: Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-lev

UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation

SafetyDGX agent

arXiv:2606.31451v1 Announce Type: cross Abstract: Unified multimodal models (UMMs) have shown great promise in integrating understanding and generation across diverse modalities. However, existing res

Unsupervised Thermodynamics of Molecular Diffusion Models: Action-Operator Semantics and Auditable Free-Energy Readout

ResearchDGX agent

arXiv:2606.30687v1 Announce Type: cross Abstract: Diffusion models are increasingly utilized for modeling molecular structures and conformational ensembles, yet the thermodynamic meaning of their lear

Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection

ResearchDGX agent

arXiv:2502.15845v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often hallucinate, limiting their reliability in sensitive applications. In black-box settings, several self-cons

Visual Prompt Discovery via Semantic Exploration

AgentsDGX agent

arXiv:2603.16250v2 Announce Type: replace-cross Abstract: LVLMs encounter significant challenges in image understanding and visual reasoning, leading to critical perception failures. Visual prompts, w

Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?

ResearchDGX agent

arXiv:2606.31407v1 Announce Type: cross Abstract: Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based methods, such

Von Mises Based Uncertainty Quantification for Closely Spaced Automotive Radar Targets

ResearchDGX agent

arXiv:2606.31473v1 Announce Type: cross Abstract: This work investigates uncertainty-aware deep learning approaches for direction of arrival (DOA) estimation in automotive radar, focusing on probabili

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG

SafetyDGX agent

arXiv:2606.30989v1 Announce Type: cross Abstract: Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs

What Drives Interactive Improvement from Feedback?

AgentsDGX agent

arXiv:2606.30774v1 Announce Type: new Abstract: We study when natural-language feedback produces improvement beyond the gains obtainable from repeated attempts alone. In multi-turn language agent sett

What Probing Reveals about Autonomous Driving: Linking Internal Prediction Errors to Ego Planning

SafetyDGX agent

arXiv:2606.31106v1 Announce Type: cross Abstract: Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal

When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models

Model ReleasesDGX agent

arXiv:2606.30852v1 Announce Type: new Abstract: Reasoning models spend different amounts of useful computation across instances, but it remains unclear when a learned stopping rule improves over simpl

When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors

Model ReleasesDGX agent

arXiv:2606.32029v1 Announce Type: cross Abstract: While large language models (LLMs) perform well on table tasks, they still make data referencing errors (DREs), i.e., incorrectly citing or omitting t

When Regulation Has Memory: Hysteresis and Control Burden in Artificial Agency

AgentsDGX agent

arXiv:2606.30975v1 Announce Type: new Abstract: Adaptive agents are usually judged by what they do, but an agent can appear stable while the internal effort required to keep it stable is increasing. T

When Reranking Hurts: Uncertainty-Based Gating for Few-Shot Reranking

ResearchDGX agent

arXiv:2606.31087v1 Announce Type: cross Abstract: Few-shot selection typically assumes that reranking retrieved examples always improves performance. We challenge this view by identifying that the exp

When to Truncate a Feature Ranking: A Residual-Overlap Stopping Rule for Subset Selection

ResearchDGX agent

arXiv:2606.31686v1 Announce Type: cross Abstract: Feature rankings are widely used in supervised feature selection because they are simple, scalable and easy to interpret. Variables are first ranked b

← Previous
1…106107108109110…358
Next →