AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
10 Jul 2026

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

Model ReleasesDGX agent

arXiv:2607.08535v1 Announce Type: cross Abstract: An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evaluator-replace

When Thinking Hurts: Epistemic Signals in the Reasoning Chains of Visual Language Models

Model ReleasesDGX agent

arXiv:2607.08059v1 Announce Type: cross Abstract: Uncertainty quantification for visual language models (VLMs) conventionally targets the answer token distribution. We provide the first three-family e

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.08054v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs

Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems

Local AiDGX agent

arXiv:2607.07989v1 Announce Type: cross Abstract: Large language model (LLM) based multi-agent systems enable complex problem solving through coordinated reasoning and action, but their distributed st

Workflow as Knowledge: Semantic Persistence for LLM-Mediated Workflows

SafetyDGX agent

arXiv:2607.08740v1 Announce Type: new Abstract: Large language model (LLM) applications increasingly use explicit workflows for tool use, retrieval, branching, checkpointing, and human approval. Exist

XFACTORS: Disentangled Information Bottleneck via Contrastive Supervision

SafetyDGX agent

arXiv:2601.21688v2 Announce Type: replace-cross Abstract: Disentangled representation learning aims to map independent factors of variation to independent representation components. On one hand, purel

9 Jul 2026

A Continual Learning Framework for Adaptive Control of Modular Soft Robots

TutorialsDGX agent

arXiv:2607.06740v1 Announce Type: cross Abstract: Soft robots have attracted significant attention in applications such as medical intervention, rehabilitation, and robotic manipulation due to their i

A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong

Model ReleasesDGX agent

arXiv:2607.06854v1 Announce Type: cross Abstract: Reinforcement learning agents for imperfect-information card games are only as strong as the opponents they train against, and they are hard to grade,

A Multi-Analyst LLM Pipeline for Auditable Rule Discovery Across 68 Public Physiological Corpora

ResearchDGX agent

arXiv:2607.06802v1 Announce Type: cross Abstract: Open physiological corpora are heterogeneous: they use different sensors, labels, sampling rates, recording settings, and clinical endpoints. They can

A Study of Commonsense Reasoning over Visual Object Properties

Model ReleasesDGX agent

arXiv:2508.10956v3 Announce Type: replace-cross Abstract: Inspired by human categorization, visual reasoning about object properties, such as physical attributes and functions, involves identifying an

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

ResearchDGX agent

arXiv:2607.07708v1 Announce Type: cross Abstract: Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge

Ad Headline Generation using Self-Critical Masked Language Model

SafetyDGX agent

arXiv:2607.06818v1 Announce Type: cross Abstract: For any E-commerce website it is a nontrivial problem to build enduring advertisements that attract shoppers. It is hard to pass the creative quality

AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org

SafetyDGX agent

arXiv:2512.11935v2 Announce Type: replace Abstract: Agentic AI systems increasingly connect large language models (LLMs) to external scientific tools, yet whether and when tool access improves predict

Agentic Data Environments

SafetyDGX agent

arXiv:2607.07397v1 Announce Type: new Abstract: Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs. Th

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.06624v1 Announce Type: new Abstract: We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the ta

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

Model ReleasesDGX agent

arXiv:2607.07690v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

Model ReleasesDGX agent

arXiv:2602.05088v4 Announce Type: replace Abstract: Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for mental heal

AirPASS: Over-the-Air Federated Learning via Pinching Antenna Systems

ResearchDGX agent

arXiv:2607.06768v1 Announce Type: cross Abstract: This paper investigates over-the-air federated learning (AirFL) in wireless systems where the access point is equipped with a multi-waveguide pinching

ALER-TI: Aligned Latent Embedding Retrieval for Time Series Imputation

Local AiDGX agent

arXiv:2607.07640v1 Announce Type: cross Abstract: Deep learning has significantly advanced time series imputation, yet most existing architectures primarily rely on localized temporal context within t

An Adaptive Differentially Private Federated Learning Framework

TutorialsDGX agent

arXiv:2602.06838v3 Announce Type: replace Abstract: Federated learning enables collaborative model training across distributed clients while preserving data privacy. However, in practical deployments,

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning

Local AiDGX agent

arXiv:2607.07033v1 Announce Type: cross Abstract: Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are

Anomaly detection in time-series via inductive biases in the latent space of conditional normalizing flows

ApplicationsDGX agent

arXiv:2603.11756v2 Announce Type: replace Abstract: Deep generative models for anomaly detection in multivariate time-series are typically trained by maximizing observed data likelihood. However, like

AT-Attn: Temporal-Aware Cross-Attention for Longitudinal Multimodal Alzheimer's Disease Diagnosis

ResearchDGX agent

arXiv:2607.07091v1 Announce Type: cross Abstract: In longitudinal Alzheimer's disease (AD) diagnosis support, clinical and imaging information is often collected at irregular visits. Integrating these

At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics

Model ReleasesDGX agent

arXiv:2607.06639v1 Announce Type: cross Abstract: On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized. Reading effective rank

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts

ResearchDGX agent

arXiv:2607.06611v1 Announce Type: cross Abstract: Automatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal inflections and

Bayesian Optimization of Genetic Algorithm Hyperparameters in a Multi-Fidelity Framework for Efficient Lattice Material Design

TutorialsDGX agent

arXiv:2607.07289v1 Announce Type: cross Abstract: This study presents a multi-fidelity framework for the systematic optimization of genetic algorithm (GA) hyperparameters. The framework integrates thr

Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report

SafetyDGX agent

arXiv:2607.07370v1 Announce Type: cross Abstract: In embodied intelligence systems, the motion controller serves as the critical bridge between semantic reasoning and physical execution. Humanoid cont

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

Model ReleasesDGX agent

arXiv:2607.07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that th

Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass

Model ReleasesDGX agent

arXiv:2607.07696v1 Announce Type: cross Abstract: Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely by the datab

ButterflyMoE: Compression-Scalable Ternary Experts via Structured Butterfly Orbits

ResearchDGX agent

arXiv:2601.13563v5 Announce Type: replace-cross Abstract: In current Mixture of Experts (MoE) architectures, linear memory scaling is present, the memory grows as the number of experts increases. N in

Can Reinforcement Learning Efficiently Discover Price Manipulation?

Model ReleasesDGX agent

arXiv:2607.06121v1 Announce Type: cross Abstract: In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditio

Can We Really Learn One Representation to Optimize All Rewards?

TutorialsDGX agent

arXiv:2602.11399v2 Announce Type: replace-cross Abstract: As unsupervised pretraining becomes increasingly ubiquitous in reinforcement learning, a more thorough theoretical understanding of these meth

CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training

ResearchDGX agent

arXiv:2607.07292v1 Announce Type: cross Abstract: Accurately estimating urban carbon emissions is critical for sustainable urban planning, yet many existing approaches remain difficult to apply consis

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis

SafetyDGX agent

arXiv:2607.07601v1 Announce Type: cross Abstract: Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize co

Co-LMLM: Continuous-Query Limited Memory Language Models

Model ReleasesDGX agent

arXiv:2607.07707v1 Announce Type: cross Abstract: Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it in their w

Collaborative Synthetic Data Generation for Knowledge Transfer in Federated Learning

ResearchDGX agent

arXiv:2607.07565v1 Announce Type: cross Abstract: One-shot federated learning (OSFL) addresses the communication overhead of federated learning by limiting training to a single round, but doing so wit

CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation

Model ReleasesDGX agent

arXiv:2603.16551v2 Announce Type: replace-cross Abstract: Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that ge

Complexity-Budgeted, Interaction-Aware Interpretable Model for Tabular Data

ApplicationsDGX agent

arXiv:2607.07060v1 Announce Type: cross Abstract: Inherently interpretable classifiers for tabular data typically rely on sparse features, rules, or patterns that users can inspect directly. The margi

Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

ResearchDGX agent

arXiv:2607.06940v1 Announce Type: cross Abstract: The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their respon

Computing with Stochastic Oracles in AI-Augmented Computation

ResearchDGX agent

arXiv:2607.06893v1 Announce Type: cross Abstract: The Stochastic-Oracle Turing Machine (SOTM) framework models AI-augmented computation as the interaction of a probabilistic Turing machine with an ora

ContrastiveCFG: Guiding Diffusion Sampling by Contrasting Positive and Negative Concepts

SafetyDGX agent

arXiv:2411.17077v2 Announce Type: replace-cross Abstract: As Classifier-Free Guidance (CFG) has proven effective in conditional diffusion model sampling for improved condition alignment, many applicat

Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

Model ReleasesDGX agent

arXiv:2607.06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary

Creativity from Friction: Human-AI Interaction for Exploratory Structural Design

ResearchDGX agent

arXiv:2607.07521v1 Announce Type: cross Abstract: AI agents that generate final answers based on user input often do not meet the needs of creative fields. Fields such as structural design and archite

Cross-Trajectory Chimera Interventions Reveal Dissociable Roles of Weight Magnitude and Direction in Grokking

ResearchDGX agent

arXiv:2607.06628v1 Announce Type: cross Abstract: Which properties of a partially trained network are causally portable to a different, independently trained network? Single-trajectory interventions s

D2PO: Optimizing Diffusion Samplers via Dynamic Preference

SafetyDGX agent

arXiv:2607.06609v1 Announce Type: cross Abstract: We propose D2PO (Dynamic Direct Preference Optimization), a principled framework for optimizing diffusion sampling policies with respect to timestep s

DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression

SafetyDGX agent

arXiv:2603.15685v2 Announce Type: replace-cross Abstract: Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make in

Deep Reinforcement Learning for Reliability Based Bi-Objective Portfolio Optimization

SafetyDGX agent

arXiv:2607.06610v1 Announce Type: cross Abstract: Portfolio optimization under uncertainty is inherently a multi-objective decision problem involving complex interactions among return, risk, market dy

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

SafetyDGX agent

arXiv:2607.07669v1 Announce Type: cross Abstract: Large language models increasingly understand dialectal English, yet still produce only standard, US-leaning English, leaving dialectal generation, th

Diffusion enabled Optimal Transport distances for graph matching

ResearchDGX agent

arXiv:2607.06646v1 Announce Type: cross Abstract: This paper introduces Diffusion Semi-Relaxed Fused Gromov-Wasserstein (DsrFGW), a novel method for graph comparison that unifies node features and str

Digital Fragmentation and Generative AI Use Across 103 Million Application Events

ResearchDGX agent

arXiv:2607.06681v1 Announce Type: cross Abstract: Knowledge workers switch between applications thousands of times per day, spending nearly a tenth of the work year transitioning between digital appli

DiPhon: Diffusion on Graphons for Scalable Graph Generation

ResearchDGX agent

arXiv:2607.07232v1 Announce Type: cross Abstract: Diffusion models represent a leading paradigm for graph generation, with notable impact in domains such as molecular design. Yet, scaling these models

Do Counterfactually Fair Image Classifiers Satisfy Group Fairness? -- A Theoretical and Empirical Study

SafetyDGX agent

arXiv:2607.06603v1 Announce Type: cross Abstract: The notion of algorithmic fairness has been actively explored from various aspects of fairness, such as counterfactual fairness (CF) and group fairnes

Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows

ResearchDGX agent

arXiv:2607.07504v1 Announce Type: new Abstract: Product data scientists often ask LLM-based agents to help with recurring execution tasks such as cleaning data, writing SQL, choosing statistical tests

Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

Model ReleasesDGX agent

arXiv:2607.07189v1 Announce Type: new Abstract: Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the

Dynamic-in-Few-Step: Unifying Dynamic Computation and Few-Step Distillation for Efficient Video Generation

ApplicationsDGX agent

arXiv:2607.06631v1 Announce Type: cross Abstract: Video Diffusion Models (VDMs) have demonstrated superior generation quality but suffer from prohibitive computational costs. While recent few-step dis

Effective Strategies for Asynchronous Software Engineering Agents

AgentsDGX agent

arXiv:2603.21489v2 Announce Type: replace-cross Abstract: AI agents have become increasingly capable at isolated software engineering (SWE) tasks such as resolving issues on Github. Yet long-horizon t

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

SafetyDGX agent

arXiv:2602.23802v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture th

End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent

AgentsDGX agent

arXiv:2607.06964v1 Announce Type: cross Abstract: Bridging the gap between human pilot intent and autonomous flight operation is critical for real-world electric vertical takeoff and landing (eVTOL) a

Enhancing deep learning models for time series classification via knowledge distillation

Model ReleasesDGX agent

arXiv:2607.06796v1 Announce Type: cross Abstract: Deep learning has achieved remarkable success in various domains including time series analysis, computer vision and natural language processing. Howe

Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2607.07178v1 Announce Type: cross Abstract: Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. However, exis

← Previous
1…7677787980…358
Next →