AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
30 Apr 2026

Assessing the Utility of Volumetric Motion Fields for Radar-based Precipitation Nowcasting with Physics-informed Deep Learning

ResearchDGX agent

arXiv:2603.13589v2 Announce Type: replace-cross Abstract: Estimating motion from spatiotemporal geoscientific data is a fundamental component of many environmental modeling and forecasting tasks. In t

ATLAS: An Annotation Tool for Long-horizon Robotic Action Segmentation

SafetyDGX agent

arXiv:2604.26637v1 Announce Type: cross Abstract: Annotating long-horizon robotic demonstrations with precise temporal action boundaries is crucial for training and evaluating action segmentation and

Atomic-Probe Governance for Skill Updates in Compositional Robot Policies

SafetyDGX agent

arXiv:2604.26689v1 Announce Type: cross Abstract: Skill libraries in deployed robotic systems are continually updated through fine-tuning, fresh demonstrations, or domain adaptation, yet existing type


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Auditing Marketing Budget Allocation with Hindsight Regret

Model ReleasesDGX agent

arXiv:2604.25977v1 Announce Type: cross Abstract: Organizations routinely make strategic budget allocations under operational constraints, but often lack a principled way to assess whether realized al

Auto-ARGUE: LLM-Based Report Generation Evaluation

ResearchDGX agent

arXiv:2509.26184v5 Announce Type: replace-cross Abstract: Generation of citation-backed reports is a primary use case for retrieval-augmented generation (RAG) systems. While open-source evaluation too

Auto-Relational Reasoning

ResearchDGX agent

arXiv:2604.26507v1 Announce Type: new Abstract: Background & Objectives: In the last decade, Machine learning research has grown rapidly, but large models are reaching their soft limits demonstrating

Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth Retrieval

AgentsDGX agent

arXiv:2601.13969v2 Announce Type: replace Abstract: Retrieving evidence for language model queries from knowledge graphs requires balancing broad search across the graph with multi-hop traversal to fo

Benchmarking Complex Multimodal Document Processing Pipelines: A Unified Evaluation Framework for Enterprise AI

Model ReleasesDGX agent

arXiv:2604.26382v1 Announce Type: cross Abstract: Most enterprise document AI today is a pipeline. Parse, index, retrieve, generate. Each of those stages has been studied to death on its own -- what's

Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

SafetyDGX agent

arXiv:2604.26577v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants, yet their safety in this

Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex

Model ReleasesDGX agent

arXiv:2604.14858v2 Announce Type: replace Abstract: As agent systems move into increasingly diverse execution settings, trajectory-level safety evaluation and diagnosis require benchmarks that evolve

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Model ReleasesDGX agent

arXiv:2508.04325v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However,

Bian Que: An Agentic Framework with Flexible Skill Arrangement for Online System Operations

AgentsDGX agent

arXiv:2604.26805v1 Announce Type: new Abstract: Operating and maintaining (O&M) large-scale online engine systems (search, recommendation, advertising) demands substantial human effort for release mon

Breaking the Autoregressive Chain: Hyper-Parallel Decoding for Efficient LLM-Based Attribute Value Extraction

ResearchDGX agent

arXiv:2604.26209v1 Announce Type: cross Abstract: Some text generation tasks, such as Attribute Value Extraction (AVE), require decoding multiple independent sequences from the same document context.

Bridging Visual and Wireless Sensing via a Unified Radiation Field for 3D Radio Map Construction

ResearchDGX agent

arXiv:2601.19216v2 Announce Type: replace-cross Abstract: The emerging applications of next-generation wireless networks demand high-fidelity environmental intelligence. 3D radio maps bridge physical

Calibrated Surprise: An Information-Theoretic Account of Creative Quality

Model ReleasesDGX agent

arXiv:2604.26269v1 Announce Type: cross Abstract: The essence of good creative writing is calibrated surprise: when constraints from all relevant dimensions act together, the feasible solution space c

Causal Learning with Neural Assemblies

TutorialsDGX agent

arXiv:2604.26919v1 Announce Type: cross Abstract: Can Neural Assemblies -- groups of neurons that fire together and strengthen through co-activation -- learn the direction of causal influence between

Ceci n'est pas une explication: Evaluating Explanation Failures as Explainability Pitfalls in Language Learning Systems

Model ReleasesDGX agent

arXiv:2604.26145v1 Announce Type: cross Abstract: AI-powered language learning tools increasingly provide instant, personalised feedback to millions of learners worldwide. However, this feedback can f

CheXthought: A global multimodal dataset of clinical chain-of-thought reasoning and visual attention for chest X-ray interpretation

ApplicationsDGX agent

arXiv:2604.26288v1 Announce Type: cross Abstract: Chest X-ray interpretation is one of the most frequently performed diagnostic tasks in medicine and a primary target for AI development, yet current v

ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language Agents

Model ReleasesDGX agent

arXiv:2412.13682v5 Announce Type: replace Abstract: Travel planning stands out among real-world applications of Language Agents because it couples significant practical demand with a rigorous constrai

ClawGym: A Scalable Framework for Building Effective Claw Agents

Model ReleasesDGX agent

arXiv:2604.26904v1 Announce Type: cross Abstract: Claw-style environments support multi-step workflows over local files, tools, and persistent workspace states. However, scalable development around th

Co-Learning Port-Hamiltonian Systems and Optimal Energy-Shaping Control

SafetyDGX agent

arXiv:2604.26172v1 Announce Type: cross Abstract: We develop a physics-informed learning framework for energy-shaping control of port-Hamiltonian (pH) systems from trajectory data. The proposed approa

CoFL: Continuous Flow Fields for Language-Conditioned Navigation

Local AiDGX agent

arXiv:2603.02854v2 Announce Type: replace-cross Abstract: Existing language-conditioned navigation systems typically rely on modular pipelines or trajectory generators, but the latter use each scene--

ComboStoc: Combinatorial Stochasticity for Diffusion Generative Models

ResearchDGX agent

arXiv:2405.13729v3 Announce Type: replace-cross Abstract: In this paper, we study an under-explored but important factor of diffusion generative models, i.e., the combinatorial complexity. Data sample

Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models

Model ReleasesDGX agent

arXiv:2604.25922v1 Announce Type: cross Abstract: We present DenialBench, a systematic benchmark measuring consciousness denial behaviors across 115 large language models from 25+ providers. Using a t

Consist-Retinex: One-Step Noise-Emphasized Consistency Training Accelerates High-Quality Retinex Enhancement

SafetyDGX agent

arXiv:2512.08982v2 Announce Type: replace-cross Abstract: Retinex-based low-light image enhancement benefits from separating reflectance and illumination, yet recent generative approaches often rely o

Correcting Performance Estimation Bias in Imbalanced Classification with Minority Subconcepts

SafetyDGX agent

arXiv:2604.26024v1 Announce Type: cross Abstract: Class-level evaluation can conceal substantial performance disparities across subconcepts within the same class, causing models that perform well on a

Culturally Aware GenAI Risks for Youth: Perspectives from Youth, Parents, and Teachers in a Non-Western Context

SafetyDGX agent

arXiv:2604.26494v1 Announce Type: cross Abstract: Generative AI tools are widely used by youth and have introduced new privacy and safety challenges. While prior research has explored youth's safety i

Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods

ResearchDGX agent

arXiv:2505.13518v2 Announce Type: replace-cross Abstract: Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing predict

Data-Centric Foundation Models in Computational Healthcare: A Survey

SafetyDGX agent

arXiv:2401.02458v3 Announce Type: replace-cross Abstract: The advent of foundation models (FMs) as an emerging suite of AI techniques has struck a wave of opportunities in computational healthcare. Th

DC-Ada: Reward-Only Decentralized Sensor Adaptation for Heterogeneous Multi-Robot Teams

SafetyDGX agent

arXiv:2604.03905v2 Announce Type: replace-cross Abstract: Heterogeneity is a defining feature of deployed multi-robot teams: platforms often differ in sensing modalities, ranges, fields of view, and f

Delineating Knowledge Boundaries for Honest Large Vision-Language Models

ResearchDGX agent

arXiv:2604.26419v1 Announce Type: cross Abstract: Large Vision-Language Models (VLMs) have achieved remarkable multimodal performance yet remain prone to factual hallucinations, particularly in long-t

DepthPilot: From Controllability to Interpretability in Colonoscopy Video Generation

Model ReleasesDGX agent

arXiv:2604.26232v1 Announce Type: cross Abstract: Controllable medical video generation has achieved remarkable progress, but it still lacks interpretability, which requires the alignment of generated

Deterministic Legal Agents: A Canonical Primitive API for Auditable Reasoning over Temporal Knowledge Graphs

Model ReleasesDGX agent

arXiv:2510.06002v3 Announce Type: replace Abstract: In high-stakes legal domains, retrieval must preserve not only semantic relevance, but also the hierarchy, temporality, and causal provenance of leg

Distill-Belief: Closed-Loop Inverse Source Localization and Characterization in Physical Fields

Local AiDGX agent

arXiv:2604.26095v1 Announce Type: new Abstract: {Closed-loop inverse source localization and characterization (ISLC) requires a mobile agent to select measurements that localize sources and infer late

Domain-Adapted Small Language Models for Reliable Clinical Triage

ResearchDGX agent

arXiv:2604.26766v1 Announce Type: cross Abstract: Accurate and consistent Emergency Severity Index (ESI) assignment remains a persistent challenge in emergency departments, where highly variable free-

DreamProver: Evolving Transferable Lemma Libraries via a Wake-Sleep Theorem-Proving Agent

AgentsDGX agent

arXiv:2604.26311v1 Announce Type: new Abstract: We introduce DreamProver, an agentic framework that leverages a 'wake-sleep' program induction paradigm to discover reusable lemmas for formal theorem p

DSIPA: Detecting LLM-Generated Texts via Sentiment-Invariant Patterns Divergence Analysis

Model ReleasesDGX agent

arXiv:2604.26328v1 Announce Type: cross Abstract: The rapid advancement of large language models (LLMs) presents new security challenges, particularly in detecting machine-generated text used for misi

DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference

Local AiDGX agent

arXiv:2604.26557v1 Announce Type: cross Abstract: The increasing deployment of Large Language Model (LLM) inference on edge AI systems demands efficient execution under tight memory budgets. A key cha

Efficient Traffic Forecasting on Large-Scale Road Network by Regularized Adaptive Graph Convolution

TutorialsDGX agent

arXiv:2506.07179v2 Announce Type: replace-cross Abstract: Traffic prediction is a critical task in spatial-temporal forecasting with broad applications in travel planning and urban management. To mode

ELIQ: A Label-Free Framework for Quality Assessment of Evolving AI-Generated Images

Model ReleasesDGX agent

arXiv:2602.03558v2 Announce Type: replace-cross Abstract: Generative text-to-image models are advancing at an unprecedented pace, continuously shifting the perceptual quality ceiling and rendering pre

Emergent Coordination in Multi-Agent Language Models

Local AiDGX agent

arXiv:2510.05174v4 Announce Type: replace-cross Abstract: When are multi-agent LLM systems merely a collection of individual agents versus an integrated collective with higher-order structure? We intr

Enforcing Benign Trajectories: A Behavioral Firewall for Structured-Workflow AI Agents

Model ReleasesDGX agent

arXiv:2604.26274v1 Announce Type: cross Abstract: Structured-workflow agents driven by large language models execute tool calls against sensitive external environments. We propose odename, a telemetry

Entropy Centroids as Intrinsic Rewards for Test-Time Scaling

Model ReleasesDGX agent

arXiv:2604.26173v1 Announce Type: cross Abstract: An effective way to scale up test-time compute of large language models is to sample multiple responses and then select the best one, as in Grok Heavy

Evaluating Strategic Reasoning in Forecasting Agents

SafetyDGX agent

arXiv:2604.26106v1 Announce Type: new Abstract: Forecasting benchmarks produce accuracy leaderboards but little insight into why some forecasters are more accurate than others. We introduce Bench to t

Evaluating the Alignment Between GeoAI Explanations and Domain Knowledge in Satellite-Based Flood Mapping

SafetyDGX agent

arXiv:2604.26051v1 Announce Type: cross Abstract: The increasing number of satellites has improved the temporal resolution of Earth observation, making satellite-based flood mapping a promising approa

Evaluating the relationship between regularity and learnability in recursive numeral systems using Reinforcement Learning

TutorialsDGX agent

arXiv:2602.21720v2 Announce Type: replace-cross Abstract: Human recursive numeral systems (i.e., counting systems such as English base-10 numerals), like many other grammatical systems, are highly reg

Evergreen: Efficient Claim Verification for Semantic Aggregates

Model ReleasesDGX agent

arXiv:2604.26180v1 Announce Type: cross Abstract: With recent semantic query processing engines, semantic aggregation has become a primitive operator, enabling the reduction of a relation into a natur

EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents

Model ReleasesDGX agent

arXiv:2511.02399v2 Announce Type: replace-cross Abstract: Recent advances in large language model agents offer the promise of automating end-to-end software development from natural language requireme

Explainable Representation of Finite-Memory Policies for POMDPs using Decision Trees

TutorialsDGX agent

arXiv:2411.13365v2 Announce Type: replace Abstract: Partially Observable Markov Decision Processes (POMDPs) are a fundamental framework for decision-making under uncertainty and partial observability.

Exploring the Potential of Probabilistic Transformer for Time Series Modeling: A Report on the ST-PT Framework

ResearchDGX agent

arXiv:2604.26762v1 Announce Type: cross Abstract: The Probabilistic Transformer (PT) establishes that the Transformer's self-attention plus its feed-forward block is mathematically equivalent to Mean-

FedPF: Accurate Target Privacy Preserving Federated Learning Balancing Fairness and Utility

SafetyDGX agent

arXiv:2510.26841v2 Announce Type: replace-cross Abstract: Federated Learning (FL) enables collaborative model training without data sharing, yet participants face a fundamental challenge, e.g., simult

From Black-Box Confidence to Measurable Trust in Clinical AI: A Framework for Evidence, Supervision, and Staged Autonomy

ResearchDGX agent

arXiv:2604.26671v1 Announce Type: cross Abstract: Trust in clinical artificial intelligence (AI) cannot be reduced to model accuracy, fluency of generation, or overall positive user impression. In med

FruitProM-V2: Robust Probabilistic Maturity Estimation and Detection of Fruits and Vegetables

ResearchDGX agent

arXiv:2604.26084v1 Announce Type: cross Abstract: Accurate fruit maturity identification is essential for determining harvest timing, as incorrect assessment directly affects yield and post-harvest qu

Fundamental Physics, Existential Risks and Human Futures

SafetyDGX agent

arXiv:2604.26530v1 Announce Type: cross Abstract: Over the past 25 years, I have been involved in some intriguing developments in the foundations of physics, exploring the quantum reality problem, the

FutureWorld: A Live Environment for Training Predictive Agents with Real-World Outcome Rewards

Model ReleasesDGX agent

arXiv:2604.26733v1 Announce Type: new Abstract: Live future prediction refers to the task of making predictions about real-world events before they unfold. This task is increasingly studied using larg

Generative AI-Based Virtual Assistant using Retrieval-Augmented Generation: An evaluation study for bachelor projects

ResearchDGX agent

arXiv:2604.25924v1 Announce Type: cross Abstract: Large Language Models have been increasingly employed in the creation of Virtual Assistants due to their ability to generate human-like text and handl

Generative models on phase space

TutorialsDGX agent

arXiv:2604.02415v2 Announce Type: replace-cross Abstract: Deep generative models such as diffusion and flow matching are powerful machine learning tools capable of learning and sampling from high-dime

Glance-or-Gaze: Incentivizing LMMs to Adaptively Focus Search via Reinforcement Learning

SafetyDGX agent

arXiv:2601.13942v2 Announce Type: replace-cross Abstract: Large Multimodal Models (LMMs) have achieved remarkable success in visual understanding, yet they struggle with knowledge-intensive queries in

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning

ApplicationsDGX agent

arXiv:2508.09547v2 Announce Type: replace-cross Abstract: We introduce Goal-Conditioned Visual Navigation Instruction Generation (GoViG), a new task that aims to generate contextually coherent navigat

Graph Construction and Matching for Imperative Programs using Neural and Structural Methods

ResearchDGX agent

arXiv:2604.26578v1 Announce Type: cross Abstract: Reusing verification artefacts requires identifying structural and semantic similarities across programs and their specifications. In this paper, we f

← Previous
1…292293294295296…354
Next →