AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
3 Jun 2026

Forgetting is Not Erasure: Recovering Latent Knowledge via Transport Keys

SafetyDGX agent

arXiv:2606.02860v1 Announce Type: cross Abstract: Catastrophic forgetting is often framed as a representational problem: after sequential training, a model appears to lose the features that supported

Formalizing the Binding Problem

TutorialsDGX agent

arXiv:2606.03976v1 Announce Type: cross Abstract: Representations of the world, arguably, contain information about features (e.g. something is blue, something is a circle) but also information about

From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2606.03660v1 Announce Type: new Abstract: Large language models are increasingly used as chemistry assistants, yet most chemistry benchmarks still score only final answers. This masks a critical


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

From Control Boundary to Insurance Claim: Reconstructing AI-Mediated Losses Through the CER Framework

AgentsDGX agent

arXiv:2606.03777v1 Announce Type: new Abstract: AI losses that arise through an insured organization's generative or agentic AI system require state reconstruction, not merely event reconstruction, be

From Long News to Accurate Forecast: Importance-Aware Fusion and PRM-Guided Reflection for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2606.03097v1 Announce Type: new Abstract: Incorporating news into time series forecasting is appealing because news can reveal abrupt exogenous events that historical values alone cannot recover

From Prompt to Service: An SLM-Based Agent Orchestration Gateway for AI-Driven Virtual Worlds

Model ReleasesDGX agent

arXiv:2606.03557v1 Announce Type: new Abstract: As generative AI capabilities expand, AI-driven virtual worlds face a growing architectural challenge. Users interact through in-world interfaces in mul

From 'What' to 'How' and 'Why': Sharing LLM-Generated Retrospective Summaries of Older Adults' Passive Tracking Data with Remote Family Members

AgentsDGX agent

arXiv:2606.03876v1 Announce Type: cross Abstract: With the growing prevalence of modern ubiquitous computing technologies, multi-modal tracking systems hold promise for providing timely awareness and

FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations

ResearchDGX agent

arXiv:2606.02615v1 Announce Type: cross Abstract: Few-shot prompting provides an effective way to adapt auditory large language models to low-resource tasks such as children's speech recognition. Howe

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models

Model ReleasesDGX agent

arXiv:2606.03165v1 Announce Type: cross Abstract: The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment). Research, mostly on Scientific Englis

FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration

AgentsDGX agent

arXiv:2512.11213v2 Announce Type: replace Abstract: Scaling test-time computation has been shown to significantly improve large language model (LLM) performance without additional training. However, e

Gender-Dependent Diagnostic Substitution in LLM Medical Triage: Same Symptoms, Unequal Urgency

Model ReleasesDGX agent

arXiv:2606.03641v1 Announce Type: new Abstract: We investigate whether large language models produce different medical triage recommendations for identical neurological symptoms when only the patient'

Generalizing Graph Foundation Models via Hyperbolic Retrieval-Augmented Generation

Local AiDGX agent

arXiv:2606.03307v1 Announce Type: cross Abstract: Graph foundation models (GFMs) emerged as a dominant paradigm in graph representation learning by leveraging large-scale pre-training for cross-domain

Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

Model ReleasesDGX agent

arXiv:2510.21011v3 Announce Type: replace-cross Abstract: As generative AI tools are increasingly used to portray people in professional roles, understanding their racial and gender representational b

Geometry-Aware Tabular Diffusion

Model ReleasesDGX agent

arXiv:2606.02607v1 Announce Type: cross Abstract: Tabular synthesis is critical for privacy-preserving sharing and augmentation, yet diffusion models rely on implicit mechanisms to capture inter-colum

GFFMERGE: Efficient Merging of Graph Neural Force Fields and Beyond

SafetyDGX agent

arXiv:2606.03232v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) have revolutionized Neural Force Fields for atomistic simulations, achieving near-quantum accuracy at reduced cost, yet a

Glass Box at Orbit: A Constitutional AI Verification Framework for Trustworthy Autonomous CubeSat Intelligence

SafetyDGX agent

arXiv:2606.02967v1 Announce Type: cross Abstract: The space industry is quietly building toward something nobody has fully reckoned with: orbital data centers running thousands of autonomous AI worklo

Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation

SafetyDGX agent

arXiv:2606.03385v1 Announce Type: cross Abstract: In robotic manipulation, the tight coupling between grasping and motion planning often obscures the true source of failure, leading to inefficient tri

Greed is Good: A Unifying Perspective on Guided Generation

ResearchDGX agent

arXiv:2502.08006v3 Announce Type: replace-cross Abstract: Training-free guided generation is a widely used and powerful technique that allows the end user to exert further control over the generative

GRZO: Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning

HardwareDGX agent

arXiv:2606.02857v1 Announce Type: cross Abstract: Zeroth-order (ZO) optimization is a memory-efficient alternative to backpropagation for fine-tuning large language models, but its deployment is limit

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory

Model ReleasesDGX agent

arXiv:2606.03144v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning as

GuidedBridge: Training-freely Improving Bridge Models with Prior Guidance

ResearchDGX agent

arXiv:2606.03119v1 Announce Type: cross Abstract: Guidance methods, such as classifier-free guidance (CFG) and auto-guidance (AG), have advanced noise-to-data generation in diffusion models. Recently,

Hallucinations as Orthogonal Noise: Inference-Time Manifold Alignment via Dynamic Contextual Orthogonalization

Model ReleasesDGX agent

arXiv:2606.03022v1 Announce Type: cross Abstract: Hallucination in Large Language Models (LLMs), characterized by the generation of content inconsistent with contextual facts or logical constraints --

Hand Trajectory Fusion for Egocentric Natural Language Query Grounding

ResearchDGX agent

arXiv:2606.02962v1 Announce Type: cross Abstract: Egocentric Natural Language Query (NLQ) grounding asks a model to localize, in a long first-person video, the temporal interval that answers a free-fo

Handoff Debt: The Rediscovery Cost When Coding Agents Take Over Interrupted Tasks

AgentsDGX agent

arXiv:2606.02875v1 Announce Type: new Abstract: Coding-agent benchmarks evaluate whether a single uninterrupted agent can resolve a repository issue. Real software work is messier: tasks are interrupt

Hedge-Bench: Benchmarking Agents on Hard, Realistic Tasks Pertaining to Financial Reasoning

Model ReleasesDGX agent

arXiv:2606.03918v1 Announce Type: new Abstract: AI agents can increasingly handle the mechanical tasks of financial analysis: retrieving documents, calculating formulas, updating spreadsheets. The har

High-Precision APT Malware Attribution with Out-of-Scope Resilience

ResearchDGX agent

arXiv:2606.03523v1 Announce Type: cross Abstract: Early attribution of Advanced Persistent Threat (APT) activity can help defenders prioritise investigation, select countermeasures, and reduce the imp

How Quantization Changes Interpretable Features: A Sparse Autoencoder Analysis of Language Models

Model ReleasesDGX agent

arXiv:2606.03002v1 Announce Type: cross Abstract: Quantization is a standard path to deploying large language models, and a quantized model is typically judged acceptable when its perplexity or downst

Human-Like Goalkeeping in a Realistic Football Simulation: a Sample-Efficient Reinforcement Learning Approach

AgentsDGX agent

arXiv:2510.23216v4 Announce Type: replace Abstract: While several high profile video games have served as testbeds for Deep Reinforcement Learning (DRL), this technique has rarely been employed by the

Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking

ResearchDGX agent

arXiv:2606.03985v1 Announce Type: cross Abstract: We introduce Humanoid-GPT, a GPT-style Transformer with causal attention trained on a billion-scale motion corpus for whole-body control. Unlike prior

Identifying Quantum Structure in AI Language: Evidence for Evolutionary Convergence of Human and Artificial Cognition

Model ReleasesDGX agent

arXiv:2511.21731v2 Announce Type: replace-cross Abstract: We present the results of cognitive tests on conceptual combinations, performed using specific Large Language Models (LLMs) as test subjects.

IdiomX A Multilingual Benchmark for Idiom Understanding, Retrieval, and Interpretation

Model ReleasesDGX agent

arXiv:2606.02584v1 Announce Type: cross Abstract: Idiomatic expressions remain a persistent challenge for natural language processing because their meanings are often non-compositional, context-depend

Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models

ResearchDGX agent

arXiv:2606.03988v1 Announce Type: new Abstract: Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly observable. Many s

'**Important** You should give me full credits!': Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems

SafetyDGX agent

arXiv:2606.03090v1 Announce Type: cross Abstract: The emergence of large language models (LLMs) has significantly accelerated recent research on LLM-based automatic grading (AG) systems. Benefiting fr

Improvise, Adapt, Overcome: An On-The-Fly Multifidelity Algorithm for Efficient Machine Learning

Model ReleasesDGX agent

arXiv:2606.02662v1 Announce Type: cross Abstract: Machine learning has accelerated quantum chemistry but is hindered by the prohibitive cost of generating high fidelity training data. Multifidelity ma

Inducing Reasoning Primitives from Agent Traces

AgentsDGX agent

arXiv:2606.02994v1 Announce Type: new Abstract: ReAct-style LLM agents often rediscover the same reasoning routines across problems, yet leave those routines trapped in transient scratchpads. We intro

Inference Cost Attacks for Retrieval-Augmented Large Language Models

SafetyDGX agent

arXiv:2606.02643v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG)-enhanced LLM systems, while powerful, introduce substantial inference costs due to the inclusion of an extra mult

InfoMem: Training Long-Context Memory Agents with Answer-Conditioned Information Gain

Local AiDGX agent

arXiv:2606.03329v1 Announce Type: new Abstract: Long-context tasks require LLMs to identify and preserve answer-relevant information from large contexts. Chunk-wise memory agents address this issue by

InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.06960v3 Announce Type: replace-cross Abstract: Large reasoning models achieve strong performance by scaling inference-time chain-of-thought, but this paradigm suffers from quadratic cost, c

Introduction to optimization methods for training SciML models

TutorialsDGX agent

arXiv:2601.10222v2 Announce Type: replace-cross Abstract: Optimization is central to both modern machine learning (ML) and scientific machine learning (SciML), yet the structure of the underlying opti

KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack Problem

ApplicationsDGX agent

arXiv:2602.20217v2 Announce Type: replace-cross Abstract: Self-speculative decoding (SSD) accelerates LLM inference by skipping layers to create an efficient draft model, yet existing methods often re

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories

SafetyDGX agent

arXiv:2606.03979v1 Announce Type: cross Abstract: The past few decades have witnessed significant advances in the design of machine learning algorithms, from early studies on task-specific shallow mod

LAP: An Agent-to-Instrument Protocol for Autonomous Science

SafetyDGX agent

arXiv:2606.03755v1 Announce Type: new Abstract: Autonomous science is moving from demonstration to infrastructure. Large language model agents now plan experiments, and self-driving laboratories execu

Large AI Models in Dental Healthcare: From General-Purpose Systems to Domain-Specific Foundation Models

AgentsDGX agent

arXiv:2606.02914v1 Announce Type: new Abstract: Background: Oral diseases affect nearly 3.5 billion people worldwide, yet the comparative clinical potential of large-scale AI models in dentistry remai

Large Byte Model: Teaching Language Models About Compiled Code

ResearchDGX agent

arXiv:2606.02834v1 Announce Type: cross Abstract: Malware analysis starts with the raw bytes of an executable program, and tools to 'lift' these to higher-level representations, such as assembly, are

LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning

ResearchDGX agent

arXiv:2602.07075v5 Announce Type: replace-cross Abstract: Current chemical large language models (LLMs) predominantly rely on explicit Chain-of-Thought (CoT) to solve complex reasoning problems. Howev

Lean-GAP: A Dataset of Formalized Graduate Algebra Problems

ResearchDGX agent

arXiv:2606.02588v1 Announce Type: cross Abstract: We present Lean-GAP (Lean-Graduate Agebra Problems), 430 formalized graduate-level algebra problems from the textbook Abstract Algebra by Dummit and F

LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

Model ReleasesDGX agent

arXiv:2606.03303v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit strong informal mathematical reasoning but struggle to generate mechanically verifiable proofs in formal languages

Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs

Local AiDGX agent

arXiv:2606.03489v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel in code generation, they remain prone to replicating subtle yet critical vulnerabilities endemic to their tra

Learn When and Where to Connect: Adaptive Virtual Nodes for Dynamic Message Passing on Graphs

TutorialsDGX agent

arXiv:2606.03068v1 Announce Type: cross Abstract: While Virtual Nodes (VNs) are often utilized in Message Passing Neural Networks (MPNNs) to facilitate effective message passing, existing VN-based met

Learned Non-Maximum Suppression for 3D Object Detection

ResearchDGX agent

arXiv:2606.03568v1 Announce Type: cross Abstract: Post-processing is a critical stage in LiDAR-based 3D object detection, where dense and overlapping proposals must be filtered for compact and reliabl

Learning Multi-Scale Hypergraph for High-Order Brain Connectivity Analysis

ResearchDGX agent

arXiv:2606.03310v1 Announce Type: cross Abstract: Understanding complex interactions between brain regions is critical for early neurodegenerative disease classification such as Alzheimer's Disease (A

Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs

SafetyDGX agent

arXiv:2602.10352v2 Announce Type: replace-cross Abstract: Self-interpretation methods prompt language models to describe their own internal states, but remain unreliable due to hyperparameter sensitiv

Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph Pretraining

ResearchDGX agent

arXiv:2509.22468v2 Announce Type: replace-cross Abstract: High-quality molecular representations are essential for property prediction and molecular design, yet large labeled datasets remain scarce. W

Learning to Refine: Spectral-Decoupled Iterative Refinement Framework for Precipitation Nowcasting

ResearchDGX agent

arXiv:2606.02661v1 Announce Type: cross Abstract: Accurate precipitation nowcasting is vital for disaster mitigation, but deep learning methods face a key trade-off: regression models produce over-smo

Leave it to the Specialist: Repair Sparse LLMs with Sparse Fine-Tuning via Sparsity Evolution

Model ReleasesDGX agent

arXiv:2505.24037v3 Announce Type: replace Abstract: Sparse large language models (LLMs) offer an attractive direction toward efficient deployment, but adapting them to downstream tasks remains challen

Leveraging BART to Assess CS1 C++ Programming Assignments using Rubric-based Criteria

SafetyDGX agent

arXiv:2606.03814v1 Announce Type: new Abstract: This paper investigates rubric-aware, multitask fine-tuning of transformer models for automated grading of introductory C++ programming assignments, wit

Libra: Efficient Resource Management for Agentic RL Post-Training

SafetyDGX agent

arXiv:2606.03077v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a standard post-training paradigm for large language models (LLMs), extending beyond preference alignment to co

Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States

TutorialsDGX agent

arXiv:2606.02907v1 Announce Type: cross Abstract: Linear probing of large language model (LLM) hidden states is widely used to claim that models learn distinct representations for different reasoning

LiveBand: Live Accompaniment Generation in the Audio Domain

Model ReleasesDGX agent

arXiv:2606.03803v1 Announce Type: cross Abstract: We present LiveBand, a real-time system that generates high-fidelity music accompaniments to live audio input, respecting strict causal constraints. O

LLM-Assisted Reranking to Operationalize Nuanced Objectives in Recommender Systems

ResearchDGX agent

arXiv:2606.02883v1 Announce Type: cross Abstract: Recommender systems have grown from content-organization tools into sophisticated systems that shape daily behavior. By controlling what we see, they

← Previous
1…165166167168169…358
Next →