AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
9 Jun 2026

Q-Delta: Beyond Key-Value Associative State Evolution

ResearchDGX agent

arXiv:2606.08804v1 Announce Type: new Abstract: Linear attention reformulates sequence modeling as recurrent state evolution, enabling efficient linear-time inference. Under the key-value associative

Quantitative Promise Theory: Intentionality and Inference in Autonomous Agents

SafetyDGX agent

arXiv:2606.08552v1 Announce Type: new Abstract: I discuss some quantitative representations of Promise Theory for processes involving autonomous agents. Agent models are common in software systems, ma

Quantum-Enhanced Similarity Measures for Polarimetric Materials Classification

ResearchDGX agent

arXiv:2606.07766v1 Announce Type: cross Abstract: We present a quantum--classical hybrid pipeline for polarimetric material classification that casts this as a point-matching problem. Voxel cubes, con


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects

ResearchDGX agent

arXiv:2606.07617v1 Announce Type: cross Abstract: While sparse autoencoders provide features more interpretable than individual neurons, reliably characterizing them remains challenging. We propose Qu

RadOT-Eval: Auditable Structured-Evidence Transport for Radiology Report Evaluation

ResearchDGX agent

arXiv:2606.08769v1 Announce Type: cross Abstract: Automatic evaluation is critical for high-stakes text generation, where errors often involve omitted findings, hallucinated content, polarity reversal

RAILS: Verification-Native Clearing For Agentic Commerce

AgentsDGX agent

arXiv:2606.08790v1 Announce Type: new Abstract: Autonomous agents negotiate, purchase, deploy code, and move funds, but no neutral mechanism determines whether they met their delegated obligation, who

RAPID: Layer-Wise Redundancy-Aware Pruning and Importance-Driven Token Merging for Efficient ViT

ResearchDGX agent

arXiv:2606.08156v1 Announce Type: cross Abstract: Vision Transformers (ViTs) achieve strong performance but suffer from high computational costs due to quadratic self-attention complexity. Although to

Reachability and asymptotics of Gaussian Transformer dynamics

ResearchDGX agent

arXiv:2606.07600v1 Announce Type: cross Abstract: We formulate data propagation through the Transformer, the machine learning architecture powering large language models, as a nonlinear control system

Real-time body pose non-verbal communication with a consistency-based reliability measure

Model ReleasesDGX agent

arXiv:2606.09390v1 Announce Type: cross Abstract: Body movement communicates intent at distances and in conditions where neither the face, nor speech can be captured. We study the recognition of commu

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short

ResearchDGX agent

arXiv:2606.09380v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a leading paradigm for improving the reasoning ability of large language models throu

Reconstructing and forecasting disease trajectories of patients with Alzheimer's disease using routine data in resource-constrained settings

ResearchDGX agent

arXiv:2606.07798v1 Announce Type: new Abstract: Alzheimer's disease is a progressive neurodegenerative disorder, and its progression varies substantially across patients. Existing work aims to forecas

Reconstructing Synthetic SDO/AIA 193 A EUV Images from He I 10830 A Observations with Diffusion Model Translator

ResearchDGX agent

arXiv:2606.08652v1 Announce Type: cross Abstract: Routine full-disk EUV imaging has been available only since the modern era, such as SOHO and SDO. To extend EUV coronal context into earlier periods,

ReCoVLA: VLM-Guided Reward Compilation for Failure Recovery in Vision-Language-Action Policies

SafetyDGX agent

arXiv:2606.09630v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies provide strong priors for language-conditioned manipulation, but remain brittle in off-nominal states requiring

RecurGuard: Runtime Monitoring for Reasoning-Token Consumption Attacks

Model ReleasesDGX agent

arXiv:2606.07968v1 Announce Type: cross Abstract: Reasoning-capable large language models can be induced to spend their generation budget on injected decoy tasks rather than answering the user's quest

REFLECT: Intervention-Supported Error Attribution for Silent Failures in LLM Agent Traces

AgentsDGX agent

arXiv:2606.09071v1 Announce Type: new Abstract: Large language model (LLM) agents now solve complex tasks through long plan-and-execution traces, yet the ability to locate errors in a completed traces

Reflection in the Dark: Exposing and Escaping the Black Box in Reflective Prompt Optimization

AgentsDGX agent

arXiv:2603.18388v2 Announce Type: replace Abstract: Automatic prompt optimization (APO) has emerged as a powerful paradigm for improving LLM performance without manual prompt engineering. Reflective A

Reinforcement Learning for Flow-Matching Policies with Density Transport

SafetyDGX agent

arXiv:2606.08602v1 Announce Type: cross Abstract: We present an online reinforcement learning (RL) algorithm for fine-tuning flow-matching policies in continuous-control problems. Our key insight is t

Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges

SafetyDGX agent

arXiv:2606.09165v1 Announce Type: new Abstract: Safety judges are increasingly deployed to evaluate model outputs against evolving criteria, yet recent meta-evaluation work shows they remain brittle u

Repair Before Veto, When Repair Is Hidden: Quantum-Accessible Features for Repair-Augmented Constraint Learning

SafetyDGX agent

arXiv:2606.08020v1 Announce Type: cross Abstract: Hard-constraint decision systems usually veto infeasible candidates. This is too rigid when the system can act: if a known affordable repair would mak

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them

Model ReleasesDGX agent

arXiv:2606.07597v1 Announce Type: cross Abstract: Pre-training data mixtures are commonly tuned by running small-scale experiments and extrapolating to the target training budget. When high-quality da

Report on CHIIR 2026 Workshop on Generative AI and Academic Search (GAI&AS)

ResearchDGX agent

arXiv:2606.08936v1 Announce Type: cross Abstract: This report summarizes the CHIIR 2026 Workshop on Generative AI and Academic Search (GAI&AS), which examined how GenAI is reshaping academic search sy

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

Model ReleasesDGX agent

arXiv:2606.07591v1 Announce Type: cross Abstract: AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We presen

Resource-aware Computation-Communication Overlap for multi-GPU ML Workloads

HardwareDGX agent

arXiv:2606.09200v1 Announce Type: cross Abstract: The rapid growth of large-scale machine learning (ML) has made distributed training across multiple GPUs a fundamental component of modern ML systems.

Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering

ApplicationsDGX agent

arXiv:2606.07523v1 Announce Type: cross Abstract: Legal domains in high-resource languages like English have widely adopted artificial intelligence for legal question answering. However, data scarcity

RetroReasoner: A Reasoning LLM for Strategic Retrosynthesis Prediction

ResearchDGX agent

arXiv:2603.12666v2 Announce Type: replace-cross Abstract: Retrosynthesis prediction aims to identify reactants that can synthesize a given product molecule. Although molecular large language models (L

Revisiting the shutdown problem

SafetyDGX agent

arXiv:2606.08296v1 Announce Type: new Abstract: A key premise in leading arguments for existential risk from artificial intelligence is that malfunctioning artificial agents could not be easily shut d

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency

Model ReleasesDGX agent

arXiv:2601.06649v2 Announce Type: replace-cross Abstract: Research in machine learning has questioned whether increases in training token counts reliably produce proportional performance gains in larg

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective

SafetyDGX agent

arXiv:2602.02572v2 Announce Type: replace-cross Abstract: Existing alignment methods directly use the reward model learned from user preference data to optimize an LLM policy, subject to KL regulariza

Rewrite to Translate, Translate to Reward: Reinforcement Learning for Source Rewriting in Machine Translation

ResearchDGX agent

arXiv:2606.08011v1 Announce Type: cross Abstract: Although directly prompting off-the-shelf Large Language Models (LLMs) to generate meaning-preserving source rewrites can effectively enhance Machine

RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations

Model ReleasesDGX agent

arXiv:2606.08376v1 Announce Type: cross Abstract: As artificial intelligence (AI) systems are increasingly deployed across socially consequential domains, reports of AI-related harms and failures have

Robust Renal Mass Segmentation on CT: A Validation Study of an AI-Based Framework

ResearchDGX agent

arXiv:2505.07573v2 Announce Type: replace-cross Abstract: Renal mass segmentation has important potential to enhance the clinical workflow, especially in settings requiring quantitative assessments. K

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

Model ReleasesDGX agent

arXiv:2606.08063v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in visual understanding, yet their performance degrades significantly un

Rosetta Memory: Adaptive Memory for Cross-LLM Agents

Model ReleasesDGX agent

arXiv:2606.07711v1 Announce Type: cross Abstract: Memory is the key component for transforming a stateless LLM into a persistent, evolving agent through experience accumulation, long-horizon planning,

RTL-BenchLS: A Large-Scale Benchmark for RTL Reasoning and Generation with Large Language Models

Model ReleasesDGX agent

arXiv:2606.08976v1 Announce Type: new Abstract: LLM-based RTL generation and reasoning is a promising direction for hardware design automation. High-quality benchmarks are critical infrastructure for

Rule-based autocorrection of Piping and Instrumentation Diagrams (P&IDs) on graphs

ApplicationsDGX agent

arXiv:2502.18493v2 Announce Type: replace-cross Abstract: A piping and instrumentation diagram (P&ID) is a central reference document in chemical process engineering. Currently, chemical engineers man

RunAgent SuperBrowser: A Theory of Autonomous Web Navigation Grounded in Human Browsing Behaviour

Model ReleasesDGX agent

arXiv:2606.09399v1 Announce Type: new Abstract: We present SUPERBROWSER, an autonomous web-navigation agent designed against a single guiding hypothesis: a web agent should browse the way a person bro

Safe-RULE: Safe Reinforcement UnLEarning

Model ReleasesDGX agent

arXiv:2606.09559v1 Announce Type: cross Abstract: Offline safe reinforcement learning (Safe RL) enables policy learning without online interactions, making it suitable for safety-critical systems such

SafeECGMatch: Calibration-Aware Joint Frequency and Time Space Semi-Supervised Learning for Open-Set ECG Classification

ResearchDGX agent

arXiv:2606.08037v1 Announce Type: cross Abstract: Electrocardiogram (ECG) classification models often suffer from severe label scarcity, making semi-supervised learning (SSL) an attractive strategy fo

SafeRun: Enabling Determinism in LLM Planning for Running

Model ReleasesDGX agent

arXiv:2606.09027v1 Announce Type: cross Abstract: Large Language Models enable flexible natural-language planning but remain unreliable in determinism-critical domains due to their probabilistic natur

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators

SafetyDGX agent

arXiv:2606.07874v1 Announce Type: new Abstract: LLMs-as-judges are the only way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement

SAGE: An LLM-driven Self Reflective Agentic Framework for Fraud Detection

AgentsDGX agent

arXiv:2606.08146v1 Announce Type: new Abstract: Fraud detection in payment, e-commerce, and telecommunications systems requires accuracy at the individual level, robustness under severe class imbalanc

SAGE: Shape-Adapting Gated Experts for Adaptive Histopathology Image Segmentation

ResearchDGX agent

arXiv:2511.18493v4 Announce Type: replace-cross Abstract: The significant variability in cell size and shape continues to pose a major obstacle in computer-assisted cancer detection on gigapixel Whole

SAILS: Surrogate-based Analysis of Interactions via Local Effect Smooths

Local AiDGX agent

arXiv:2606.09404v1 Announce Type: cross Abstract: Feature interactions drive much of the predictive power of machine learning models, yet existing explanation methods only detect and quantify interact

Sample-Efficient LLM-Based Detection of Malicious Web Server Logs with Forensically Explainable Reasoning

TutorialsDGX agent

arXiv:2606.08649v1 Announce Type: cross Abstract: Forensic analysis of web server logs demands both accurate detection and human-readable explanations that can satisfy legal requirements. We present C

Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning

SafetyDGX agent

arXiv:2606.07602v1 Announce Type: cross Abstract: LLM-based LEGO assembly generation requires both semantic grounding and physical feasibility. We identify a data-induced failure mode, PhysHack, in wh

SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models

SafetyDGX agent

arXiv:2606.07705v1 Announce Type: cross Abstract: Although multi-objective reinforcement learning (MORL) is central to aligning large language models with complex human preferences, the prevailing pra

Scaffold Effects on GAIA: A Controlled Comparison

Model ReleasesDGX agent

arXiv:2606.08529v1 Announce Type: new Abstract: Published agent capability scores conflate what a model can do with what its scaffold lets it do, and the magnitude of this elicitation gap is not well

ScaleSweep: Accurate NVFP4 Post-Training Quantization of LLMs via Block Scale Initialization

Model ReleasesDGX agent

arXiv:2606.07618v1 Announce Type: cross Abstract: NVFP4 is a recently introduced hardware-supported FP4 format that improves the fidelity of 4-bit quantization through fine-grained block scales. Howev

Scaling Decision-Focused Learning to Large Problems with Lagrangian Decomposition

ResearchDGX agent

arXiv:2606.08797v1 Announce Type: cross Abstract: Decision-focused learning has shown great promise for addressing predict-then-optimize problems, particularly in the presence of under-specified model

Scaling Neural Network Verification with Tensor Parallelism and Fully Sharded Data Parallelism

SafetyDGX agent

arXiv:2606.09377v1 Announce Type: cross Abstract: Formal neural network verification -- proving that a network satisfies safety properties for all inputs in a specified domain -- is bounded in practic

Scaling Participation in Modular AI Systems

ResearchDGX agent

arXiv:2606.07812v1 Announce Type: new Abstract: Humanity is a mosaic of multifaceted talents and needs, and any truly intelligent AI must reflect that richness. Yet the LLMs used by all are built by t

SceneConductor: 3D Scene Generation from Single Image with Multi-Agent Orchestration

Model ReleasesDGX agent

arXiv:2606.08402v1 Announce Type: cross Abstract: Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context fro

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems

Model ReleasesDGX agent

arXiv:2606.08034v1 Announce Type: cross Abstract: Symbolic benchmarks have emerged as a key approach to assess model robustness under minor modifications to STEM-related questions. However, existing s

SciTrace: Trajectory-Aware Safety Reasoning for Scientific Discovery Agents

SafetyDGX agent

arXiv:2606.08234v1 Announce Type: new Abstract: LLM-based scientific agents have shown strong capacity for autonomous research, yet their safety layers remain structurally divorced from core reasoning

SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research

AgentsDGX agent

arXiv:2606.09730v1 Announce Type: new Abstract: Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model

SecureClaw: Clawing Back Control of LLM Agents

SafetyDGX agent

arXiv:2606.09549v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents face two distinct security failures: unauthorized external actions and exposure of sensitive plaintext in

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding

Model ReleasesDGX agent

arXiv:2606.09064v1 Announce Type: cross Abstract: Recent advances in Video Large Language Models (Video-LLMs) have enabled performance on long-video understanding tasks. However, existing methods stil

Seeing is Believing: Aligning Prompt Rewriting with Visual Anchors for Text-to-Image Generation

ResearchDGX agent

arXiv:2606.08492v1 Announce Type: cross Abstract: Despite the impressive capabilities of text-to-image (T2I) models, an intent-generation gap often persists due to the brevity and ambiguity of user pr

Seeing the Hivemind: A Consensus-Aware Interaction Technique for Mitigating AI Homogenization

Local AiDGX agent

arXiv:2606.09587v1 Announce Type: cross Abstract: People are increasingly using AI for creative tasks such as writing. While adoption continues to grow, this form of use risks undermining individual c

SEF-CLGC at SemEval-2026 Task 11: Logical Notation Impact on Language Model Performance

SafetyDGX agent

arXiv:2606.09157v1 Announce Type: cross Abstract: This paper revisits our pipeline called Syllogistic Evaluation Framework-Common Logic Grammar Construction (SEF-CLGC). We combine formal logical notat

← Previous
1…147148149150151…358
Next →