AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
10 Jun 2026

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

ResearchDGX agent

arXiv:2602.12424v2 Announce Type: replace-cross Abstract: Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitatin

RAT: Reference-Augmented Training for ASV Anti-Spoofing

Model ReleasesDGX agent

arXiv:2606.10908v1 Announce Type: cross Abstract: We introduce a spoofing countermeasure architecture conditioned on speaker-reference recordings, but observe that it converges to a solution that effe

READER: Robust Evidence-based Authorship Decoding via Extracted Representations

Model ReleasesDGX agent

arXiv:2606.10794v1 Announce Type: new Abstract: As agentic applications increasingly route user tasks through official and third-party LLM APIs, provenance becomes an operational question: which model


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

Model ReleasesDGX agent

arXiv:2606.10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in solving high-school mathematics, their ability to evaluate the diverse reas

ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models

Model ReleasesDGX agent

arXiv:2606.11164v1 Announce Type: new Abstract: Long chain-of-thought (CoT) trajectories in large language model (LLM) reasoning cause severe inference bottlenecks due to rapid key-value (KV) cache gr

Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning

SafetyDGX agent

arXiv:2606.10346v1 Announce Type: new Abstract: Reinforcement learning has become a key paradigm for eliciting reasoning abilities in large language models, where exploration is crucial for discoverin

Reasoning over Semantic IDs Enhances Generative Recommendation

SafetyDGX agent

arXiv:2603.23183v2 Announce Type: replace-cross Abstract: Recent advances in generative recommendation have leveraged pretrained LLMs by formulating sequential recommendation as autoregressive generat

Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models

Model ReleasesDGX agent

arXiv:2606.10949v1 Announce Type: new Abstract: Persistent memory systems promise to make LLMs more helpful by storing user beliefs over time. We show they also make models less correct by systematica

Recoverable but Not Stationary:Local Linear Structures in Weights and Activations

Model ReleasesDGX agent

arXiv:2606.10929v1 Announce Type: cross Abstract: Task vectors, LoRA, activation steering, and random search around pretrained weights all suggest that learned behaviour can be controlled by linear di

ReflectiChain: Epistemic Grounding in LLM-Driven World Models for Supply Chain Resilience

Model ReleasesDGX agent

arXiv:2606.10359v1 Announce Type: new Abstract: AI agents in supply chains face a fundamental epistemic gap: large language models (LLMs) interpret policies but lack physical grounding, while reinforc

Regimes: An Auditable, Held-Out-Gated Improvement Loop Demonstrated on LongMemEval with ActiveGraph

AgentsDGX agent

arXiv:2606.10241v1 Announce Type: new Abstract: Autonomous improvement loops are hard to trust because the improvement process is usually external scaffolding bolted onto the agent: failures go unlogg

Representation Curriculum: Stagewise Training for Robust Ranking and Allocation

SafetyDGX agent

arXiv:2606.09891v1 Announce Type: cross Abstract: Ranking in digital marketplaces is a dynamic exposure-allocation mechanism: displayed items shape discovery trajectories and success events logged by

RKSC: Reasoning-Aware KV Cache Sharing and Confident Early Exit for Multi-Step LLM Inference

ResearchDGX agent

arXiv:2606.09937v1 Announce Type: cross Abstract: We introduce RKSC (Reasoning-Aware KV Cache Sharing), a training-free inference framework that eliminates two structural redundancies in multi-branch

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2510.14828v3 Announce Type: replace Abstract: Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks

RoboNaldo: Accurate, Stable and Powerful Humanoid Soccer Shooting via Motion-Guided Curriculum Reinforcement Learning

SafetyDGX agent

arXiv:2606.11092v1 Announce Type: cross Abstract: Elite humanoid soccer shooting requires whole-body stability, high-impulse whole-body interactions, and accuracy to targets. Motion tracking-driven re

Robust Deep Reinforcement Learning Through Adversarial Attacks and Training : A Survey

AgentsDGX agent

arXiv:2403.00420v3 Announce Type: replace-cross Abstract: Deep Reinforcement Learning (DRL) is a subfield of machine learning for training autonomous agents that take sequential actions across complex

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

SafetyDGX agent

arXiv:2606.10917v1 Announce Type: new Abstract: Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interac

Rotate2Think: Geometric Priming via Orthogonal Rotation to Improve Language Model Reasoning

Model ReleasesDGX agent

arXiv:2606.09873v1 Announce Type: cross Abstract: Reasoning models achieve strong performance on challenging tasks by generating explicit intermediate reasoning traces before producing a final answer.

Routing-Aware Expert Calibration for Machine Unlearning in Mixture-of-Experts Language Models

ResearchDGX agent

arXiv:2606.10338v1 Announce Type: cross Abstract: Machine unlearning is increasingly important for large language models, yet unlearning in Mixture-of-Experts (MoE) architectures remains underexplored

SAFE: An LLM-as-Verifier Framework for Evidence-Grounded Multi-Hop Reasoning

Model ReleasesDGX agent

arXiv:2604.01993v2 Announce Type: replace-cross Abstract: Multi-hop QA benchmarks often reward Large Language Models (LLMs) for spurious correctness, where models reach correct answers through invalid

Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling

Model ReleasesDGX agent

arXiv:2606.09926v1 Announce Type: cross Abstract: Sampling from the sequence-level power distribution p^alpha elicits RL-level reasoning from base language models without any parameter updates, but th

SCOPE: Sequential Causal Optimization of Process Interventions

Model ReleasesDGX agent

arXiv:2512.17629v4 Announce Type: replace-cross Abstract: Prescriptive Process Monitoring (PresPM) recommends interventions during running business processes to optimize key performance indicators (KP

SD-GRPO: Verifiable Segment Decomposition for Long-Form Vision-Language Generation

SafetyDGX agent

arXiv:2606.09871v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) and its variants, originally developed for Large Language Models (LLMs), have recently been applied to Multi

Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts

SafetyDGX agent

arXiv:2606.10334v1 Announce Type: new Abstract: Code-generating large language models (LLMs) increasingly produce visual artifacts such as charts, web pages, and slides by writing programs that are ex

Self-EmoQ: Plutchik-Guided Value-based Planning to Drive Streaming Emotional TTS

SafetyDGX agent

arXiv:2606.09837v1 Announce Type: cross Abstract: Emotional interaction is increasingly crucial for conversational AI, yet current systems lack a self-emotion determination mechanism to drive the stre

SHAPE: Coalition-Aware Expert Pruning for Sparse Mixture-of-Experts LLMs

Model ReleasesDGX agent

arXiv:2606.09886v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (MoE) large language models achieve strong quality with low per-token compute, yet their deployment is often limited by the

SHAPO: Sharpness-Aware Policy Optimization for Safe Exploration

Model ReleasesDGX agent

arXiv:2606.10228v1 Announce Type: cross Abstract: Safe exploration is a prerequisite for deploying reinforcement learning (RL) agents in safety-critical domains. In this paper, we approach safe explor

Sigma-Branch: Hierarchical Single-Path Network Reconstruction for Dynamic Inference with Reduced Active Parameters

Model ReleasesDGX agent

arXiv:2606.09924v1 Announce Type: cross Abstract: Deploying deep neural networks on memory-constrained edge accelerators is bottlenecked by per-inference off-chip weight transfer rather than computati

Sim2Schedule: A Simulator-Guided LLM Framework for Autonomous Open-Pit Mine Scheduling

Model ReleasesDGX agent

arXiv:2606.10286v1 Announce Type: new Abstract: Open-pit mine scheduling is a critical process for maximizing economic return under complex geotechnical and operational constraints. While Mixed-Intege

SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval

Model ReleasesDGX agent

arXiv:2606.10388v1 Announce Type: cross Abstract: Agent skill libraries are becoming routable software assets: a retrieved skill can contribute instructions, scripts, resource bindings, and execution

SocraticPO: Policy Optimization via Interactive Guidance

SafetyDGX agent

arXiv:2606.09887v1 Announce Type: cross Abstract: Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewar

Soul Computing: A Theoretical Framework and Technical Architecture for Intelligent Agents with Independent Consciousness

ResearchDGX agent

arXiv:2606.10413v1 Announce Type: new Abstract: Breakthroughs in large language models and multimodal generation technologies have propelled the digital reconstruction of human mental traits, emotiona

SPACE: Source-free Proxy Anchor Concept Erasure for MLLMs

Model ReleasesDGX agent

arXiv:2606.09868v1 Announce Type: cross Abstract: As Multimodal Large Language Models (MLLMs) face growing privacy risks and regulatory constraints, machine unlearning (MU) has emerged as a crucial so

Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

Local AiDGX agent

arXiv:2606.10738v1 Announce Type: cross Abstract: Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for s

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation

ResearchDGX agent

arXiv:2606.10368v1 Announce Type: cross Abstract: Speech-to-text (S2T) systems for recognition (ASR) and translation (S2TT) typically generate discrete text tokens. In contrast, continuous-target lang

STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios

Model ReleasesDGX agent

arXiv:2606.10394v1 Announce Type: new Abstract: Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existin

Stop Early, Spend Less: Hidden-State Probes as a Practical Recipe for Streaming Moderation of LLM Outputs

SafetyDGX agent

arXiv:2606.10487v1 Announce Type: cross Abstract: Deploying large language models in user-facing systems requires efficient output safety filtering. Existing approaches typically rely on a separate mo

STORM: Stepwise Token Optimization with Reward-Guided Beam Search

ResearchDGX agent

arXiv:2606.10621v1 Announce Type: cross Abstract: Modern retrieval increasingly relies on dense and learned-sparse neural models that are effective but require encoding the entire corpus into a specia

Structure from Reasoning, Numbers from Search: On-Premise Open LLMs as Structural Priors for Coupled MIMO Controller Tuning

Model ReleasesDGX agent

arXiv:2606.11015v1 Announce Type: new Abstract: Tuning controllers for strongly coupled multi-input multi-output (MIMO) industrial processes is hard: decentralized classical auto-tuning ignores loop i

Structure-Preserving Learning Improves Geometry Generalization in Neural PDEs

SafetyDGX agent

arXiv:2602.02788v2 Announce Type: replace-cross Abstract: We aim to develop physics foundation models for science and engineering that provide real-time solutions to Partial Differential Equations (PD

Superficial Beliefs in LLM Decision-Making

ResearchDGX agent

arXiv:2606.11016v1 Announce Type: new Abstract: We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic u

Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction

ApplicationsDGX agent

arXiv:2606.10279v1 Announce Type: new Abstract: Supervised fine-tuning with synthetic rationale data is widely assumed to improve language model performance on clinical prediction tasks by teaching mo

Support sufficiency as action-sufficient compression: a single-cycle rate-regret formulation

SafetyDGX agent

arXiv:2606.09858v1 Announce Type: cross Abstract: Robust decision-making requires compression. A system that forms a rich support state cannot usually preserve its full structure at the point of actio

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

Model ReleasesDGX agent

arXiv:2606.11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However,

TaCarla: A comprehensive benchmarking dataset for end-to-end autonomous driving

AgentsDGX agent

arXiv:2602.23499v4 Announce Type: replace-cross Abstract: Collecting a high-quality dataset is a critical task that demands meticulous attention to detail, as overlooking certain aspects can render th

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition

SafetyDGX agent

arXiv:2606.09883v1 Announce Type: cross Abstract: Large language models (LLMs) have made remarkable progress in reasoning tasks, largely driven by post-training paradigms, especially reinforcement lea

Temporal Context Conditioning for Seasonality-Aware Precipitation Nowcasting of High-Intensity Rainfall

Model ReleasesDGX agent

arXiv:2606.09959v1 Announce Type: cross Abstract: Precipitation nowcasting is increasingly being approached with deep learning models that learn directly from recent radar observations. Although such

Temporal Sheaf Neural Networks with Dynamic Orthogonal Transport

Model ReleasesDGX agent

arXiv:2606.10071v1 Announce Type: cross Abstract: We introduce Temporal Sheaf Neural Networks (TSNN), a temporal link prediction framework that equips each node with a time-varying orthogonal frame an

Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies

SafetyDGX agent

arXiv:2606.10371v1 Announce Type: cross Abstract: Diffusion-based action generation has become a foundational component of embodied AI, but its reliance on visual conditioning leaves deployed visuomot

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning

SafetyDGX agent

arXiv:2606.11087v1 Announce Type: cross Abstract: Expressive continuous control policies, such as diffusion and flow models, form the backbone of recent advances in scaling imitation learning for simu

The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment

AgentsDGX agent

arXiv:2606.10747v1 Announce Type: new Abstract: As AI systems built from multiple language-model agents become more common, they are increasingly used to make decisions together: discussing, negotiati

The Bioelectrical Information Theory: Investigating the theoretical compression limit of bioelectrical signals under artificial intelligence

ResearchDGX agent

arXiv:2606.09922v1 Announce Type: cross Abstract: Bioelectrical signals are increasingly acquired at scales that challenge the bandwidth of brain-computer interfaces. However, their compression is sti

The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge

AgentsDGX agent

arXiv:2606.10296v1 Announce Type: cross Abstract: Multi-agent debate systems are typically evaluated only on whether the final answer is correct, overlooking the quality of the intermediate reasoning

The Distributed Detectability Band Against Marginal-Preserving Attacks

AgentsDGX agent

arXiv:2606.10456v1 Announce Type: cross Abstract: AI-control monitors score individual agent actions to detect misbehavior, but real harm can be distributed across many benign-looking steps, each indi

The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans

Model ReleasesDGX agent

arXiv:2606.09844v1 Announce Type: cross Abstract: Large Language Models (LLMs) alter their privacy behavior based on the perceived identity of their interlocutor. While safety mechanisms typically pre

The Role of Feedback Alignment in Self-Distillation

SafetyDGX agent

arXiv:2606.11173v1 Announce Type: new Abstract: Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains t

The Whale That Outswam Evolution: Swarm Intelligence Maximises Memory in Connectome Reservoirs

SafetyDGX agent

arXiv:2606.09902v1 Announce Type: cross Abstract: Reservoir computing exploits the fixed dynamics of a recurrent network for temporal processing, requiring only a trained linear readout. Biological ne

Time Series as Language: A Universal Tokenizer for General-Purpose Time Series Foundation Models

ResearchDGX agent

arXiv:2606.09861v1 Announce Type: cross Abstract: While Next-Token Prediction (NTP) has unified LLM pretraining, its adaptation to unbounded, continuous time series (TS) remains open. To bridge the ga

torch-sla: Differentiable Sparse Linear Algebra with Adjoint Solvers and Sparse Tensor Parallelism for PyTorch

HardwareDGX agent

arXiv:2601.13994v3 Announce Type: replace-cross Abstract: Differentiable sparse linear algebra is foundational for scientific machine learning, yet PyTorch lacks a unified library for it: torch.sparse

Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation

AgentsDGX agent

arXiv:2606.10749v1 Announce Type: cross Abstract: Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintain memory, a

← Previous
1…138139140141142…358
Next →