AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
20,981 results
11 Aug 2026

PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

Model ReleasesDGX agent

arXiv:2608.07631v1 Announce Type: cross Abstract: LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue

Parameter Exploration for RLVR via Variational Learning

Model ReleasesDGX agent

arXiv:2608.09805v1 Announce Type: cross Abstract: Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an importan

PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation

SafetyDGX agent

arXiv:2608.08726v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts. Yet each rollou


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

PATH: Next-Interval Prediction via Autoregressive Tree Hierarchy on Tabular Data

ResearchDGX agent

arXiv:2608.08078v1 Announce Type: new Abstract: Interval prediction aims to achieve a target coverage level while producing intervals that are as short as possible. Many conformal regression pipelines

Performance of large language models in the optical diagnosis of colorectal polyps

Model ReleasesDGX agent

arXiv:2608.07543v1 Announce Type: cross Abstract: Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language

Persistent Semantic Entities in Tool-Augmented LLM Systems

Model ReleasesDGX agent

arXiv:2608.07952v1 Announce Type: cross Abstract: Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries---

Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models

SafetyDGX agent

arXiv:2608.08199v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in group decision-making with other LLMs and humans. Yet it remains unclear whether their influen

PIVOT: Preference-based Intervention Vectors for Pedagogical Tutor Steering

SafetyDGX agent

arXiv:2608.07509v1 Announce Type: cross Abstract: LLMs are increasingly used for conversational tutoring, but effective tutoring requires more than correct answers. Tutors must choose when to scaffold

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling

Model ReleasesDGX agent

arXiv:2608.08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents. Current benchmarks face three struct

PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs

Model ReleasesDGX agent

arXiv:2608.09028v1 Announce Type: new Abstract: Institutional policies stay in natural language while the systems that check compliance demand machine-readable constraints. Bridging that gap is still

PolypSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering

ResearchDGX agent

arXiv:2603.07066v2 Announce Type: replace-cross Abstract: Generative diffusion models are increasingly used for medical imaging data augmentation, but text prompting cannot produce causal training dat

Population-Scalable Multi-Agent World Modeling

AgentsDGX agent

arXiv:2608.08600v1 Announce Type: cross Abstract: World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent environment

Position: Certifiable State Integrity Should Be Built from Local Validity, Not Global Scale

Local AiDGX agent

arXiv:2601.21249v2 Announce Type: replace Abstract: Breakthroughs in language and vision have motivated increasingly general foundation models for time series and physical dynamics, where evidence is

Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use

ResearchDGX agent

arXiv:2608.07475v1 Announce Type: cross Abstract: Generative Artificial Intelligence (GenAI) presents a governance challenge for STEM assessment. Unrestricted access can enable task outsourcing that u

Predictive safety filter enhanced curriculum learning control for efficient vehicle dynamics controller

Model ReleasesDGX agent

arXiv:2608.09653v1 Announce Type: cross Abstract: Recent advances in learning-based control have enabled impressive achievements in solving complex control problems in various domains. However, since

PRISM: A Predictive Protocol for Permutation Optimization via Landscape Diagnostics

ResearchDGX agent

arXiv:2608.08344v1 Announce Type: cross Abstract: Permutation optimization arises whenever the components of a system are fixed but their ordering affects performance. We introduce PRISM, a predictive

Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via Proxy Representations

SafetyDGX agent

arXiv:2608.08245v1 Announce Type: cross Abstract: LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making it difficu

Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation: Guarantees and Empirical Limits

SafetyDGX agent

arXiv:2608.07913v1 Announce Type: cross Abstract: Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate for federated

Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction

Local AiDGX agent

arXiv:2608.08443v1 Announce Type: cross Abstract: Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other parts of a

Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation

SafetyDGX agent

arXiv:2608.09263v1 Announce Type: new Abstract: Outcome verifiers score completed reasoning traces but do not assign credit to intermediate tokens. Privileged self-distillation attempts to fill this g

Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation

SafetyDGX agent

arXiv:2608.09228v1 Announce Type: cross Abstract: On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the

Probabilistic Circuits for Knowledge Graph Completion with Reduced Rule Sets

Model ReleasesDGX agent

arXiv:2508.06706v2 Announce Type: replace Abstract: Rule-based methods for knowledge graph completion provide explainable results, but often require tens of thousands of rules to achieve competitive p

ProbSPARQL: Querying Knowledge Graphs with Multi-dimensional, Uncertain Numeric Data

ResearchDGX agent

arXiv:2607.18262v2 Announce Type: replace Abstract: The SFB 1574 Circular Factory is building a shared knowledge graph infrastructure for integrating data about returned products. A central challenge

Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States

ResearchDGX agent

arXiv:2608.08024v1 Announce Type: cross Abstract: Large language models (LLMs) can generate fluent and useful responses but remain prone to hallucinations. We introduce Prompt Embedding Probes (PEP),

PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary

Model ReleasesDGX agent

arXiv:2608.08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically frame

Protecting patient privacy in clinical foundation models: Technical and legal perspectives

ApplicationsDGX agent

arXiv:2608.07705v1 Announce Type: new Abstract: Clinical foundation models trained on large-scale patient data are increasingly used for decision support, screening, and public health. As deployment e

Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update

SafetyDGX agent

arXiv:2607.11505v2 Announce Type: replace-cross Abstract: Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward behav

Quantization Degradation in Large Language Models: A Signal-Noise Perspective

ResearchDGX agent

arXiv:2608.08188v1 Announce Type: new Abstract: Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-wi

QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing

AgentsDGX agent

arXiv:2608.07743v1 Announce Type: new Abstract: Identifying a meaningful quantum speedup requires more than matching a classical problem to a familiar quantum primitive: the claim must preserve the ta

QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture

Model ReleasesDGX agent

arXiv:2510.22087v3 Announce Type: replace-cross Abstract: The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent from

Query-Only Backdoor Attacks on Self-Evolving Skills via Trajectory Poisoning

AgentsDGX agent

arXiv:2608.08303v1 Announce Type: new Abstract: Agentic skills improve large language model (LLM) agents by encoding reusable procedures for complex tasks. However, manually authored skills often adap

Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis

Model ReleasesDGX agent

arXiv:2509.21629v4 Announce Type: replace-cross Abstract: Program verification relies on loop invariants, yet automatically discovering strong invariants remains a long-standing challenge. We investig

RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation

ApplicationsDGX agent

arXiv:2601.10168v3 Announce Type: replace-cross Abstract: Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet

RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction

ResearchDGX agent

arXiv:2608.09331v1 Announce Type: cross Abstract: Brain-to-audio reconstruction is limited by prior domination: when a pretrained generator is conditioned on a weak neural signal, it produces realisti

RAG-Based Auto-Configuration for Industrial Fieldbus Devices

Model ReleasesDGX agent

arXiv:2608.08618v1 Announce Type: cross Abstract: Industrial device commissioning requires engineers to manually extract hundreds of protocol-specific parameters from heterogeneous PDF manuals and tra

RangeFactory: Scalable Construction of Multi-Hop Cyber Ranges

AgentsDGX agent

arXiv:2608.09526v1 Announce Type: cross Abstract: Real-world cyberattacks often require sustained progress across multiple hosts and network segments, making multi-hop cyber ranges essential infrastru

RankGuide: Tensor-Rank-Guided Routing and Steering for Efficient Reasoning

ResearchDGX agent

arXiv:2604.16694v2 Announce Type: replace Abstract: Large reasoning models (LRMs) enhance problem-solving capabilities by generating explicit multi-step chains of thought (CoT) reasoning; however, the

RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement

ResearchDGX agent

arXiv:2608.09111v1 Announce Type: new Abstract: AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-a

Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression

SafetyDGX agent

arXiv:2608.08960v1 Announce Type: new Abstract: Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduce

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

Model ReleasesDGX agent

arXiv:2608.07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on e

RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2608.09467v1 Announce Type: cross Abstract: Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable

Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response

SafetyDGX agent

arXiv:2506.07223v2 Announce Type: replace Abstract: Large language models (LLMs) have substantially improved the planning capabilities of embodied agents, enabling their deployment in dynamic and safe

REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

SafetyDGX agent

arXiv:2608.07931v1 Announce Type: new Abstract: Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in

REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation

Model ReleasesDGX agent

arXiv:2503.22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require

ReMIND: Orchestrating Modular Large Language Models for Controllable Serendipity A REM-Inspired System Design for Emergent Creative Ideation

ResearchDGX agent

arXiv:2601.07121v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used not only for problem solving but also for creative ideation; however, generating ideas that

Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification

ResearchDGX agent

arXiv:2608.09512v1 Announce Type: new Abstract: Active inference offers a unified framework for perception, learning, and action, but scaling discrete active-inference models to rich spatial and tempo

Representation Matters in Longitudinal Affective Computing

Local AiDGX agent

arXiv:2608.07518v1 Announce Type: cross Abstract: Longitudinal, in-the-wild, wearable sensing yields day-level physiology, sleep, activity, and environmental streams, whereas affect and cognition are

Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing

Model ReleasesDGX agent

arXiv:2608.08514v1 Announce Type: new Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and mod

Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation

ResearchDGX agent

arXiv:2608.08713v1 Announce Type: cross Abstract: Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial

Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach

Local AiDGX agent

arXiv:2608.09742v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., A and B, providing an efficient way to

Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction

Model ReleasesDGX agent

arXiv:2608.09182v1 Announce Type: cross Abstract: Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Existing localiz

Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?

Model ReleasesDGX agent

arXiv:2608.09629v1 Announce Type: new Abstract: Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artif

REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering

AgentsDGX agent

arXiv:2608.08612v1 Announce Type: cross Abstract: Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering. However, existin

Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines

ResearchDGX agent

arXiv:2604.01029v2 Announce Type: replace-cross Abstract: Multi-LLM revision pipelines, in which a second model reviews and improves a draft produced by a first, are widely assumed to derive their gai

RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation

Model ReleasesDGX agent

arXiv:2608.08684v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. Existing met

RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning

Model ReleasesDGX agent

arXiv:2608.09123v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a s

RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation

SafetyDGX agent

arXiv:2608.09226v1 Announce Type: cross Abstract: Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures ar

Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models

Model ReleasesDGX agent

arXiv:2608.07890v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert. Recent theory sh

SafeQL: Search-based Refinement for Safe and Efficient LLM-based Text-to-SQL

TutorialsDGX agent

arXiv:2608.09260v1 Announce Type: cross Abstract: Large language models (LLMs) have advanced Text-to-SQL by enabling natural language interfaces to databases without task-specific fine-tuning. However

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

Model ReleasesDGX agent

arXiv:2608.09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance,

← Previous
1…1011121314…350
Next →