AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
10 Aug 2026

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

Model ReleasesDGX agent

arXiv:2608.06931v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory

Seeking SOTA: Time-Series Forecasting Must Adopt Taxonomy-Specific Evaluation to Dispel Illusory Gains

Model ReleasesDGX agent

arXiv:2603.15506v2 Announce Type: replace-cross Abstract: We argue that the current practice of evaluating AI/ML time-series forecasting models, predominantly on benchmarks characterized by strong, pe

Semantic Adapter Routing with Fine-Tuning Task Embeddings

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.19079v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning (PEFT) has led to model ecosystems in which a single backbone is paired with many task-specialized adapters. Given s

Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models

ResearchDGX agent

arXiv:2608.06409v1 Announce Type: cross Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures a

SetEasy: A Multi-Modal Classroom Engagement Assessment and Seating Optimization Framework

ResearchDGX agent

arXiv:2608.07188v1 Announce Type: new Abstract: SetEasy optimizes classroom engagement in fixed seating grids. It fuses multimodal sensing (wristband physiology, 4K video, environmental data) and trai

Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation

SafetyDGX agent

arXiv:2608.06632v1 Announce Type: new Abstract: Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.g., cl

SkillEval: Decomposing Agent Skill Quality into Interpretable Signals

AgentsDGX agent

arXiv:2608.06891v1 Announce Type: new Abstract: Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill quality becomes inc

SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent

AgentsDGX agent

arXiv:2608.07449v1 Announce Type: new Abstract: LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifact

SLED: Scalable Location Encoding via Distillation

Model ReleasesDGX agent

arXiv:2608.06612v1 Announce Type: cross Abstract: The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer siz

Social World Models

Model ReleasesDGX agent

arXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited informat

Soft Redaction of Image Provenance via Zero-Knowledge Proofs

ResearchDGX agent

arXiv:2608.07063v1 Announce Type: cross Abstract: Content provenance standards, such as C2PA, are increasingly used to attach signed records of origin, editing history, and rights to digital images. H

SoRoMoX: Fast, Differentiable, and Parallelizable Soft Robot Models

SafetyDGX agent

arXiv:2608.06650v1 Announce Type: cross Abstract: Reduced-order models based on Cosserat-rod theory are now well established, and modeling theory is no longer the primary bottleneck in soft-robot cont

Stability of Transformers under Layer Normalization

ResearchDGX agent

arXiv:2510.09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable. Layer normalization, a standard component, improves training stabili

StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection

Model ReleasesDGX agent

arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as

Strategy-first synthesis planning for complex natural products

AgentsDGX agent

arXiv:2608.07454v1 Announce Type: cross Abstract: The total synthesis of a complex molecule is among the most demanding intellectual and experimental feats in chemistry: a chemist must plan many steps

Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts

Model ReleasesDGX agent

arXiv:2608.06770v1 Announce Type: new Abstract: Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic ins

SyncSBC: Decentralized Swarm Behavior Prediction for Synchronized Autonomous Control

AgentsDGX agent

arXiv:2608.06587v1 Announce Type: cross Abstract: Robot swarms utilize many independent limited-sensing agents to produce complex emergent behaviors without requiring centralized control. However, lit

TaskSense: Focusing on What Matters in World Models

TutorialsDGX agent

arXiv:2608.06544v1 Announce Type: new Abstract: World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools

ApplicationsDGX agent

arXiv:2608.07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As generative AI app

TEPA: Revoking Stale Memories for Conflict-Robust Language Agents

ResearchDGX agent

arXiv:2608.07429v1 Announce Type: new Abstract: Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also creates a central falsifiability proble

TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation

ResearchDGX agent

arXiv:2608.06396v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models route each token through a small subset of experts, making routing patterns useful for identifying task-relev

The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows

SafetyDGX agent

arXiv:2608.06714v1 Announce Type: new Abstract: Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits

The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products

AgentsDGX agent

arXiv:2606.15485v2 Announce Type: replace-cross Abstract: Agentic AI systems act autonomously, use tools, adapt to context, and operate in complex real-world environments. However, these same characte

TOFD: Target-Oriented Feature Decoupling against Poisoning Attacks in Split Federated Learning

ResearchDGX agent

arXiv:2608.07274v1 Announce Type: cross Abstract: Split Federated Learning (SFL) facilitates privacy-preserving collaborative training with reduced client-side overhead. However, its split architectur

Toward a Causal Data Management Ecosystem for Decision Making and Agentic AI

AgentsDGX agent

arXiv:2608.07214v1 Announce Type: cross Abstract: Modern AI is no longer a single model but an ecosystem: classical ML predictors, deep and multimodal models, large language models, and agents, each t

Towards a Theoretical Understanding of Two Tower Recommendation Models

ApplicationsDGX agent

arXiv:2403.00802v2 Announce Type: replace-cross Abstract: Production-grade recommender systems rely heavily on a large-scale corpus used by online media services, including Netflix, Pinterest, and Ama

Towards Assurance Closure in AI-Native Large-Scale Agile Software Development

AgentsDGX agent

arXiv:2608.07317v1 Announce Type: cross Abstract: The AI-Native Manifesto envisions large-scale agile software development in which humans increasingly govern intent, risk, and exceptions while agents

Towards Multi-Label Graph Foundation Models: from Single-Vector Representation Learning to Multi-Semantic Basis Learning

ResearchDGX agent

arXiv:2608.06394v1 Announce Type: new Abstract: Multi-label node classification is an important yet challenging task in graph learning, where nodes exhibit multiple semantics simultaneously. Existing

TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure

Model ReleasesDGX agent

arXiv:2608.06657v1 Announce Type: new Abstract: Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustw

TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade

Model ReleasesDGX agent

arXiv:2608.06549v1 Announce Type: cross Abstract: LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents

Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking

Model ReleasesDGX agent

arXiv:2608.07077v1 Announce Type: new Abstract: The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). Current models solve the sta

TransSLR: A Lightweight Transformer for Sign Language Recognition

Model ReleasesDGX agent

arXiv:2608.06407v1 Announce Type: cross Abstract: Automated Sign Language Recognition for under-represented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifi

TRIBE: Predicting Team Performance via Communication Behavior Ensembles

Model ReleasesDGX agent

arXiv:2608.06926v1 Announce Type: new Abstract: Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present

Unsupervised Adaptation of PDE Foundation Models

ResearchDGX agent

arXiv:2608.07053v1 Announce Type: new Abstract: Pretrained partial differential equation (PDE) foundation models can generalize across different equations, but adapting them to unseen PDE systems typi

Vehicle routing problem using deep reinforcement learning - A case study about truck planning in the industry

AgentsDGX agent

arXiv:2608.06668v1 Announce Type: new Abstract: As an important component of the supply chain industry, transportation has experienced rapid development in the past decade with the assistance of digit

WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader

Model ReleasesDGX agent

arXiv:2608.06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approa

WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance

Model ReleasesDGX agent

arXiv:2608.06704v1 Announce Type: new Abstract: Delegating a web task involves more than asking a question; it requires transferring a policy: what to verify, how to handle uncertainty, which preferen

Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons

Model ReleasesDGX agent

arXiv:2608.07303v1 Announce Type: new Abstract: Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and the

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

SafetyDGX agent

arXiv:2608.07267v1 Announce Type: new Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies t

WorldMark: A Plug-and-Play World Knowledge Interface for Cross-Host Language Model Watermarking

Model ReleasesDGX agent

arXiv:2608.06416v1 Announce Type: cross Abstract: Watermarking traces the provenance of text produced by large language models by embedding statistically detectable signals during decoding. Existing s

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

Model ReleasesDGX agent

arXiv:2608.07341v1 Announce Type: cross Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. extbf{Contamination mitigation

ZIPBrain: Can EEG Foundation Models Be Faster, Locally Deployable, but Accurate?

Local AiDGX agent

arXiv:2608.07033v1 Announce Type: new Abstract: This work investigates whether Electroencephalograph (EEG) foundation models (EFMs) can be made faster and locally deployable without sacrificing accura

7 Aug 2026

A Lexical Analysis of online Reviews on Human-AI Interactions

ResearchDGX agent

arXiv:2511.13480v2 Announce Type: replace-cross Abstract: This study focuses on understanding the complex dynamics between humans and AI systems by analyzing user reviews. While previous research has

A note on conditional PAC-efficient reasoning in large language model routing

ResearchDGX agent

arXiv:2512.03057v2 Announce Type: replace-cross Abstract: We study distribution-free risk control for model routing, motivated by large language model reasoning. We formalize pointwise conditional eff

A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper

ResearchDGX agent

arXiv:2608.05165v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we study the use o

A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2608.05791v1 Announce Type: cross Abstract: Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and thei

A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decomposition

ResearchDGX agent

arXiv:2608.05673v1 Announce Type: new Abstract: Trajectory prediction has shifted toward structured formulations with explicit social modeling. However, existing methods inadequately distinguish the f

ABC: Numerical Data Collection under Local Differential Privacy without Prior Knowledge

ResearchDGX agent

arXiv:2608.05737v1 Announce Type: cross Abstract: Local Differential Privacy (LDP) provides strong privacy guarantees for collecting numerical data. A fundamental challenge, however, is that existing

Abstract Event Causal Rules: Induction and Application

Model ReleasesDGX agent

arXiv:2608.05205v1 Announce Type: new Abstract: Event-centric intelligent analytical systems heavily depend on explicit causal event knowledge for risk early warning, decision-making support and narra

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

Local AiDGX agent

arXiv:2608.05784v1 Announce Type: new Abstract: Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent's memory today records what the

Adaptive Arena-based Contestable Argumentative Network-of-Experts for Open-Ended Care Plan Coordination

SafetyDGX agent

arXiv:2608.05391v1 Announce Type: new Abstract: Care plan coordination demands synthesizing heterogeneous clinical, functional, and psychosocial information across multiple professional disciplines, w

AegisShield: Democratizing Cyber Threat Modeling with Generative AI

ResearchDGX agent

arXiv:2509.10482v2 Announce Type: replace-cross Abstract: The increasing sophistication of technology systems makes traditional threat modeling hard to scale, especially for small organizations with l

Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services

AgentsDGX agent

arXiv:2608.05159v1 Announce Type: new Abstract: Enterprise operations extensively rely on multiple heterogeneous business systems and information applications, which also result in severe data silos a

Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks

Model ReleasesDGX agent

arXiv:2608.05266v1 Announce Type: new Abstract: Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and sync

Agentic Software Issue Resolution with Large Language Models: A Survey

AgentsDGX agent

arXiv:2512.22256v2 Announce Type: replace-cross Abstract: Software issue resolution aims to address real-world issues in software repositories based on natural language descriptions provided by users,

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2608.05987v1 Announce Type: new Abstract: Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisi

AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations

Model ReleasesDGX agent

arXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr

All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Generalizable Training

SafetyDGX agent

arXiv:2601.03895v2 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs). How

An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics

Model ReleasesDGX agent

arXiv:2604.15145v2 Announce Type: replace Abstract: The rigorous evaluation of the novelty of a scientific paper is, even for human scientists, a challenging task. With the increasing interest in AI s

An Emerging Retail Portfolio Management Application: Personalized, Tax-Aware Reinforcement Learning with Natural Language Goals

SafetyDGX agent

arXiv:2608.05255v1 Announce Type: cross Abstract: Retail investors lack access to the kind of personalized, tax-aware portfolio management that institutional clients take for granted -- existing robo-

← Previous
1…2122232425…354
Next →