AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,153 results
Safety

The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem

DGX agent

arXiv:2607.26068v1 Announce Type: cross Abstract: Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize

safetyarxiv-cs-ai
31 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG

DGX agent

arXiv:2607.26470v1 Announce Type: new Abstract: Multi-turn information-seeking conversations require both multi-hop reasoning and long-range dependency tracking across turns. However, existing RAG sys

model-releasesarxiv-cs-cl
30 Jul 2026
Applications

ContactFlow: A video action conditioning that transfers across embodiments

DGX agent

arXiv:2607.26579v1 Announce Type: cross Abstract: World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. Howe

applicationsarxiv-cs-cv
30 Jul 2026
Safety

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

DGX agent

arXiv:2607.26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk

safetyarxiv-cs-lg
30 Jul 2026
Safety

Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

DGX agent

arXiv:2607.26060v1 Announce Type: new Abstract: LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical

safetyarxiv-cs-cl
30 Jul 2026
Safety

Parameterized Fair Resource Allocation under Diversity Constraints

DGX agent

arXiv:2607.26485v1 Announce Type: cross Abstract: Resource allocation across multiple agent groups arises in many applications including e-commerce recommendation systems, housing assignment, and cour

safetyarxiv-cs-lg
30 Jul 2026
Safety

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

DGX agent

arXiv:2607.25207v1 Announce Type: new Abstract: This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal pol

safetyarxiv-cs-lg
29 Jul 2026
Model Releases

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

DGX agent

arXiv:2607.24821v1 Announce Type: cross Abstract: While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modali

model-releasesarxiv-cs-cv
29 Jul 2026
Model Releases

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

DGX agent

arXiv:2607.26041v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success o

model-releasesarxiv-cs-ai
29 Jul 2026
Research

Sheet As Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding

DGX agent

arXiv:2605.05811v2 Announce Type: replace Abstract: Workbook-scale spreadsheet understanding is increasingly important for language-model-based data analysis agents, but remains challenging because re

researcharxiv-cs-ai
29 Jul 2026
Model Releases

UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams

DGX agent

arXiv:2607.26017v1 Announce Type: new Abstract: Memory is essential for LLM agents to accumulate task experience and reuse task-specific execution strategies. However, real-world deployment over bound

model-releasesarxiv-cs-cl
29 Jul 2026
Research

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

DGX agent

arXiv:2607.24343v1 Announce Type: cross Abstract: Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body bu

researcharxiv-cs-ai
28 Jul 2026
Model Releases

Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework

DGX agent

arXiv:2607.22715v1 Announce Type: new Abstract: The rapid growth of Artificial Intelligence-generated content (AIGC) is reshaping video production and circulation, exposing children to an increasing v

model-releasesarxiv-cs-cv
28 Jul 2026
Safety

Constrained Reinforcement Learning Using Successor Representations

DGX agent

arXiv:2607.24057v1 Announce Type: new Abstract: Real-world Reinforcement Learning depends on the ability to formulate safety constraints into a policy. A common way to model such constraints is to int

safetyarxiv-cs-lg
28 Jul 2026
Safety

Data Pyramid for Embodied Manipulation

DGX agent

arXiv:2607.24744v1 Announce Type: cross Abstract: Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require d

safetyarxiv-cs-cv
28 Jul 2026
Model Releases

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

DGX agent

arXiv:2607.24241v1 Announce Type: cross Abstract: Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw p

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric

DGX agent

arXiv:2607.23532v1 Announce Type: cross Abstract: Swarms of LLM-assisted autonomous robots are increasingly proposed for cooperative intelligence, surveillance, and reconnaissance (ISR) in contested e

safetyarxiv-cs-ai
28 Jul 2026
Model Releases

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

DGX agent

arXiv:2607.22951v1 Announce Type: cross Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified ope

model-releasesarxiv-cs-ai
28 Jul 2026
Hardware

Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

DGX agent

arXiv:2607.22997v1 Announce Type: cross Abstract: Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next

hardwarearxiv-cs-ai
28 Jul 2026
Model Releases

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

DGX agent

arXiv:2607.24063v1 Announce Type: new Abstract: On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is tow

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost

DGX agent

arXiv:2607.24010v1 Announce Type: new Abstract: Active RAG systems decide when to retrieve external knowledge during generation, making them a budget-sensitive case of agentic RAG and self-adaptive re

safetyarxiv-cs-lg
28 Jul 2026
Safety

Cross-reality location privacy protection in 6G-enabled vehicular metaverses: an LLM-enhanced hybrid generative diffusion model-based approach

DGX agent

arXiv:2601.12311v2 Announce Type: replace-cross Abstract: The emergence of 6G-enabled vehicular metaverses enables Autonomous Vehicles (AVs) to operate across physical and virtual spaces through space

safetyarxiv-cs-lg
27 Jul 2026
Safety

From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE

DGX agent

arXiv:2607.21608v1 Announce Type: cross Abstract: With the EU AI Act entering into force, organizations developing or operating AI systems face new obligations on transparency, risk management, and tr

safetyarxiv-cs-cl
27 Jul 2026
Safety

Learning Spatiotemporal Decision Priors for Efficient Path Planning under Partial Observability

DGX agent

arXiv:2607.22166v1 Announce Type: new Abstract: Path planning under partial observability remains challenging because an agent must make long-horizon navigation decisions from only locally bounded obs

safetyarxiv-cs-ro
27 Jul 2026
Safety

Approximate Quantum State Preparation Through Proximal Policy Optimization

DGX agent

arXiv:2607.21121v1 Announce Type: cross Abstract: In this work, a quantum architecture search framework for approximate quantum state preparation (QSP) is proposed. QSP is a challenging task, since th

safetyarxiv-cs-lg
24 Jul 2026
Model Releases

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

DGX agent

arXiv:2607.20465v1 Announce Type: cross Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how

model-releasesarxiv-cs-cl
24 Jul 2026
Applications

Learning to Detect UI Principle Violations via Reinforcement Learning

DGX agent

arXiv:2607.20690v1 Announce Type: new Abstract: Small language models and coding agents increasingly generate web front-end code, yet their outputs are typically evaluated primarily for functional cor

applicationsarxiv-cs-cl
24 Jul 2026
Model Releases

LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

DGX agent

arXiv:2607.20872v1 Announce Type: new Abstract: Long-form legal research reports increasingly rely on LLMs and agentic research systems, but their reliability depends not only on answering the task, b

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

DGX agent

arXiv:2607.21111v1 Announce Type: cross Abstract: Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

DGX agent

arXiv:2607.19747v1 Announce Type: new Abstract: As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream gen

model-releasesarxiv-cs-cl
23 Jul 2026
Applications

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

DGX agent

arXiv:2607.19228v1 Announce Type: new Abstract: Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear

applicationsarxiv-cs-cv
23 Jul 2026
Model Releases

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation

DGX agent

arXiv:2607.19695v1 Announce Type: new Abstract: Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode

model-releasesarxiv-cs-ro
23 Jul 2026
Safety

REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning

DGX agent

arXiv:2607.19450v1 Announce Type: cross Abstract: Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool

safetyarxiv-cs-ai
23 Jul 2026
Model Releases

The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL

DGX agent

arXiv:2607.19749v1 Announce Type: cross Abstract: Model-based reinforcement-learning agents of the DreamerV3 family forget catastrophically when trained on task sequences, even when an unbounded repla

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

Unified Prediction and Planning via Conflict-Aware Disjoint Parameter Training

DGX agent

arXiv:2607.19971v1 Announce Type: new Abstract: Accurate motion prediction of surrounding agents and safe motion planning are two closely coupled key tasks for social robot navigation in crowded envir

model-releasesarxiv-cs-ro
23 Jul 2026
Safety

A Causal Model of Theory of Mind in Conflict for Artificial Intelligence

DGX agent

arXiv:2606.16944v2 Announce Type: replace Abstract: Theory of mind (ToM), the capacity to ascribe mental states to others and use those ascriptions for prediction and inference, is widely assumed to b

safetyarxiv-cs-ai
16 Jul 2026
Safety

Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference

DGX agent

arXiv:2607.13205v1 Announce Type: cross Abstract: Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accum

safetyarxiv-cs-ai
16 Jul 2026
Safety

Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management

DGX agent

arXiv:2607.13239v1 Announce Type: new Abstract: Foundation models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used for transportation management center

safetyarxiv-cs-ai
16 Jul 2026
Safety

Joint On-and-Off Policy Learning for Vision-and-Language Navigation

DGX agent

arXiv:2607.13461v1 Announce Type: new Abstract: Vision-and-Language Navigation (VLN) necessitates an embodied agent to navigate in the physical world by adhering to natural language instructions. Rece

safetyarxiv-cs-ro
16 Jul 2026
Research

Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics

DGX agent

arXiv:2607.13245v1 Announce Type: new Abstract: While 3D Scene Graphs (3DSGs) provide crucial structured representations for embodied agents, conventional Ahead-of-Time, build-everything-then-filter p

researcharxiv-cs-cv
16 Jul 2026
Model Releases

Beyond Perfect Priors: Adaptive Gaussian Graph for 4D Driving Reconstruction in the Wild

DGX agent

arXiv:2607.12214v1 Announce Type: new Abstract: Reconstructing 4D driving scenes in the wild (e.g., internet and AI-generated videos) is critical for diverse autonomous driving simulation. While recen

model-releasesarxiv-cs-cv
15 Jul 2026
Safety

Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction

DGX agent

arXiv:2607.12835v1 Announce Type: new Abstract: Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, whe

safetyarxiv-cs-cl
15 Jul 2026
Safety

Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

DGX agent

arXiv:2607.12273v1 Announce Type: cross Abstract: As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, w

safetyarxiv-cs-ai
15 Jul 2026
Research

Decentralized Gradient Descent: Bottleneck Regimes and Budget Complexity

DGX agent

arXiv:2607.12172v1 Announce Type: cross Abstract: Decentralized gradient descent (DGD) is widely used for solving distributed optimization problems over networks of agents. While its convergence prope

researcharxiv-cs-lg
15 Jul 2026
Tutorials

Diversity-Enriched Option-Critic

DGX agent

arXiv:2011.02565v2 Announce Type: replace-cross Abstract: Temporal abstraction allows reinforcement learning agents to represent knowledge and develop strategies over different temporal scales. The op

tutorialsarxiv-cs-ai
15 Jul 2026
Model Releases

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

DGX agent

arXiv:2512.01241v4 Announce Type: replace-cross Abstract: Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

How Inference Compute Shapes Frontier LLM Evaluation

DGX agent

arXiv:2606.17930v2 Announce Type: replace Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result,

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

DGX agent

arXiv:2607.12924v1 Announce Type: new Abstract: In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic acti

model-releasesarxiv-cs-ai
15 Jul 2026
← Previous
1…173174175176177…233
Next →