AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
11,040 results
Research

Are Large Language Models Suitable for Graph Computation? Progress and Prospects

DGX agent

arXiv:2606.06865v1 Announce Type: new Abstract: Large language models (LLMs) have been increasingly explored for graph computation, where tasks require reasoning over structured relationships and algo

researcharxiv-cs-cl
8 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

CHDP: Cooperative Hybrid Diffusion Policies for Reinforcement Learning in Parameterized Action Space

DGX agent

arXiv:2601.05675v2 Announce Type: replace Abstract: Hybrid action space, which combines discrete choices and continuous parameters, is prevalent in domains such as robot control and game AI. However,

safetyarxiv-cs-ai
8 Jun 2026
Research

ChronoForest: Closed-Loop Multi-Tree Diffusion Planning for Efficient Bridge Search and Route Composition

DGX agent

arXiv:2606.06618v1 Announce Type: cross Abstract: How can we plan long-horizon routes that reach designated goals, visit required waypoints, and remain short when only short-horizon offline trajectori

researcharxiv-cs-ai
8 Jun 2026
Safety

Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios

DGX agent

arXiv:2606.06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know. Existing benchmarks emphasize domain

safetyarxiv-cs-lg
8 Jun 2026
Safety

Learning All-Terrain Locomotion for a Planetary Rover with Actively Articulated Suspension

DGX agent

arXiv:2606.06790v1 Announce Type: cross Abstract: This paper presents ERNEST, a four-wheeled planetary rover concept equipped with a two-degree-of-freedom Active Gimbal Suspension that combines yaw an

safetyarxiv-cs-lg
8 Jun 2026
Safety

LLM-Augmented Digital Twin for Policy Evaluation in Short-Video Platforms

DGX agent

arXiv:2603.11333v2 Announce Type: replace Abstract: Short-video platforms are closed-loop, human-in-the-loop ecosystems where platform policy, creator incentives, and user behavior co-evolve. This fee

safetyarxiv-cs-ai
8 Jun 2026
Model Releases

MMAE: A Massive Multitask Audio Editing Benchmark

DGX agent

arXiv:2606.07229v1 Announce Type: cross Abstract: We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose ins

model-releasesarxiv-cs-cl
8 Jun 2026
Research

Performance Variation in Deep Reinforcement Learning

DGX agent

arXiv:2606.06746v1 Announce Type: new Abstract: Deep reinforcement learning (RL) algorithms often suffer from low run-to-run robustness, manifesting as significant performance variation across indepen

researcharxiv-cs-lg
8 Jun 2026
Safety

Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability

DGX agent

arXiv:2606.07437v1 Announce Type: cross Abstract: The ISO 26262 standard defines functional safety for road vehicles through risk assessments based on Severity, Exposure, and Controllability, grounded

safetyarxiv-cs-ai
8 Jun 2026
Model Releases

RealDocBench: A Benchmark for Field-Level QA and Layout Understanding on Real-World Regulated Documents

DGX agent

arXiv:2606.07401v1 Announce Type: new Abstract: Document parsing systems are increasingly deployed in high-stakes, regulated workflows such as mortgage underwriting, financial reporting, supply-chain

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

ScenicRules: An Autonomous Driving Benchmark with Multi-Objective Specifications and Abstract Scenarios

DGX agent

arXiv:2602.16073v2 Announce Type: replace-cross Abstract: Developing autonomous driving systems for complex traffic environments requires balancing multiple objectives, such as avoiding collisions, ob

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Benchmarking Counterfactual Prediction in Epidemic Time Series with Time-Varying Interventions

DGX agent

arXiv:2606.05692v1 Announce Type: cross Abstract: Deep learning has enabled significant advances in time-series causal inference, yet progress remains constrained by the lack of realistic benchmarks w

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

DGX agent

arXiv:2606.05463v1 Announce Type: new Abstract: Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically perf

model-releasesarxiv-cs-ai
6 Jun 2026
Research

Stable Deep Reinforcement Learning via Isotropic Gaussian Representations

DGX agent

arXiv:2602.19373v3 Announce Type: replace-cross Abstract: Deep reinforcement learning systems often suffer from unstable training dynamics due to non-stationarity, where learning objectives and data d

researcharxiv-cs-ai
6 Jun 2026
Model Releases

Epidemiology of Model Collapse: Modeling Synthetic Data Contamination via Bilayer SIR Dynamics

DGX agent

arXiv:2606.05168v1 Announce Type: new Abstract: Training on synthetic data causes model collapse, but existing analyses treat this as single-chain degradation. In reality, the AI ecosystem involves cr

model-releasesarxiv-cs-cl
5 Jun 2026
Safety

EVE: A Generator-Verifier System for Generative Policies

DGX agent

arXiv:2512.21430v2 Announce Type: replace Abstract: Visuomotor policies based on generative such as diffusion and flow-matching have shown strong performance for robotics applications but degrade unde

safetyarxiv-cs-ro
5 Jun 2026
Safety

LadderMan: Learning Humanoid Perceptive Ladder Climbing

DGX agent

arXiv:2606.05873v1 Announce Type: cross Abstract: Humanoid robots hold great promise for operating in human-centered environments, yet ladder climbing remains one of the most challenging tasks due to

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

DGX agent

arXiv:2606.05744v1 Announce Type: new Abstract: Spatial planning maps are central to territorial governance, translating planning objectives, regulations, and spatial strategies into visual forms for

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)

DGX agent

arXiv:2606.05901v1 Announce Type: new Abstract: Large language models (LLMs) have fundamentally transformed the landscape of Natural Language Processing. Despite these advances, LLMs and LLM-based sys

model-releasesarxiv-cs-cl
5 Jun 2026
Safety

Robust Scene Transfer for PointGoal Navigation via Privileged Sensor Guided Contrastive Learning

DGX agent

arXiv:2606.05506v1 Announce Type: new Abstract: We propose a sensor-guided adaptive contrastive learning framework for visual representation learning in PointGoal navigation. During training, privileg

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation

DGX agent

arXiv:2606.05660v1 Announce Type: new Abstract: Embodied AI systems are increasingly expected to reason and act over extended horizons in physical environments. This growing capability brings safety t

model-releasesarxiv-cs-ro
5 Jun 2026
Model Releases

SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations

DGX agent

arXiv:2606.05563v1 Announce Type: cross Abstract: Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison

DGX agent

arXiv:2606.05436v1 Announce Type: cross Abstract: Summarizing the latest medical literature to guide clinical decision-making is essential for evidence-based medicine and high-quality patient care. Ye

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Affordance2Action: Task-Conditioned Scene-level Affordance Grounding for Real-Time Manipulation

DGX agent

arXiv:2606.04172v1 Announce Type: new Abstract: Task-conditioned manipulation requires grounding instructions to task-relevant functional parts rather than object categories. This setting is scene-dep

model-releasesarxiv-cs-ro
4 Jun 2026
Research

Beyond Pixel Histories: World Models with Persistent 3D State

DGX agent

arXiv:2603.03482v2 Announce Type: replace-cross Abstract: Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities. However, e

researcharxiv-cs-ai
4 Jun 2026
Model Releases

Bilevel Autoresearch: Meta-Autoresearching Itself

DGX agent

arXiv:2603.23420v2 Announce Type: replace Abstract: If autoresearch is itself a form of research, then autoresearch can be applied to research itself. We present Bilevel Autoresearch, a bilevel framew

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

DiffAero: A GPU-Accelerated Differentiable Simulation Framework for Efficient Quadrotor Policy Learning

DGX agent

arXiv:2509.10247v1 Announce Type: cross Abstract: This letter introduces DiffAero, a lightweight, GPU-accelerated, and fully differentiable simulation framework designed for efficient quadrotor contro

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

DLO-Lab: Benchmarking Deformable Linear Object Manipulations with Differentiable Physics

DGX agent

arXiv:2606.04206v1 Announce Type: new Abstract: We address the challenge of enabling robots to manipulate deformable linear objects (DLOs), such as ropes, cables, and rubber bands. Prior work has prim

model-releasesarxiv-cs-ro
4 Jun 2026
Model Releases

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

DGX agent

arXiv:2606.04282v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are predominantly evaluated on free-form vision-language tasks such as visual question answering, captioning, a

model-releasesarxiv-cs-cv
4 Jun 2026
Local Ai

MAD: Mapping-Aware World Models for Agile Quadrotor Flight

DGX agent

arXiv:2606.04534v1 Announce Type: new Abstract: Agile quadrotor flight in cluttered scenes requires more than a reactive mapping from a depth image to a control command: the vehicle must remember whic

local-aiarxiv-cs-ro
4 Jun 2026
Model Releases

MineXplore: An Open-Source Reinforcement Learning Exploration Benchmark for GNSS-Denied Underground Environment

DGX agent

arXiv:2606.04569v1 Announce Type: new Abstract: Underground mines present extreme conditions for autonomous robot navigation: GPS is denied, lighting is degraded, and tunnel topology is loop-rich and

model-releasesarxiv-cs-ro
4 Jun 2026
Research

Reasoning Shift: How Context Silently Shortens LLM Reasoning

DGX agent

arXiv:2604.01161v2 Announce Type: replace Abstract: Large language models (LLMs) exhibiting test-time scaling behavior, such as extended reasoning traces and self-verification, have demonstrated remar

researcharxiv-cs-lg
4 Jun 2026
Safety

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

DGX agent

arXiv:2606.04923v1 Announce Type: cross Abstract: Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

Self-Evolving Deep Research via Joint Generation and Evaluation

DGX agent

arXiv:2606.04507v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capab

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning

DGX agent

arXiv:2606.04735v1 Announce Type: cross Abstract: Temporal credit assignment is central to both biological and artificial intelligence, yet its interaction with non-linear function approximation is po

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

Acceptance-Test-Driven Evaluation Protocols for Business-Centric LLM Systems

DGX agent

arXiv:2606.02755v1 Announce Type: cross Abstract: Large language model (LLM) applications are increasingly expected to satisfy deterministic institutional requirements while relying on probabilistic g

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

AURA: Action-Gated Memory for Robot Policies at Constant VRAM

DGX agent

arXiv:2606.02775v1 Announce Type: new Abstract: The KV-cache is the right memory for datacenters but the wrong memory for robots. Datacenter inference batches many short requests and resets them, amor

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Cosmos 3: Omnimodal World Models for Physical AI

DGX agent

arXiv:2606.02800v1 Announce Type: cross Abstract: We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Identifying Quantum Structure in AI Language: Evidence for Evolutionary Convergence of Human and Artificial Cognition

DGX agent

arXiv:2511.21731v2 Announce Type: replace-cross Abstract: We present the results of cognitive tests on conceptual combinations, performed using specific Large Language Models (LLMs) as test subjects.

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Learning to Bet for Horizon-Aware Anytime-Valid Testing

DGX agent

arXiv:2603.19551v2 Announce Type: replace-cross Abstract: We develop horizon-aware anytime-valid tests and confidence sequences for bounded means under a strict deadline N. Using the betting/e-process

safetyarxiv-cs-lg
3 Jun 2026
Safety

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving

DGX agent

arXiv:2606.02964v1 Announce Type: cross Abstract: Large Language Model (LLM) inference relies on key-value (KV) caches to avoid redundant attention computation. While approximate KV cache retention te

safetyarxiv-cs-cl
3 Jun 2026
Safety

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation

DGX agent

arXiv:2606.03159v1 Announce Type: cross Abstract: As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-lo

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

Proof-Refactor: Refactoring Generated Formal Proofs into Modular Artifacts

DGX agent

arXiv:2606.03743v1 Announce Type: new Abstract: While Large Language Models (LLMs) have shown strong performance in generating formal proofs, their outputs often remain less readable, modular, maintai

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series

DGX agent

arXiv:2606.03301v1 Announce Type: new Abstract: We introduce SagaQA, a long-form video benchmark for multi-hop reasoning over full-length TV series. Existing video reasoning benchmarks often emphasize

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Spike-Aware C++ INT8 Inference for Sparse Spiking Language Models on Commodity CPUs

DGX agent

arXiv:2606.03026v1 Announce Type: cross Abstract: Spiking language models expose activation sparsity that dense Transformer runtimes do not directly exploit. This paper studies that property from a sy

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning

DGX agent

arXiv:2606.03962v1 Announce Type: cross Abstract: Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward. Yet, modern applicati

safetyarxiv-cs-ai
3 Jun 2026
Research

A Theoretical Framework for Self-Play Theorem Proving Algorithms

DGX agent

arXiv:2606.01861v1 Announce Type: new Abstract: Self-play, a type of training algorithm that enables a model to self-improve, has recently shown promising empirical results in the context of formal th

researcharxiv-cs-lg
2 Jun 2026
Applications

Algorithmic algorithm development with LLMs: A Case Study on LLM-Usage for Contraction Order Optimization in Tensor Networks

DGX agent

arXiv:2606.01975v1 Announce Type: new Abstract: We consider LLM-based algorithm development through a case study on contractionorder optimisation for tensor networks with OpenEvolve. We pay particular

applicationsarxiv-cs-ai
2 Jun 2026
← Previous
1…211212213214215…230
Next →