AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,724 results
Safety

Repair Before Veto, When Repair Is Hidden: Quantum-Accessible Features for Repair-Augmented Constraint Learning

DGX agent

arXiv:2606.08020v1 Announce Type: cross Abstract: Hard-constraint decision systems usually veto infeasible candidates. This is too rigid when the system can act: if a known affordable repair would mak

safetyarxiv-cs-ai
9 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Report: GKE Inference Gateway delivers up to 92% faster AI responses

DGX agent

As generative AI moves from experimental pilots to massive production environments, the efficiency of your infrastructure becomes the ultimate differentiator. One way to get the most out of it and min

model-releasesgoogle-cloud-ai
9 Jun 2026
Model Releases

SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing

DGX agent

arXiv:2602.09809v2 Announce Type: replace Abstract: Scientific diagrams convey explicit structural information, yet modern text-to-image models often produce visually plausible but structurally incorr

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

SLMJury: Can Small Language Models Judge as Well as Large Ones?

DGX agent

arXiv:2606.07810v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used as judges for evaluating model outputs, but their high cost, latency, and opacity limit scalability. We i

model-releasesarxiv-cs-ai
9 Jun 2026
Research

SynManDex: Synthesizing Human-like Dexterous Grasps from Synthetic Human Pre-Grasps

DGX agent

arXiv:2606.09798v1 Announce Type: new Abstract: Human hand-object interactions encode functional intent, but direct transfer to robotic hands often fails under morphology, contact, and reachability co

researcharxiv-cs-ro
9 Jun 2026
Safety

Systems-Level Planning and Coordination of Truck-Drone Collaborative Delivery Networks

DGX agent

arXiv:2606.08738v1 Announce Type: cross Abstract: Urban last-mile parcel delivery increasingly relies on heterogeneous fleets whose performance depends on timely coordination, reliable communication,

safetyarxiv-cs-ro
9 Jun 2026
Safety

The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In

DGX agent

arXiv:2606.08172v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate high-stakes interactions in finance, medicine, and mental-health support, yet users have limited con

safetyarxiv-cs-ai
9 Jun 2026
Local Ai

this model is the opposite of mythos. Its small, cost effective, apache 2.0, and locally deployable. This is the way LLMs should go. small, …

DGX agent

this model is the opposite of mythos. Its small, cost effective, apache 2.0, and locally deployable. This is the way LLMs should go. small, open source, transparent and sovereign vs large, expensive,

local-aiclem-delangue--x
9 Jun 2026
Research

Are Large Language Models Suitable for Graph Computation? Progress and Prospects

DGX agent

arXiv:2606.06865v1 Announce Type: new Abstract: Large language models (LLMs) have been increasingly explored for graph computation, where tasks require reasoning over structured relationships and algo

researcharxiv-cs-cl
8 Jun 2026
Safety

CHDP: Cooperative Hybrid Diffusion Policies for Reinforcement Learning in Parameterized Action Space

DGX agent

arXiv:2601.05675v2 Announce Type: replace Abstract: Hybrid action space, which combines discrete choices and continuous parameters, is prevalent in domains such as robot control and game AI. However,

safetyarxiv-cs-ai
8 Jun 2026
Research

ChronoForest: Closed-Loop Multi-Tree Diffusion Planning for Efficient Bridge Search and Route Composition

DGX agent

arXiv:2606.06618v1 Announce Type: cross Abstract: How can we plan long-horizon routes that reach designated goals, visit required waypoints, and remain short when only short-horizon offline trajectori

researcharxiv-cs-ai
8 Jun 2026
Safety

Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios

DGX agent

arXiv:2606.06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know. Existing benchmarks emphasize domain

safetyarxiv-cs-lg
8 Jun 2026
Hardware

Good take My guess is - demand for intelligence is near infinite - but 80% of workloads will be running on 99% cheaper models within 12-18 m…

DGX agent

Good take My guess is - demand for intelligence is near infinite - but 80% of workloads will be running on 99% cheaper models within 12-18 months - 20% of workloads will still run on latest gen models

hardwareclem-delangue--x
8 Jun 2026
Safety

Learning All-Terrain Locomotion for a Planetary Rover with Actively Articulated Suspension

DGX agent

arXiv:2606.06790v1 Announce Type: cross Abstract: This paper presents ERNEST, a four-wheeled planetary rover concept equipped with a two-degree-of-freedom Active Gimbal Suspension that combines yaw an

safetyarxiv-cs-lg
8 Jun 2026
Safety

LLM-Augmented Digital Twin for Policy Evaluation in Short-Video Platforms

DGX agent

arXiv:2603.11333v2 Announce Type: replace Abstract: Short-video platforms are closed-loop, human-in-the-loop ecosystems where platform policy, creator incentives, and user behavior co-evolve. This fee

safetyarxiv-cs-ai
8 Jun 2026
Model Releases

MMAE: A Massive Multitask Audio Editing Benchmark

DGX agent

arXiv:2606.07229v1 Announce Type: cross Abstract: We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose ins

model-releasesarxiv-cs-cl
8 Jun 2026
Research

Performance Variation in Deep Reinforcement Learning

DGX agent

arXiv:2606.06746v1 Announce Type: new Abstract: Deep reinforcement learning (RL) algorithms often suffer from low run-to-run robustness, manifesting as significant performance variation across indepen

researcharxiv-cs-lg
8 Jun 2026
Tools

Phoenix at 10,000 stars on GitHub: How an open source AI observability project grew by following its community

DGX agent

Phoenix crossed 10,000 GitHub stars. Here is how the open-source AI observability project grew from a Jupyter notebook extension into a community-shaped platform for traces, evals, OpenInference, and

toolsarize-ai
8 Jun 2026
Safety

Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability

DGX agent

arXiv:2606.07437v1 Announce Type: cross Abstract: The ISO 26262 standard defines functional safety for road vehicles through risk assessments based on Severity, Exposure, and Controllability, grounded

safetyarxiv-cs-ai
8 Jun 2026
Model Releases

RealDocBench: A Benchmark for Field-Level QA and Layout Understanding on Real-World Regulated Documents

DGX agent

arXiv:2606.07401v1 Announce Type: new Abstract: Document parsing systems are increasingly deployed in high-stakes, regulated workflows such as mortgage underwriting, financial reporting, supply-chain

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

ScenicRules: An Autonomous Driving Benchmark with Multi-Objective Specifications and Abstract Scenarios

DGX agent

arXiv:2602.16073v2 Announce Type: replace-cross Abstract: Developing autonomous driving systems for complex traffic environments requires balancing multiple objectives, such as avoiding collisions, ob

model-releasesarxiv-cs-ai
8 Jun 2026
Applications

For some industries and dev orgs, this will happen already in 2026. For others, it may take until 2030. A short list of key differences for …

DGX agent

For some industries and dev orgs, this will happen already in 2026. For others, it may take until 2030. A short list of key differences for whether orgs will reach that level soon or later: • How much

applicationsitamar-friedman--x
7 Jun 2026
Tutorials

Super-powerful AI models will launch in the coming weeks. We are looking at a potential step change in model capabilities. The biggest mista…

DGX agent

Super-powerful AI models will launch in the coming weeks. We are looking at a potential step change in model capabilities. The biggest mistake right now is to lock into one vendor. I say this not only

tutorialsdair-ai--x
7 Jun 2026
Model Releases

Benchmarking Counterfactual Prediction in Epidemic Time Series with Time-Varying Interventions

DGX agent

arXiv:2606.05692v1 Announce Type: cross Abstract: Deep learning has enabled significant advances in time-series causal inference, yet progress remains constrained by the lack of realistic benchmarks w

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

DGX agent

arXiv:2606.05463v1 Announce Type: new Abstract: Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically perf

model-releasesarxiv-cs-ai
6 Jun 2026
Research

Stable Deep Reinforcement Learning via Isotropic Gaussian Representations

DGX agent

arXiv:2602.19373v3 Announce Type: replace-cross Abstract: Deep reinforcement learning systems often suffer from unstable training dynamics due to non-stationarity, where learning objectives and data d

researcharxiv-cs-ai
6 Jun 2026
Model Releases

Your margin is my opportunity: AI version… The biggest surprise of 2026 is that the capability gap between the best open-weight/source model…

DGX agent

Your margin is my opportunity: AI version… The biggest surprise of 2026 is that the capability gap between the best open-weight/source models and the best closed models has narrowed much faster than t

model-releasesclem-delangue--x
6 Jun 2026
Model Releases

Epidemiology of Model Collapse: Modeling Synthetic Data Contamination via Bilayer SIR Dynamics

DGX agent

arXiv:2606.05168v1 Announce Type: new Abstract: Training on synthetic data causes model collapse, but existing analyses treat this as single-chain degradation. In reality, the AI ecosystem involves cr

model-releasesarxiv-cs-cl
5 Jun 2026
Safety

EVE: A Generator-Verifier System for Generative Policies

DGX agent

arXiv:2512.21430v2 Announce Type: replace Abstract: Visuomotor policies based on generative such as diffusion and flow-matching have shown strong performance for robotics applications but degrade unde

safetyarxiv-cs-ro
5 Jun 2026
Safety

LadderMan: Learning Humanoid Perceptive Ladder Climbing

DGX agent

arXiv:2606.05873v1 Announce Type: cross Abstract: Humanoid robots hold great promise for operating in human-centered environments, yet ladder climbing remains one of the most challenging tasks due to

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

DGX agent

arXiv:2606.05744v1 Announce Type: new Abstract: Spatial planning maps are central to territorial governance, translating planning objectives, regulations, and spatial strategies into visual forms for

model-releasesarxiv-cs-cl
5 Jun 2026
Tools

'Reality: The Final Eval' — 現実タスク完了率こそが最終評価指標(@swyx / Andon Labs)。 複数のAI実装を並列で回していると、実感として正確だと思う。SWE-Benchの数字より「本番で動くか」が判断軸。エージェント設計で最初に決めるの…

DGX agent

# Reality: The Final Eval This post argues that real-world task completion rate is the ultimate metric for evaluating AI systems, prioritizing actual production performance over benchmark scores like

toolsswyx--x
5 Jun 2026
Model Releases

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)

DGX agent

arXiv:2606.05901v1 Announce Type: new Abstract: Large language models (LLMs) have fundamentally transformed the landscape of Natural Language Processing. Despite these advances, LLMs and LLM-based sys

model-releasesarxiv-cs-cl
5 Jun 2026
Safety

Robust Scene Transfer for PointGoal Navigation via Privileged Sensor Guided Contrastive Learning

DGX agent

arXiv:2606.05506v1 Announce Type: new Abstract: We propose a sensor-guided adaptive contrastive learning framework for visual representation learning in PointGoal navigation. During training, privileg

safetyarxiv-cs-cv
5 Jun 2026
Model Releases

Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation

DGX agent

arXiv:2606.05660v1 Announce Type: new Abstract: Embodied AI systems are increasingly expected to reason and act over extended horizons in physical environments. This growing capability brings safety t

model-releasesarxiv-cs-ro
5 Jun 2026
Model Releases

SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations

DGX agent

arXiv:2606.05563v1 Announce Type: cross Abstract: Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison

DGX agent

arXiv:2606.05436v1 Announce Type: cross Abstract: Summarizing the latest medical literature to guide clinical decision-making is essential for evidence-based medicine and high-quality patient care. Ye

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

The NVIDIA Nemotron Coalition continues to grow. We're excited to welcome new members: @hcompany_ai, @NousResearch, and @PrimeIntellect. And…

DGX agent

The NVIDIA Nemotron Coalition continues to grow. We're excited to welcome new members: @hcompany_ai, @NousResearch, and @PrimeIntellect. And a big thank you to our existing members: @bfl_ai, @cursor_a

model-releasesnous-research--x
5 Jun 2026
Model Releases

Affordance2Action: Task-Conditioned Scene-level Affordance Grounding for Real-Time Manipulation

DGX agent

arXiv:2606.04172v1 Announce Type: new Abstract: Task-conditioned manipulation requires grounding instructions to task-relevant functional parts rather than object categories. This setting is scene-dep

model-releasesarxiv-cs-ro
4 Jun 2026
Safety

AI enthusiasts are in a race against time, AI skeptics are in a race against entropy

DGX agent

AI enthusiasts are in a race against time, AI skeptics are in a race against entropy Charity Majors neatly captures the dynamic between AI enthusiasts and AI skeptics, both of whom are trying to build

safetysimon-willison
4 Jun 2026
Applications

As enterprise AI matures, data and context emerge as new competitive edge

DGX agent

Enterprise AI is entering a new phase where competitive advantage depends less on foundation models and more on the ability to connect data and business knowledge through an enterprise context layer.

applicationssiliconangle
4 Jun 2026
Research

Beyond Pixel Histories: World Models with Persistent 3D State

DGX agent

arXiv:2603.03482v2 Announce Type: replace-cross Abstract: Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities. However, e

researcharxiv-cs-ai
4 Jun 2026
Model Releases

Bilevel Autoresearch: Meta-Autoresearching Itself

DGX agent

arXiv:2603.23420v2 Announce Type: replace Abstract: If autoresearch is itself a form of research, then autoresearch can be applied to research itself. We present Bilevel Autoresearch, a bilevel framew

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

DiffAero: A GPU-Accelerated Differentiable Simulation Framework for Efficient Quadrotor Policy Learning

DGX agent

arXiv:2509.10247v1 Announce Type: cross Abstract: This letter introduces DiffAero, a lightweight, GPU-accelerated, and fully differentiable simulation framework designed for efficient quadrotor contro

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

DLO-Lab: Benchmarking Deformable Linear Object Manipulations with Differentiable Physics

DGX agent

arXiv:2606.04206v1 Announce Type: new Abstract: We address the challenge of enabling robots to manipulate deformable linear objects (DLOs), such as ropes, cables, and rubber bands. Prior work has prim

model-releasesarxiv-cs-ro
4 Jun 2026
Applications

Enterprise AI usage leaderboards are BAD and lead to the wrong behaviors. Employees feel pressure to hit higher token usage numbers without …

DGX agent

Enterprise AI usage leaderboards are BAD and lead to the wrong behaviors. Employees feel pressure to hit higher token usage numbers without any of the positive work transformation. I’ve heard directly

applicationsallie-k--miller--x
4 Jun 2026
Model Releases

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

DGX agent

arXiv:2606.04282v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are predominantly evaluated on free-form vision-language tasks such as visual question answering, captioning, a

model-releasesarxiv-cs-cv
4 Jun 2026
Model Releases

Highlighting recent advances in multi-GPU and tensor parallel support in llama.cpp Over the last few months llama.cpp maintainers and engine…

DGX agent

Highlighting recent advances in multi-GPU and tensor parallel support in llama.cpp Over the last few months llama.cpp maintainers and engineers from NVIDIA collaborated to improve the multi-GPU perfor

model-releasesgeorgi-gerganov--x
4 Jun 2026
← Previous
1…342343344345346…370
Next →