AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
Tutorials

The advanced civilizations of sci-fi legend (Banks, Asimov, etc) have some form of simulation to guide society. @joon_s_pk is taking a crack…

DGX agent

The advanced civilizations of sci-fi legend (Banks, Asimov, etc) have some form of simulation to guide society. @joon_s_pk is taking a crack at building that simulator with @simile_ai. As Joon's cofou

tutorialssonya-huang--x
22 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Use Case 1: Autonomous ML Research Can an AI autonomously improve another AI’s training recipe? We tasked Fugu Ultra with improving a small …

DGX agent

Use Case 1: Autonomous ML Research Can an AI autonomously improve another AI’s training recipe? We tasked Fugu Ultra with improving a small GPT model using AutoResearch. Over 14 hours on a single H100

model-releasesdavid-ha--x
22 Jun 2026
Model Releases

From Content to Knowledge: Lightning Fast Long-Video Understanding with Neural Knowledge Representations

DGX agent

arXiv:2606.11913v1 Announce Type: new Abstract: We propose a new paradigm for long video understanding by treating a long video as a Neural Knowledge Representation (NKR). NKR represents video content

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

MPC-Patch-Bench: Security-Aware LLM Code Patch for Multi-Party Computation

DGX agent

arXiv:2606.11416v1 Announce Type: cross Abstract: Repository-level benchmarks for evaluating Large Language Model (LLM) code repair on Secure Multi-Party Computation (MPC) software do not yet exist, a

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

DGX agent

arXiv:2606.11926v1 Announce Type: cross Abstract: Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction. Researchers test candidate directions, interpret the

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness

DGX agent

arXiv:2606.10577v1 Announce Type: new Abstract: Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs). Howe

model-releasesarxiv-cs-ro
10 Jun 2026
Model Releases

Evaluating Research-Level Math Proofs via Strict Step-Level Verification

DGX agent

arXiv:2606.10799v1 Announce Type: new Abstract: Large Language Models (LLMs) struggle to rigorously verify complex mathematical proofs. Standard global evaluation approaches suffer from 'context poiso

model-releasesarxiv-cs-ai
10 Jun 2026
Safety

Event-Driven Reinforcement Learning Enables Long-Horizon Control in Semiconductor Fabrication

DGX agent

arXiv:2606.10705v1 Announce Type: cross Abstract: Reinforcement learning promises to optimize sequential decisions in large-scale systems. Semiconductor manufacturing systems are stochastic and highly

safetyarxiv-cs-ai
10 Jun 2026
Model Releases

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses

DGX agent

arXiv:2606.09878v1 Announce Type: new Abstract: Standard benchmarks report aggregate accuracy, but practitioners need to know which specific capabilities a model lacks. We introduce FailureScope, a be

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

Geometry-Aware Reinforcement Learning for 2D Irregular Nesting

DGX agent

arXiv:2606.10611v1 Announce Type: cross Abstract: Traditional heuristic solvers for the 2D irregular nesting problem share a fundamental limitation: they are blind to polygon geometry, relying on guid

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake

DGX agent

arXiv:2606.10460v1 Announce Type: cross Abstract: Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can b

model-releasesarxiv-cs-ai
10 Jun 2026
Local Ai

Multi-task LLMs for Bug Classification: Efficient Inference with Auxiliary Decoding Heads

DGX agent

arXiv:2606.09956v1 Announce Type: cross Abstract: The rapid adoption of LLM-powered code generation has dramatically accelerated software development, yet effective verification methods remain severel

local-aiarxiv-cs-lg
10 Jun 2026
Applications

Proud to have Pinecone Nexus be a part of this along with our friends at @LangChain, @tavilyai, and @guardrails_ai!

DGX agent

Proud to have Pinecone Nexus be a part of this along with our friends at @LangChain, @tavilyai, and @guardrails_ai! Most #AIAgents don't fail because of the model. They fail because of the infrastruct

applicationspinecone--x
10 Jun 2026
Model Releases

Sim2Schedule: A Simulator-Guided LLM Framework for Autonomous Open-Pit Mine Scheduling

DGX agent

arXiv:2606.10286v1 Announce Type: new Abstract: Open-pit mine scheduling is a critical process for maximizing economic return under complex geotechnical and operational constraints. While Mixed-Intege

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Auditable Graph-Guided Root Cause Analysis for Kubernetes Incidents

DGX agent

arXiv:2606.08590v1 Announce Type: cross Abstract: Kubernetes incidents are diagnosed reliably only when a root-cause system's reported gains come from incident evidence rather than scenario-specific s

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

ComplexConstraints and Beyond: Expert Rubrics for RLVR

DGX agent

arXiv:2606.09118v1 Announce Type: new Abstract: As LLM capabilities advance rapidly, the evaluation methods used to assess them increasingly lag behind. Traditional benchmarks relied on programmatic v

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Decentralized End-to-End Multi-AAV Pursuit Using Predictive Spatio-Temporal Observation via Deep Reinforcement Learning

DGX agent

arXiv:2603.24238v2 Announce Type: replace Abstract: Decentralized cooperative pursuit in cluttered environments is challenging for autonomous aerial swarms, especially under partial and noisy percepti

safetyarxiv-cs-ro
9 Jun 2026
Safety

Decoupling Semantics and Logic: A Training-Free Coarse-to-Fine Pipeline for Video Retrieval-Augmented Generation

DGX agent

arXiv:2606.07924v1 Announce Type: cross Abstract: This paper presents our system description for the 2nd Workshop on Multimodal Augmented Generation via MultimodAl Retrieval (MAGMaR). Addressing the c

safetyarxiv-cs-ai
9 Jun 2026
Safety

Distilling LLM Reasoning into an Interpretable Policy Tree for Human-AI Collaboration

DGX agent

arXiv:2606.08596v1 Announce Type: new Abstract: Constructing efficient and reliable policies to assist humans is indispensable for human-AI collaboration. Existing methods mainly follow two lines of w

safetyarxiv-cs-ai
9 Jun 2026
Local Ai

Dynamic Distributed Constraint Optimization and Metareasoning for Continual, Large-Scale Satellite Operations

DGX agent

arXiv:2601.06188v3 Announce Type: replace Abstract: As Earth-observing satellite constellations grow in size and capability, distributed onboard control offers a pathway to novel responses and time-se

local-aiarxiv-cs-ai
9 Jun 2026
Model Releases

End-to-End Context Compression at Scale

DGX agent

arXiv:2606.09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Implementing Grassroots Logic Programs with Multiagent Transition Systems and AI (Full Version)

DGX agent

arXiv:2602.06934v4 Announce Type: replace-cross Abstract: Grassroots Logic Programs (GLP) is a concurrent logic programming language in which logic variables are partitioned into paired readers and wr

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Instrumental convergence and power-seeking

DGX agent

arXiv:2606.08832v1 Announce Type: new Abstract: Recent years have seen increasing concern that artificial intelligence may soon pose an existential risk to humanity. One leading ground for concern is

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation

DGX agent

arXiv:2510.20182v2 Announce Type: replace Abstract: Pedestrian simulation traditionally relies on expert-tuned, hand-crafted models that limit scalability and generalization. Meanwhile, large-scale vi

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

DGX agent

arXiv:2606.09038v1 Announce Type: new Abstract: Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. H

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher Systems

DGX agent

arXiv:2606.08481v1 Announce Type: cross Abstract: Enterprise property graphs vary widely in schema structure, internal terminology, domain assumptions, governance constraints, and user interaction pat

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Systematic LLM Translation of Legacy Scientific Code to Differentiable Frameworks: Application to a Land Surface Model

DGX agent

arXiv:2606.07681v1 Announce Type: cross Abstract: Differentiable programming offers transformative capabilities for scientific modeling, enabling gradient-based parameter estimation, sensitivity analy

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

DGX agent

arXiv:2606.07723v1 Announce Type: new Abstract: Open-vocabulary long-horizon manipulation requires robots to reason over flexible instructions and complex multi-object scenes while adaptively planning

model-releasesarxiv-cs-ro
9 Jun 2026
Model Releases

Learn to Match: Two-Sided Matching with Temporally Extended Feedback

DGX agent

arXiv:2606.06744v1 Announce Type: new Abstract: Two-sided matching markets often involve information that unfolds over time through interviews, repeated interaction, learning, and separation. Existing

model-releasesarxiv-cs-lg
8 Jun 2026
Model Releases

Modernizing Healthcare: How Alcidion achieved greater stability and performance with AlloyDB

DGX agent

In clinical informatics, every second counts. For Alcidion, a global leader in smart health solutions, the mission is simple but critical: use technology to reduce cognitive load for clinicians and pr

model-releasesgoogle-cloud-ai
8 Jun 2026
Model Releases

Seeing a number of benchmarks showing Opus is the best model for long-running work. Five tips for running Opus autonomously for hours/days: …

DGX agent

Seeing a number of benchmarks showing Opus is the best model for long-running work. Five tips for running Opus autonomously for hours/days: 1. Use auto mode for permissions, so Claude doesn’t ask for

model-releasesboris-cherny--x
8 Jun 2026
Research

The point is that you should start implementing ways to encode instructions/prompts with clear goals inside automations. Nothing new but new…

DGX agent

The point is that you should start implementing ways to encode instructions/prompts with clear goals inside automations. Nothing new but newer LLMs are being trained to perform for longer duration uni

researchdair-ai--x
8 Jun 2026
Model Releases

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning

DGX agent

arXiv:2606.06673v1 Announce Type: new Abstract: Sparse rewards and heterogeneous task sequences remain persistent challenges in Reinforcement Learning (RL), often resulting in slow convergence, weak g

model-releasesarxiv-cs-lg
8 Jun 2026
Syntheses

Wiki Lint Report — 2026-06-07

DGX agent

Automated lint: 47 errors, 12 warnings, 3 info

linthealth-checkautomated
7 Jun 2026
Industry

This is so true. What people fail to realize is that when new technology tools become available, they tend to be used to create new capabili…

DGX agent

This is so true. What people fail to realize is that when new technology tools become available, they tend to be used to create new capabilities up the stack, while down the stack, archaeological laye

industryclem-delangue--x
7 Jun 2026
Research

Learning Adaptive Parallel Execution for Efficient Code Localization

DGX agent

arXiv:2601.19568v2 Announce Type: replace Abstract: Code localization constitutes a key bottleneck in automated software development pipelines. While concurrent tool execution can enhance discovery sp

researcharxiv-cs-ai
6 Jun 2026
Model Releases

WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation

DGX agent

arXiv:2606.06147v1 Announce Type: new Abstract: End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation. However, existing approaches typically rely on historical observati

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

CLFEC: A New Task for Unified Linguistic and Factual Error Correction in paragraph-level Chinese Professional Writing

DGX agent

arXiv:2602.23845v2 Announce Type: replace Abstract: Chinese text correction has traditionally focused on spelling and grammar, while factual error correction is usually treated separately. However, in

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments

DGX agent

arXiv:2606.05661v1 Announce Type: cross Abstract: Continual learning, the ability of AI systems to improve through sequential experience, has attracted substantial interest, but no high-quality benchm

model-releasesarxiv-cs-cl
5 Jun 2026
Hardware

I absolutely agree that there really is this 10x opportunity for companies to be $40 trillion in market cap and beyond—perhaps Nvidia, Googl…

DGX agent

I absolutely agree that there really is this 10x opportunity for companies to be 40 trillion in market cap and beyond—perhaps Nvidia, Google, and beyond. Really fascinating to consider what that could

hardwareswyx--x
5 Jun 2026
Safety

Seeking Counsel: Ongoing Targeted Campaign Against US Law Firms

DGX agent

Written by: Chad Reams, Tufail Ahmed, Keith Knapp, Ashley Frazer, Tyler McLellan Introduction From January through May 2026, Mandiant identified a financially motivated data theft extortion campaign e

safetygoogle-cloud-ai
5 Jun 2026
Model Releases

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset

DGX agent

arXiv:2606.06338v1 Announce Type: new Abstract: Video question answering (VideoQA) aims to answer questions about given videos. While existing approaches excel on factoid VideoQA, they struggle with d

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

We've made a breakthrough in self-evolving AI scientists moving from 'search' to 'principled discovery': Scientific discovery requires that …

DGX agent

We've made a breakthrough in self-evolving AI scientists moving from 'search' to 'principled discovery': Scientific discovery requires that the search space itself changes, and an AI scientist must pe

model-releasesgary-marcus--x
5 Jun 2026
Model Releases

Andon Labs' Real-World AI Evals: Claude calls the FBI, AI CEOs, price cartels, Butter-Bench, & Luna https://latent.space/p/andon @andonlabs …

DGX agent

Andon Labs' Real-World AI Evals: Claude calls the FBI, AI CEOs, price cartels, Butter-Bench, & Luna https://latent.space/p/andon @andonlabs cofounders @lukaspet and @axelbacklund explain why dollar-de

model-releasesswyx--x
4 Jun 2026
Safety

Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments

DGX agent

arXiv:2209.15448v3 Announce Type: replace Abstract: As AI becomes more prevalent throughout society, effective methods of integrating humans and AI systems that leverage their respective strengths and

safetyarxiv-cs-lg
4 Jun 2026
Safety

Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning

DGX agent

arXiv:2606.04574v1 Announce Type: new Abstract: This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in

safetyarxiv-cs-lg
4 Jun 2026
Model Releases

@nvidia @nebiustf More info on the new Nemotron 3 Ultra: https://x.com/NVIDIAAI/status/2062521325076299981?s=20

DGX agent

@nvidia @nebiustf More info on the new Nemotron 3 Ultra: https://x.com/NVIDIAAI/status/2062521325076299981?s=20 Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built

model-releasesnous-research--x
4 Jun 2026
Model Releases

Quantum entanglement provides a competitive advantage in adversarial games

DGX agent

arXiv:2603.10289v2 Announce Type: replace-cross Abstract: Whether uniquely quantum resources confer advantages in fully classical, competitive environments remains an open question. Competitive zero-s

model-releasesarxiv-cs-ai
4 Jun 2026
← Previous
1…308309310311312…371
Next →