AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,762 results
23 Jun 2026

Select-to-Act: Hierarchical Reinforcement Learning via Adaptive Language Guidance

Model ReleasesDGX agent

arXiv:2606.22350v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been widely applied to sequential decision-making, yet it often suffers from poor sample efficiency due to costly intera

Two-Bridge: Exclusive Objectives and Extended Horizon StarCraft II Benchmark

Model ReleasesDGX agent

arXiv:2603.06608v2 Announce Type: replace-cross Abstract: The research community lacks a middle ground between StarCraft II full game and its mini-games. The full-game's sprawling state-action space r

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.23543v1 Announce Type: cross Abstract: Scaling reinforcement learning for visual mathematical reasoning requires more than generating harder questions: as data volume grows, the reward labe

22 Jun 2026

Boost BigQuery with Python: Managed Python UDFs now generally available

Model ReleasesDGX agent

SQL is the industry standard for high-performance structured data analysis. However, expressing complex procedural logic, scientific computations, advanced string manipulations, or machine learning wo

Wiki Lint Report — 2026-06-22

SynthesesDGX agent

Automated lint: 48 errors, 13 warnings, 3 info

The advanced civilizations of sci-fi legend (Banks, Asimov, etc) have some form of simulation to guide society. @joon_s_pk is taking a crack…

TutorialsDGX agent

The advanced civilizations of sci-fi legend (Banks, Asimov, etc) have some form of simulation to guide society. @joon_s_pk is taking a crack at building that simulator with @simile_ai. As Joon's cofou

Use Case 1: Autonomous ML Research Can an AI autonomously improve another AI’s training recipe? We tasked Fugu Ultra with improving a small …

Model ReleasesDGX agent

Use Case 1: Autonomous ML Research Can an AI autonomously improve another AI’s training recipe? We tasked Fugu Ultra with improving a small GPT model using AutoResearch. Over 14 hours on a single H100

11 Jun 2026

From Content to Knowledge: Lightning Fast Long-Video Understanding with Neural Knowledge Representations

Model ReleasesDGX agent

arXiv:2606.11913v1 Announce Type: new Abstract: We propose a new paradigm for long video understanding by treating a long video as a Neural Knowledge Representation (NKR). NKR represents video content

MPC-Patch-Bench: Security-Aware LLM Code Patch for Multi-Party Computation

Model ReleasesDGX agent

arXiv:2606.11416v1 Announce Type: cross Abstract: Repository-level benchmarks for evaluating Large Language Model (LLM) code repair on Secure Multi-Party Computation (MPC) software do not yet exist, a

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

Model ReleasesDGX agent

arXiv:2606.11926v1 Announce Type: cross Abstract: Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction. Researchers test candidate directions, interpret the

10 Jun 2026

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness

Model ReleasesDGX agent

arXiv:2606.10577v1 Announce Type: new Abstract: Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs). Howe

Evaluating Research-Level Math Proofs via Strict Step-Level Verification

Model ReleasesDGX agent

arXiv:2606.10799v1 Announce Type: new Abstract: Large Language Models (LLMs) struggle to rigorously verify complex mathematical proofs. Standard global evaluation approaches suffer from 'context poiso

Event-Driven Reinforcement Learning Enables Long-Horizon Control in Semiconductor Fabrication

SafetyDGX agent

arXiv:2606.10705v1 Announce Type: cross Abstract: Reinforcement learning promises to optimize sequential decisions in large-scale systems. Semiconductor manufacturing systems are stochastic and highly

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses

Model ReleasesDGX agent

arXiv:2606.09878v1 Announce Type: new Abstract: Standard benchmarks report aggregate accuracy, but practitioners need to know which specific capabilities a model lacks. We introduce FailureScope, a be

Geometry-Aware Reinforcement Learning for 2D Irregular Nesting

Model ReleasesDGX agent

arXiv:2606.10611v1 Announce Type: cross Abstract: Traditional heuristic solvers for the 2D irregular nesting problem share a fundamental limitation: they are blind to polygon geometry, relying on guid

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake

Model ReleasesDGX agent

arXiv:2606.10460v1 Announce Type: cross Abstract: Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can b

Multi-task LLMs for Bug Classification: Efficient Inference with Auxiliary Decoding Heads

Local AiDGX agent

arXiv:2606.09956v1 Announce Type: cross Abstract: The rapid adoption of LLM-powered code generation has dramatically accelerated software development, yet effective verification methods remain severel

Proud to have Pinecone Nexus be a part of this along with our friends at @LangChain, @tavilyai, and @guardrails_ai!

ApplicationsDGX agent

Proud to have Pinecone Nexus be a part of this along with our friends at @LangChain, @tavilyai, and @guardrails_ai! Most #AIAgents don't fail because of the model. They fail because of the infrastruct

Sim2Schedule: A Simulator-Guided LLM Framework for Autonomous Open-Pit Mine Scheduling

Model ReleasesDGX agent

arXiv:2606.10286v1 Announce Type: new Abstract: Open-pit mine scheduling is a critical process for maximizing economic return under complex geotechnical and operational constraints. While Mixed-Intege

9 Jun 2026

Auditable Graph-Guided Root Cause Analysis for Kubernetes Incidents

Model ReleasesDGX agent

arXiv:2606.08590v1 Announce Type: cross Abstract: Kubernetes incidents are diagnosed reliably only when a root-cause system's reported gains come from incident evidence rather than scenario-specific s

ComplexConstraints and Beyond: Expert Rubrics for RLVR

Model ReleasesDGX agent

arXiv:2606.09118v1 Announce Type: new Abstract: As LLM capabilities advance rapidly, the evaluation methods used to assess them increasingly lag behind. Traditional benchmarks relied on programmatic v

Decentralized End-to-End Multi-AAV Pursuit Using Predictive Spatio-Temporal Observation via Deep Reinforcement Learning

SafetyDGX agent

arXiv:2603.24238v2 Announce Type: replace Abstract: Decentralized cooperative pursuit in cluttered environments is challenging for autonomous aerial swarms, especially under partial and noisy percepti

Decoupling Semantics and Logic: A Training-Free Coarse-to-Fine Pipeline for Video Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2606.07924v1 Announce Type: cross Abstract: This paper presents our system description for the 2nd Workshop on Multimodal Augmented Generation via MultimodAl Retrieval (MAGMaR). Addressing the c

Distilling LLM Reasoning into an Interpretable Policy Tree for Human-AI Collaboration

SafetyDGX agent

arXiv:2606.08596v1 Announce Type: new Abstract: Constructing efficient and reliable policies to assist humans is indispensable for human-AI collaboration. Existing methods mainly follow two lines of w

Dynamic Distributed Constraint Optimization and Metareasoning for Continual, Large-Scale Satellite Operations

Local AiDGX agent

arXiv:2601.06188v3 Announce Type: replace Abstract: As Earth-observing satellite constellations grow in size and capability, distributed onboard control offers a pathway to novel responses and time-se

End-to-End Context Compression at Scale

Model ReleasesDGX agent

arXiv:2606.09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache

Implementing Grassroots Logic Programs with Multiagent Transition Systems and AI (Full Version)

Model ReleasesDGX agent

arXiv:2602.06934v4 Announce Type: replace-cross Abstract: Grassroots Logic Programs (GLP) is a concurrent logic programming language in which logic variables are partitioned into paired readers and wr

Instrumental convergence and power-seeking

SafetyDGX agent

arXiv:2606.08832v1 Announce Type: new Abstract: Recent years have seen increasing concern that artificial intelligence may soon pose an existential risk to humanity. One leading ground for concern is

PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation

Model ReleasesDGX agent

arXiv:2510.20182v2 Announce Type: replace Abstract: Pedestrian simulation traditionally relies on expert-tuned, hand-crafted models that limit scalability and generalization. Meanwhile, large-scale vi

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

Model ReleasesDGX agent

arXiv:2606.09038v1 Announce Type: new Abstract: Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. H

PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher Systems

Model ReleasesDGX agent

arXiv:2606.08481v1 Announce Type: cross Abstract: Enterprise property graphs vary widely in schema structure, internal terminology, domain assumptions, governance constraints, and user interaction pat

Systematic LLM Translation of Legacy Scientific Code to Differentiable Frameworks: Application to a Land Surface Model

Model ReleasesDGX agent

arXiv:2606.07681v1 Announce Type: cross Abstract: Differentiable programming offers transformative capabilities for scientific modeling, enabling gradient-based parameter estimation, sensitivity analy

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

Model ReleasesDGX agent

arXiv:2606.07723v1 Announce Type: new Abstract: Open-vocabulary long-horizon manipulation requires robots to reason over flexible instructions and complex multi-object scenes while adaptively planning

8 Jun 2026

Learn to Match: Two-Sided Matching with Temporally Extended Feedback

Model ReleasesDGX agent

arXiv:2606.06744v1 Announce Type: new Abstract: Two-sided matching markets often involve information that unfolds over time through interviews, repeated interaction, learning, and separation. Existing

Modernizing Healthcare: How Alcidion achieved greater stability and performance with AlloyDB

Model ReleasesDGX agent

In clinical informatics, every second counts. For Alcidion, a global leader in smart health solutions, the mission is simple but critical: use technology to reduce cognitive load for clinicians and pr

Seeing a number of benchmarks showing Opus is the best model for long-running work. Five tips for running Opus autonomously for hours/days: …

Model ReleasesDGX agent

Seeing a number of benchmarks showing Opus is the best model for long-running work. Five tips for running Opus autonomously for hours/days: 1. Use auto mode for permissions, so Claude doesn’t ask for

The point is that you should start implementing ways to encode instructions/prompts with clear goals inside automations. Nothing new but new…

ResearchDGX agent

The point is that you should start implementing ways to encode instructions/prompts with clear goals inside automations. Nothing new but newer LLMs are being trained to perform for longer duration uni

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.06673v1 Announce Type: new Abstract: Sparse rewards and heterogeneous task sequences remain persistent challenges in Reinforcement Learning (RL), often resulting in slow convergence, weak g

7 Jun 2026

Wiki Lint Report — 2026-06-07

SynthesesDGX agent

Automated lint: 47 errors, 12 warnings, 3 info

This is so true. What people fail to realize is that when new technology tools become available, they tend to be used to create new capabili…

IndustryDGX agent

This is so true. What people fail to realize is that when new technology tools become available, they tend to be used to create new capabilities up the stack, while down the stack, archaeological laye

6 Jun 2026

Learning Adaptive Parallel Execution for Efficient Code Localization

ResearchDGX agent

arXiv:2601.19568v2 Announce Type: replace Abstract: Code localization constitutes a key bottleneck in automated software development pipelines. While concurrent tool execution can enhance discovery sp

WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation

Model ReleasesDGX agent

arXiv:2606.06147v1 Announce Type: new Abstract: End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation. However, existing approaches typically rely on historical observati

5 Jun 2026

CLFEC: A New Task for Unified Linguistic and Factual Error Correction in paragraph-level Chinese Professional Writing

Model ReleasesDGX agent

arXiv:2602.23845v2 Announce Type: replace Abstract: Chinese text correction has traditionally focused on spelling and grammar, while factual error correction is usually treated separately. However, in

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments

Model ReleasesDGX agent

arXiv:2606.05661v1 Announce Type: cross Abstract: Continual learning, the ability of AI systems to improve through sequential experience, has attracted substantial interest, but no high-quality benchm

I absolutely agree that there really is this 10x opportunity for companies to be $40 trillion in market cap and beyond—perhaps Nvidia, Googl…

HardwareDGX agent

I absolutely agree that there really is this 10x opportunity for companies to be 40 trillion in market cap and beyond—perhaps Nvidia, Google, and beyond. Really fascinating to consider what that could

Seeking Counsel: Ongoing Targeted Campaign Against US Law Firms

SafetyDGX agent

Written by: Chad Reams, Tufail Ahmed, Keith Knapp, Ashley Frazer, Tyler McLellan Introduction From January through May 2026, Mandiant identified a financially motivated data theft extortion campaign e

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset

Model ReleasesDGX agent

arXiv:2606.06338v1 Announce Type: new Abstract: Video question answering (VideoQA) aims to answer questions about given videos. While existing approaches excel on factoid VideoQA, they struggle with d

We've made a breakthrough in self-evolving AI scientists moving from 'search' to 'principled discovery': Scientific discovery requires that …

Model ReleasesDGX agent

We've made a breakthrough in self-evolving AI scientists moving from 'search' to 'principled discovery': Scientific discovery requires that the search space itself changes, and an AI scientist must pe

4 Jun 2026

Andon Labs' Real-World AI Evals: Claude calls the FBI, AI CEOs, price cartels, Butter-Bench, & Luna https://latent.space/p/andon @andonlabs …

Model ReleasesDGX agent

Andon Labs' Real-World AI Evals: Claude calls the FBI, AI CEOs, price cartels, Butter-Bench, & Luna https://latent.space/p/andon @andonlabs cofounders @lukaspet and @axelbacklund explain why dollar-de

Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments

SafetyDGX agent

arXiv:2209.15448v3 Announce Type: replace Abstract: As AI becomes more prevalent throughout society, effective methods of integrating humans and AI systems that leverage their respective strengths and

Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning

SafetyDGX agent

arXiv:2606.04574v1 Announce Type: new Abstract: This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in

@nvidia @nebiustf More info on the new Nemotron 3 Ultra: https://x.com/NVIDIAAI/status/2062521325076299981?s=20

Model ReleasesDGX agent

@nvidia @nebiustf More info on the new Nemotron 3 Ultra: https://x.com/NVIDIAAI/status/2062521325076299981?s=20 Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built

Quantum entanglement provides a competitive advantage in adversarial games

Model ReleasesDGX agent

arXiv:2603.10289v2 Announce Type: replace-cross Abstract: Whether uniquely quantum resources confer advantages in fully classical, competitive environments remains an open question. Competitive zero-s

Simulate, Reason, Decide: Scientific Reasoning with LLMs for Simulation-Driven Decision Making

ResearchDGX agent

arXiv:2606.04505v1 Announce Type: new Abstract: Scientific simulators are increasingly being integrated into LLM-driven systems for high-stakes simulation-driven decision-making. However, existing fra

3 Jun 2026

Automated Report-Derived Oncology VQA Benchmark for Evaluating Vision-Language Models on 3D Medical Imaging

Model ReleasesDGX agent

arXiv:2606.02809v1 Announce Type: new Abstract: Evaluating vision-language models (VLMs) on medical images requires benchmarks that are clinically grounded, scalable, and controlled for evaluation con

ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models

Model ReleasesDGX agent

arXiv:2606.03157v1 Announce Type: new Abstract: Large language models (LLMs) have been widely adopted in healthcare, yet they still encounter significant challenges in complex clinical decision-making

Discovering autonomous quantum error correction via deep reinforcement learning

SafetyDGX agent

arXiv:2511.12482v2 Announce Type: replace-cross Abstract: Quantum error correction is essential for fault-tolerant quantum computing. However, standard methods relying on active measurements may intro

Human-in-the-Loop Contextual Bandits for Short-Term Rental Dynamic Pricing: Structural Equivalence of Historical Warm-Up and Approval-Gated Live Learning

SafetyDGX agent

arXiv:2606.02595v1 Announce Type: new Abstract: Dynamic pricing in short-term rental (STR) markets presents a distinctive challenge for online learning algorithms: pricing decisions carry significant

Inference Cost Attacks for Retrieval-Augmented Large Language Models

SafetyDGX agent

arXiv:2606.02643v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG)-enhanced LLM systems, while powerful, introduce substantial inference costs due to the inclusion of an extra mult

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

Model ReleasesDGX agent

arXiv:2606.02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-regis

← Previous
1…246247248249250…297
Next →