AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,607 results
Model Releases

Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test

DGX agent

arXiv:2605.22841v1 Announce Type: cross Abstract: What happens when the strongest alliance member pressures a weaker member over territory and strategic control? We examine the Greenland sovereignty c

model-releasesarxiv-cs-ai
25 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems

DGX agent

arXiv:2605.23024v1 Announce Type: new Abstract: Large language models now write software, draft legal documents, and produce clinical notes, yet fundamental limits, from Turing and Arrow to the No Fre

applicationsarxiv-cs-ai
25 May 2026
Tutorials

// The Efficiency Frontier in LLMs // (bookmark this one) How much are you overpaying for context you do not need? It turns out that context…

DGX agent

// The Efficiency Frontier in LLMs // (bookmark this one) How much are you overpaying for context you do not need? It turns out that context costs dominate production LLM bills, and the right strategy

tutorialsdair-ai--x
25 May 2026
Research

USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots

DGX agent

arXiv:2510.07869v4 Announce Type: replace Abstract: Underwater environments pose unique challenges for robotic navigation and manipulation. While existing research has primarily focused on task-specif

researcharxiv-cs-ro
25 May 2026
Safety

⚠️⚠️⚠️i don’t think most people understand the implications of the mood shift below, so I will spell them out. they are serious, and eventua…

DGX agent

⚠️⚠️⚠️i don’t think most people understand the implications of the mood shift below, so I will spell them out. they are serious, and eventually will affect the global economy. when a serious coder as

safetygary-marcus--x
24 May 2026
Safety

neurosymbolic by @swarat et al for the Erdos win, with much more careful, quantitative work than openai’s in hindsight i wonder whether Open…

DGX agent

neurosymbolic by @swarat et al for the Erdos win, with much more careful, quantitative work than openai’s in hindsight i wonder whether OpenAI rushed theirs out, knowing this was coming? Another 9 ope

safetygary-marcus--x
24 May 2026
Tools

Quoting Armin Ronacher

DGX agent

The most frustrating failure mode right now is that people submit issues that are not in their own voice. They contain an observed problem somewhere, but it has been thrown into a clanker and the clan

toolssimon-willison
24 May 2026
Applications

“We’re focused on providing the best in class document understanding and OCR technologies.” In partnership with @Wing_VC's 2026 Enterprise T…

DGX agent

“We’re focused on providing the best in class document understanding and OCR technologies.” In partnership with @Wing_VC's 2026 Enterprise Tech 30, @jerryjliu0, Co-Founder & CEO of @Llama_Index, visit

applicationsjerry-liu--x
24 May 2026
Safety

CCLab: Adversarial Testing of Learning- and Non-Learning-Based Congestion Controllers

DGX agent

arXiv:2605.21915v1 Announce Type: cross Abstract: Congestion controllers (CCs) are critical to network performance, and yet their robustness under adverse conditions remains insufficiently understood.

safetyarxiv-cs-lg
23 May 2026
Model Releases

Chebyshev Policies and the Mountain Car Problem: Reinforcement Learning for Low-Dimensional Control Tasks

DGX agent

arXiv:2605.22305v1 Announce Type: new Abstract: We analytically solve the Mountain Car problem, a canonical benchmark in RL, and derive an optimal control solution, closing a gap after 36 years. This

model-releasesarxiv-cs-lg
23 May 2026
Safety

Long-term Fairness with Selective Labels

DGX agent

arXiv:2605.22291v1 Announce Type: new Abstract: Long-term fairness algorithms aim to satisfy fairness beyond static and short-term notions by accounting for the dynamics between decision-making polici

safetyarxiv-cs-lg
23 May 2026
Safety

Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games

DGX agent

arXiv:2602.10894v2 Announce Type: replace Abstract: Two-player games such as board games have long been used as traditional benchmarks for reinforcement learning. This work revisits a policy optimizat

safetyarxiv-cs-lg
23 May 2026
Model Releases

Short-Term-to-Long-Term Memory Transfer for Knowledge Graphs under Partial Observability

DGX agent

arXiv:2605.22142v1 Announce Type: new Abstract: Reinforcement learning under partial observability requires deciding what information to retain, yet most memory-based approaches do not explicitly mode

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models

DGX agent

arXiv:2605.22732v1 Announce Type: cross Abstract: We investigate whether acoustic emotion recognition models can serve as proxies for the Pathos dimension in political speech analysis, as operationali

model-releasesarxiv-cs-cl
22 May 2026
Tutorials

Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks

DGX agent

arXiv:2601.23086v2 Announce Type: replace Abstract: Chain-of-thought (CoT) reasoning provides a significant performance uplift to LLMs by enabling planning, exploration, and deliberation of their acti

tutorialsarxiv-cs-ai
22 May 2026
Model Releases

ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning

DGX agent

arXiv:2605.22734v1 Announce Type: new Abstract: Biomedical knowledge graphs (KGs) treat disease associations as static facts, but temporal information is crucial for clinical reasoning, e.g., a sympto

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Closed and gated surveillance economy SaaS systems and models will get replaced by open source in two tiers: - The models themselves - The e…

DGX agent

Closed and gated surveillance economy SaaS systems and models will get replaced by open source in two tiers: - The models themselves - The everything-ultra-app harness If an American open source champ

model-releasesyann-lecun--x
22 May 2026
Research

Mind the Gaps: Multi-Robot Feedback-Driven Ergodic Coverage in Unknown Environments

DGX agent

arXiv:2605.21719v1 Announce Type: new Abstract: In this work, we address the problem of multi-robot adaptive coverage, where teams of robots perform dynamic sampling by continuously adjusting their po

researcharxiv-cs-ro
22 May 2026
Model Releases

OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind

DGX agent

arXiv:2605.20423v1 Announce Type: new Abstract: Large Language Models (LLMs) perform well on many language tasks, but their Theory of Mind (ToM) reasoning is still uneven in complex social settings. E

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

DGX agent

arXiv:2605.22109v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchm

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Putnam 2025 Problems in Rocq using Opus 4.6 and Rocq-MCP

DGX agent

arXiv:2603.20405v2 Announce Type: replace-cross Abstract: We report on an experiment in which Claude Opus~4.6, equipped with a suite of Model Context Protocol (MCP) tools for the Rocq proof assistant,

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Residual Skill Optimization for Text-to-SQL Ensembles

DGX agent

arXiv:2605.21792v1 Announce Type: new Abstract: Text-to-SQL ensembles improve over single-candidate generation by drawing multiple SQL candidates and selecting one, but their effectiveness is bounded

model-releasesarxiv-cs-cl
22 May 2026
Safety

ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving

DGX agent

arXiv:2605.21168v1 Announce Type: new Abstract: Safety-critical scenarios are central to evaluating autonomous driving systems, yet their rarity in naturalistic logs makes simulation-based stress test

safetyarxiv-cs-ai
22 May 2026
Safety

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning

DGX agent

arXiv:2506.14648v2 Announce Type: replace Abstract: Preference-based Reinforcement Learning (PbRL) methods provide a solution to avoid reward engineering by learning reward models based on human prefe

safetyarxiv-cs-ro
22 May 2026
Model Releases

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

DGX agent

arXiv:2605.22570v1 Announce Type: new Abstract: Spatio-temporal reasoning is a core capability for Multimodal Large Language Models (MLLMs) operating in the real world. As such, evaluating it precisel

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Causal Path Alignment: Anchoring the Optimization Trajectory for Controllable In-Parameter Knowledge Editing

DGX agent

arXiv:2506.04042v2 Announce Type: replace Abstract: Knowledge editing is pivotal for efficiently updating the parametric memory of Large Language Models (LLMs), enabling them to function as evolving a

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

FedCritic: Serverless Federated Critic Learning-based Resource Allocation for Multi-Cell OFDMA in 6G

DGX agent

arXiv:2605.21418v1 Announce Type: cross Abstract: In sixth-generation (6G) ultra-dense networks, aggressive frequency reuse amplifies inter-cell interference (ICI), making multi-cell orthogonal freque

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum

DGX agent

arXiv:2605.21133v1 Announce Type: new Abstract: In this paper, we explore spatial-aware humanoid whole-body manipulation task. Compared with tabletop settings, this task poses two key challenges: 1) S

model-releasesarxiv-cs-ro
21 May 2026
Model Releases

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling

DGX agent

arXiv:2508.08636v2 Announce Type: replace Abstract: Large language models (LLMs) have revolutionized artificial intelligence by enabling complex reasoning capabilities. While recent advancements in re

model-releasesarxiv-cs-cl
21 May 2026
Tutorials

Learn how to build LLM Wikis and LLM Artifacts.

DGX agent

Learn how to build LLM Wikis and LLM Artifacts. New VIDEO: From LLM Wikis to LLM Artifacts Shared all my thoughts on why LLM wikis and HTML artifacts are a big deal. Plus, new tools to help you build

tutorialsdair-ai--x
21 May 2026
Model Releases

MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks

DGX agent

arXiv:2605.20729v1 Announce Type: new Abstract: Accurate evaluation of conversational retrieval is pivotal for advancing Retrieval-Augmented Generation (RAG) systems. However, existing conversational

model-releasesarxiv-cs-cl
21 May 2026
Safety

Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity

DGX agent

arXiv:2605.20271v1 Announce Type: cross Abstract: We develop a rigorous statistical theory of multi-head attention (MHA) as an ensemble of Nadaraya-Watson (NW) kernel regression estimators. Building o

safetyarxiv-cs-lg
21 May 2026
Hardware

NVIDIA GTC Taipei at COMPUTEX: Live Updates on What’s Next in AI

DGX agent

At NVIDIA GTC Taipei at COMPUTEX, the world’s developers, researchers and industry leaders are converging to dive into the latest breakthroughs shaping every industry, covering topics spanning AI fact

hardwarenvidia-blog
21 May 2026
Model Releases

On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists

DGX agent

arXiv:2605.20668v1 Announce Type: new Abstract: With the advancement of AI capabilities, AI reviewers are beginning to be deployed in scientific peer review, yet their capability and credibility remai

model-releasesarxiv-cs-cl
21 May 2026
Safety

Q-SpiRL: Quantum Spiking Reinforcement Learning for Adaptive Robot Navigation

DGX agent

arXiv:2605.20801v1 Announce Type: new Abstract: Adaptive robot navigation in dynamic environments requires policies that can reach the target reliably while producing efficient and stable trajectories

safetyarxiv-cs-ro
21 May 2026
Research

Reinforcement Learning-based Control via Y-wise Affine Neural Networks: Comparative Case Studies for Chemical Processes

DGX agent

arXiv:2605.21211v1 Announce Type: cross Abstract: In this work we present an efficient and practically implementable approach for the application of reinforcement learning (RL)-based control in chemic

researcharxiv-cs-lg
21 May 2026
Safety

Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces

DGX agent

arXiv:2509.22963v3 Announce Type: replace Abstract: Reinforcement learning (RL) struggles to scale to large, combinatorial action spaces common in many real-world problems. This paper introduces a nov

safetyarxiv-cs-lg
21 May 2026
Model Releases

roto 2.0: The Robot Tactile Olympiad

DGX agent

arXiv:2605.21429v1 Announce Type: cross Abstract: Tactile-based reinforcement learning (RL) is currently hindered by fragmented research and a focus on over-saturated orientation tasks. We introduce v

model-releasesarxiv-cs-lg
21 May 2026
Safety

The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study

DGX agent

arXiv:2605.20767v1 Announce Type: new Abstract: Large language models (LLMs) show potential as simulators of human behavior, offering a scalable way to study responses to interventions. However, becau

safetyarxiv-cs-cl
21 May 2026
Applications

Validating Navmesh using Geometry: Voxel-Based Analysis with Prioritized Exploration

DGX agent

arXiv:2605.21397v1 Announce Type: cross Abstract: Navigation mesh (Navmesh) inconsistencies affect the player experience by directly impacting the navigation systems used by non-playable characters (N

applicationsarxiv-cs-ro
21 May 2026
Model Releases

Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls

DGX agent

arXiv:2605.19728v1 Announce Type: new Abstract: Foundation video models produce visually impressive results, but their use in embodied AI remains limited because they are primarily trained on natural

model-releasesarxiv-cs-cv
20 May 2026
Safety

Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay

DGX agent

arXiv:2605.19352v1 Announce Type: cross Abstract: Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the

safetyarxiv-cs-ai
20 May 2026
Tutorials

btw we did a bake off of Exa vs competitors and it took all of 1.5 hrs for the team to unanimously converge on exa lol. so proud to see my f…

DGX agent

btw we did a bake off of Exa vs competitors and it took all of 1.5 hrs for the team to unanimously converge on exa lol. so proud to see my former landlords crush it - time travel back to last year and

tutorialsswyx--x
20 May 2026
Model Releases

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data

DGX agent

arXiv:2605.19130v1 Announce Type: cross Abstract: Children acquire language grounding with remarkable robustness from limited visuo-linguistic input in ways that surpass today's best large multimodal

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models

DGX agent

arXiv:2605.18824v1 Announce Type: cross Abstract: Evaluation of foundation models often rely on aggregate scores from benchmarks that lack comprehensive coverage and metadata for a fine-grained evalua

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Got to play with a little of this before launch as well. My experience as a social scientist was that it was more bioscience focused right n…

DGX agent

Got to play with a little of this before launch as well. My experience as a social scientist was that it was more bioscience focused right now, but I think Google has been the leading lab in releasing

model-releasesethan-mollick--x
20 May 2026
Model Releases

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models

DGX agent

arXiv:2605.19341v1 Announce Type: cross Abstract: Hallucination remains a central failure mode of large language models, but existing benchmarks operationalize it inconsistently across summarization,

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

JAXenstein: Accelerated Benchmarking for First-Person Environments

DGX agent

arXiv:2605.19926v1 Announce Type: new Abstract: The progression of reinforcement learning algorithms have been driven by challenging benchmarks. The rate in which a researcher can iterate on a problem

model-releasesarxiv-cs-lg
20 May 2026
← Previous
1…345346347348349…367
Next →