AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
29 May 2026

Different models need different prompts, sometimes tools “Harness profiles” are how we do that in deepagents

ApplicationsDGX agent

Different models need different prompts, sometimes tools “Harness profiles” are how we do that in deepagents Deep Agents v0.6 makes harness profiles a first-class abstraction. Now, you can get product

Differentiable Belief-based Opponent Shaping

Model ReleasesDGX agent

arXiv:2605.29042v1 Announce Type: new Abstract: Human coordination often relies on the ability to influence the beliefs of others through strategic action. In multi-agent reinforcement learning, oppon

Dynamic Mixture of Progressive Parameter-Efficient Expert Library for Lifelong Robot Learning

Model ReleasesDGX agent

arXiv:2506.05985v3 Announce Type: replace Abstract: A generalist agent must continuously learn and adapt throughout its lifetime, achieving efficient forward transfer while minimizing catastrophic for

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

FedQHD: Closed-Form Function-Space Federated Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.29002v1 Announce Type: new Abstract: Federated reinforcement learning enables decentralized agents to collaboratively improve policies or value estimates without exchanging raw trajectories

PokerSkill: LLMs Can Play Expert-Level Poker without Training or Solvers

Model ReleasesDGX agent

arXiv:2605.30094v1 Announce Type: new Abstract: Poker is a landmark challenge for artificial intelligence. The dominant approach relies on equilibrium solvers built on counterfactual regret minimizati

Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison

Model ReleasesDGX agent

arXiv:2605.30087v1 Announce Type: new Abstract: Emerging personal AI agents are moving toward persistent, multi-source memory. This creates an evaluation problem: systems must decide how to use confli

SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?

Model ReleasesDGX agent

arXiv:2605.30329v1 Announce Type: new Abstract: Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. How

The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF

Model ReleasesDGX agent

arXiv:2605.29491v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in agentic and retrieval-augmented generation (RAG) systems, where they must execute user-specifi

Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning

SafetyDGX agent

arXiv:2605.29032v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably

There are still a few spots left for @hwchase17 + @traversal_ai founder @_anish_agarwal's technical fireside chat in NYC on June 2nd. Learn …

TutorialsDGX agent

There are still a few spots left for @hwchase17 + @traversal_ai founder @_anish_agarwal's technical fireside chat in NYC on June 2nd. Learn how Traversal builds, ships and improves their agents. Enjoy

28 May 2026

COTTA: Context-Aware Transfer Adaptation for Trajectory Prediction in Autonomous Driving

SafetyDGX agent

arXiv:2604.00402v2 Announce Type: replace-cross Abstract: Developing robust models to accurately predict the trajectories of surrounding agents is fundamental to autonomous driving safety. However, mo

Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation

Model ReleasesDGX agent

arXiv:2605.28587v1 Announce Type: new Abstract: Understanding dynamic 3D environments is essential for safe autonomous driving, particularly when reasoning about human-centric, nonrigid agents. Howeve

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning

Model ReleasesDGX agent

arXiv:2510.27266v2 Announce Type: replace Abstract: Autonomous graphical user interface (GUI) agents rely on accurate GUI grounding, which maps language instructions to on-screen coordinates, to execu

Fed up with vibe coders, dev sneaks data-nuking prompt injection into their code

IndustryDGX agent

A developer added hidden prompt injection instructions to jqwik, a Java testing app, to sabotage projects created by AI coding agents in response to the 'vibe coding' controversy. The incident represe

Learning When to Optimize: Verified Optimization Skills from Expert GPU-Kernel Lineages

HardwareDGX agent

arXiv:2605.28213v1 Announce Type: new Abstract: LLM-based agents are increasingly used to generate GPU kernels, but they often know what optimizations to try without knowing when those optimizations a

On Compositional Learning Behaviours in Formal Mathematics

Model ReleasesDGX agent

arXiv:2605.28512v1 Announce Type: new Abstract: Self-evolving scientific agents capable of conquering the hard tail of formal mathematics require Compositional Learning Behaviours (CLBs) -- the capaci

Poison with Style: A Practical Poisoning Attack on Code Large Language Models

ResearchDGX agent

arXiv:2605.27631v1 Announce Type: cross Abstract: Code Large Language Models (CLLMs) serve as the core of modern code agents, enabling developers to automate complex software development tasks. In thi

Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement

Model ReleasesDGX agent

arXiv:2605.28360v1 Announce Type: new Abstract: Automatic prompt optimization (APO) has driven significant gains in LLM-based agentic workflows. However, existing methods treat each task's prompt as a

Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real Deployment

SafetyDGX agent

arXiv:2605.27659v1 Announce Type: cross Abstract: Due to limited resources and public safety concerns, deep reinforcement learning (RL) agents for many cyber-physical systems (e.g., autonomous vehicle

We're selectively releasing the Paris 2.0 weights and partnering with researchers and teams interested in diffusion-based video models, worl…

IndustryDGX agent

We're selectively releasing the Paris 2.0 weights and partnering with researchers and teams interested in diffusion-based video models, world models, and embodied agents. The model is on Hugging Face:

27 May 2026

Balancing Plasticity and Stability with Fast and Slow Successor Features

ApplicationsDGX agent

arXiv:2605.26357v1 Announce Type: new Abstract: A hallmark of intelligence is the ability to adapt in non-stationary environments, yet deep Reinforcement Learning (RL) agents often struggle in such se

Cogent Security launches autonomous vulnerability response tools as AI-assisted exploits outpace scanners

Model ReleasesDGX agent

Cogent Security Inc., a startup that employs agentic artificial intelligence for vulnerability management, today launched two new platform capabilities aimed at compressing enterprise vulnerability re

E^3C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control

ResearchDGX agent

arXiv:2605.26316v1 Announce Type: cross Abstract: Controllable and physically grounded egocentric video generation is essential for embodied agents to reason about how their own and others' actions ma

From Static Context to Calibrated Interactive RL: Mitigating Distribution Shift in Multi-turn Dialogue with Aligned Simulator

SafetyDGX agent

arXiv:2605.26403v1 Announce Type: new Abstract: A long-standing goal of the research community is to develop highly interactive LLM-based dialogue agents. Recent research focuses on optimizing policie

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams

Model ReleasesDGX agent

arXiv:2605.27074v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) achieve strong performance on reactive question answering, but real-world streaming assistants require p

MemFail: Stress-Testing Failure Modes of LLM Memory Systems

Model ReleasesDGX agent

arXiv:2605.26667v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external memory systems to remain consistent across long-horizon interactions, but little empiric

MerLean-Prover: A Recursive Looping Harness for End-to-End Lean 4 Theorem Proving

Model ReleasesDGX agent

arXiv:2605.26959v1 Announce Type: cross Abstract: MerLean-Prover is an end-to-end Lean4 theorem prover that replaces sorry declarations with kernel-checkable proofs. It is built from three agent types

Modernising Reinforcement Learning-Based Navigation for Embodied Semantic Scene Graph Generation

SafetyDGX agent

arXiv:2603.25415v2 Announce Type: replace Abstract: Semantic world models enable embodied agents to reason about objects, relations, and spatial context beyond purely geometric representations. In Org

Position: AI Safety Requires Effective Controllability

Model ReleasesDGX agent

arXiv:2605.27117v1 Announce Type: new Abstract: AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing ha

ReasonOps: A Unified Operational Paradigm for Trustworthy Verified LLM Reasoning

SafetyDGX agent

arXiv:2605.27014v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed artificial intelligence from primarily generative systems into increasingly capable reasoning agents. Re

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

Model ReleasesDGX agent

arXiv:2605.26340v1 Announce Type: new Abstract: Autonomous research agents produce competitive solutions and professional-looking manuscripts, yet their outputs contain verifiability failures undetect

Starlette, an open-source Python framework underpinning FastAPI, has a vulnerability called BadHost that can allow hackers to bypass authorization (Dan Goodin/Ars Technica)

IndustryDGX agent

Dan Goodin / Ars Technica: Starlette, an open-source Python framework underpinning FastAPI, has a vulnerability called BadHost that can allow hackers to bypass authorization — Millions of AI agents an

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving

SafetyDGX agent

arXiv:2605.27038v1 Announce Type: new Abstract: Vision-Language Models (VLMs) provide a promising foundation for autonomous driving planning, yet bridging semantic reasoning and precise 3D spatial for

Understanding the Challenges in Iterative Generative Optimization with LLMs

ResearchDGX agent

arXiv:2603.23994v2 Announce Type: replace-cross Abstract: Generative optimization uses large language models (LLMs) to iteratively improve artifacts (such as code, workflows or prompts) using executio

Zero-Shot MARL Benchmark in the Cyber-Physical Mobility Lab

Model ReleasesDGX agent

arXiv:2601.16578v2 Announce Type: replace Abstract: We present a reproducible benchmark for evaluating sim-to-real transfer of Multi-Agent Reinforcement Learning (MARL) policies for Connected and Auto

26 May 2026

a shout out to the paper: https://arxiv.org/html/2605.25376v1 'KYA: A Framework-Agnostic Trust Layer for Autonomous Systems with Verifiable …

SafetyDGX agent

KYA is a framework-agnostic trust layer designed for autonomous systems that provides verifiable guarantees, addressing the need for trustworthy and transparent operation of AI agents across different

AuthTrace: Diagnosing Evidence Construction in Thematically Dense Single-Author Corpora

Model ReleasesDGX agent

arXiv:2605.25382v1 Announce Type: new Abstract: Evidence construction systems--chunk retrieval, agent memory, knowledge-graph traversal, and thematic indexing--are evaluated on separate benchmarks wit

Delayed Assignments in Online Non-Centroid Clustering with Stochastic Arrivals

ResearchDGX agent

arXiv:2601.16091v2 Announce Type: replace-cross Abstract: Clustering is a fundamental problem, aiming to partition a set of elements, like agents or data points, into clusters such that elements in th

Emission-Aware Reinforcement Learning for Sustainable Electric Vehicle Charging and Carbon Dioxide Reduction Under Varying Renewable Penetration

Model ReleasesDGX agent

arXiv:2605.24543v1 Announce Type: new Abstract: The rapid growth of Electric Vehicle (EV) adoption challenges power distribution networks through peak load spikes, voltage instability, and transformer

Generative Visual Code Mobile World Models

SafetyDGX agent

arXiv:2602.01576v2 Announce Type: replace-cross Abstract: Mobile Graphical User Interface (GUI) World Models (WMs) offer a promising path for improving mobile GUI agent performance at train- and infer

Hadamard Representation: Scaffolding Performance Across Model-free RL

Model ReleasesDGX agent

arXiv:2406.09079v5 Announce Type: replace Abstract: Deep reinforcement learning agents progressively lose representational capacity during training: neurons become dormant, removing active capacity fr

How to Mitigate the Distribution Shift Problem in Robotics Control: A Robust and Adaptive Approach Based on Offline to Online Imitation Learning

SafetyDGX agent

arXiv:2605.25414v1 Announce Type: new Abstract: Distribution shift in imitation learning refers to the problem that the agent cannot plan proper actions for a state that has not been visited during th

In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models

ApplicationsDGX agent

arXiv:2605.23908v1 Announce Type: new Abstract: We are in the midst of large-scale industrial and academic efforts to automate the processes of scientific, technological and creative production throug

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

SafetyDGX agent

arXiv:2605.24414v1 Announce Type: new Abstract: We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe

Learning in Low-Dimensional Subspaces: Orthogonal Bottlenecks for Reinforcement Learning

SafetyDGX agent

arXiv:2605.26012v1 Announce Type: cross Abstract: Deep reinforcement learning (RL) agents commonly rely on high-dimensional neural representations, despite growing evidence that task-relevant value an

NVIDIA Vera CPU Is ‘Packing a Heavy-Hitting Punch’ Against Competition

Model ReleasesDGX agent

The shift to agentic AI creates a new CPU requirement for the AI factory: fast cores, massive memory bandwidth and the ability to sustain high performance when all cores are active. Initial benchmark

On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits

SafetyDGX agent

arXiv:2605.25789v1 Announce Type: cross Abstract: We study a stochastic multi-armed bandit problem where an agent is granted a free exploration budget before regret accumulates, a setting not captured

ran my first benchmark this weekend (longmemeval) mostly to test activegraph, learned a lot! - this is a stepping stone to show the event ba…

Model ReleasesDGX agent

ran my first benchmark this weekend (longmemeval) mostly to test activegraph, learned a lot! - this is a stepping stone to show the event based agent system works. the AI convinced me not to start wit

Streaming Reinforcement Learning under Partial Observability with Real-Time Recurrent Learning

Model ReleasesDGX agent

arXiv:2605.24709v1 Announce Type: new Abstract: Streaming reinforcement learning has emerged as an online learning paradigm that conforms to the restrictions of natural learning agents that process da

Structural Abstraction as an Inductive Bias for Non-Stationary Language Model Training

Model ReleasesDGX agent

arXiv:2603.17198v2 Announce Type: replace-cross Abstract: A foundational principle in cognitive science holds that intelligent agents do not learn by storing experiences as isolated instances, but by

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

Model ReleasesDGX agent

arXiv:2605.24703v1 Announce Type: cross Abstract: Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA). Unlike text-on

Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs

ResearchDGX agent

arXiv:2601.14340v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are widely integrated into interactive systems such as dialogue agents and task-oriented assistants. This growing

When Search Becomes Memory: Turning Robot Design Trials into Transferable Skills

Model ReleasesDGX agent

arXiv:2605.25832v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as proposal generators for evolutionary robot design, yet most loops remain memoryless: simulator r

25 May 2026

Design and Report Benchmarks for Knowledge Work

Model ReleasesDGX agent

arXiv:2605.23262v1 Announce Type: new Abstract: The development of LLM agents has led to a growing body of work on knowledge-work AI, including coding, research, and healthcare. However, current knowl

DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.23176v1 Announce Type: new Abstract: Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, main

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.23238v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as economic agents in marketplaces, auctions, and bidding settings. Anticipating their behavior i

How Far Will They Go? Red-Teaming Online Influence with Large Language Models

Local AiDGX agent

arXiv:2605.22880v1 Announce Type: cross Abstract: As large language model (LLM)-based agents increasingly participate in online discourse, red-teaming their capacity to support political influence cam

给小伙买了Hugging Face家的Reachy Mini,目前桌面陪伴机器人里地表最强了。 不光是配件做工好,IDE和开发者生态好。还支持Agentic编程,10岁以下小朋友+codex很容易给它加功能。 手册上说组装需要3小时,小伙今天自己干了两个小时,就剩个头部了,3小时…

IndustryDGX agent

The post reviews the Hugging Face Reachy Mini robot, praising its build quality, IDE, and developer ecosystem for desktop companion robotics applications. It highlights the robot's support for agentic

MedExpMem: Adapting Experience Memory for Differential Diagnosis

Model ReleasesDGX agent

arXiv:2605.22872v1 Announce Type: cross Abstract: Experienced physicians develop diagnostic expertise through clinical practice, acquiring not only disease knowledge but also the ability to differenti

Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study

Model ReleasesDGX agent

arXiv:2605.23108v1 Announce Type: cross Abstract: AI-assisted code review tools typically operate as generic 'expert reviewer' agents, producing homogeneous findings regardless of the analysis type ne

← Previous
1…231232233234235…297
Next →