AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,771 results
12 May 2026

Learning Strategic Value and Cooperation in Multi-Player Stochastic Games through Side Payments

ResearchDGX agent

arXiv:2303.05307v2 Announce Type: replace-cross Abstract: We study general-sum, multi-player stochastic games with transferable utility, motivated by settings where agents can use side payments to mak

LLM Advertisement based on Neuron Auctions

SafetyDGX agent

arXiv:2605.08326v1 Announce Type: cross Abstract: As Large Language Models (LLMs) transition into conversational agents, generative advertising emerges as a crucial monetization strategy. However, emb

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

Model ReleasesDGX agent

arXiv:2605.08678v1 Announce Type: new Abstract: Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonst

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Neural Co-state Policies: Structuring Hidden States in Recurrent Reinforcement Learning

SafetyDGX agent

arXiv:2605.05373v2 Announce Type: replace Abstract: A key capability of intelligent agents is operating under partial observability: reasoning and acting effectively despite missing or incomplete stat

PDEAgent-Bench: A Multi-Metric, Multi-Library Benchmark for PDE Solver Generation

Model ReleasesDGX agent

arXiv:2605.09636v1 Announce Type: new Abstract: PDE-to-solver code generation aims to automatically synthesize executable numerical solvers from partial differential equation (PDE) specifications. Thi

Playing games with knowledge: AI-Induced delusions need game theoretic interventions

Local AiDGX agent

arXiv:2605.08409v1 Announce Type: new Abstract: Conversational AI has a fundamental flaw as a knowledge interface: sycophantic chatbots induce epistemic entrenchment and delusional belief spirals even

Position: AI Security Policy Should Target Systems, Not Models

Model ReleasesDGX agent

arXiv:2605.09504v1 Announce Type: cross Abstract: We present swarm-attack, an open-source adversarial testing framework in which multiple lightweight LLM agents coordinate through shared memory, paral

Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias

SafetyDGX agent

arXiv:2605.08315v1 Announce Type: new Abstract: Existing LLM-based policy optimizers see only scalar rewards: that a policy scored 0.45, but not whether the agent got stuck in a loop, fell into a hole

SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators

SafetyDGX agent

arXiv:2605.08334v1 Announce Type: new Abstract: We present SalesSim, a framework and testbed for evaluating the ability of Multimodal Large Language Models (MLLMs) to simulate realistic, persona-drive

Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind

Model ReleasesDGX agent

arXiv:2603.26089v2 Announce Type: replace-cross Abstract: The ability to represent oneself and others as agents with knowledge, intentions, and belief states that guide their behavior - Theory of Mind

Shields to Guarantee Probabilistic Safety in MDPs

SafetyDGX agent

arXiv:2605.10888v1 Announce Type: cross Abstract: Shielding is a prominent model-based technique to ensure safety of autonomous agents. Classical shielding aims to ensure that nothing bad ever happens

Sign up if you're in LA!

TutorialsDGX agent

Sign up if you're in LA! The last AI Agents Happy Hour was so fun, @jvedi and I are going to do it again. This time, co-hosted by @pinecone! We want to see what you're building and learn from your exp

Statistical Model Checking of the Keynes+Schumpeter Model: A Transient Sensitivity Analysis of a Macroeconomic ABM

Model ReleasesDGX agent

arXiv:2605.10447v1 Announce Type: cross Abstract: Agent-based models (ABMs) are increasingly used in macroeconomics, but their analysis still often relies on ad hoc Monte Carlo campaigns with heteroge

Step Rejection Fine-Tuning: A Practical Distillation Recipe

Model ReleasesDGX agent

arXiv:2605.10674v1 Announce Type: cross Abstract: Rejection Fine-Tuning (RFT) is a standard method for training LLM agents, where unsuccessful trajectories are discarded from the training set. In the

Talk to Your Slides: High-Efficiency Slide Editing via Language-Driven Structured Data Manipulation

Model ReleasesDGX agent

arXiv:2505.11604v5 Announce Type: replace Abstract: Editing presentation slides is a frequent yet tedious task, ranging from creative layout design to repetitive text maintenance. While recent GUI-bas

TIE: Time Interval Encoding for Video Generation over Events

SafetyDGX agent

arXiv:2605.10543v1 Announce Type: new Abstract: Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which

Towards Conversational Medical AI with Eyes, Ears and a Voice

Model ReleasesDGX agent

arXiv:2605.09272v1 Announce Type: new Abstract: The practice of medicine relies not only upon skillful dialogue but also on the nuanced exchange and interpretation of rich auditory and visual cues bet

Try Grok Voice

Model ReleasesDGX agent

Try Grok Voice Grok Voice Think Fast 1.0 ranks #1 on the Artificial Analysis τ-Voice benchmark for real-world agentic customer service resolution Absolutely outperforming GPT-Realtime-2 (High) and Gem

What Parameter Golf taught us about AI-assisted research

Model ReleasesDGX agent

Parameter Golf brought together 1,000+ participants and 2,000+ submissions to explore AI-assisted machine learning research, coding agents, quantization, and novel model design under strict constraint

What to expect during KB4-CON: Join theCUBE May 14

ApplicationsDGX agent

Human risk management is becoming a practical measure of enterprise security. The old playbook treated employees as the weak link; the new one has to account for people, AI agents and automated decisi

When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews

Model ReleasesDGX agent

arXiv:2605.10171v1 Announce Type: cross Abstract: Scientific peer reviews frequently contain conflicting expert judgments, and the increasing scale of conference submissions makes it challenging for A

11 May 2026

AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models

ApplicationsDGX agent

arXiv:2605.07308v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have significantly advanced the capabilities of robotic agents in executing diverse tasks; however, they still face

Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.07333v1 Announce Type: new Abstract: In-context reinforcement learning (ICRL) studies agents that, after pretraining, adapt to new tasks by conditioning on additional context without parame

Contrast-X: A Multi-Modal Contrast Image Synthesis Benchmark and Universal Modality Flow Matching

Model ReleasesDGX agent

arXiv:2601.15884v2 Announce Type: replace Abstract: Contrast-enhanced imaging is central to oncologic diagnosis, but contrast agents can be contraindicated for many of the patients who need them most.

Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought

Model ReleasesDGX agent

arXiv:2605.07123v1 Announce Type: new Abstract: In-context reinforcement learning (ICRL) refers to the ability of RL agents to adapt to new tasks at inference time without parameter updates by conditi

Decentralized Time-Varying Optimization for Streaming Data via Temporal Weighting

SafetyDGX agent

arXiv:2605.06971v1 Announce Type: cross Abstract: Classical optimization theory largely focuses on fixed objective functions, whereas many modern learning systems operate in dynamic environments where

Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation

Model ReleasesDGX agent

arXiv:2605.07323v1 Announce Type: new Abstract: Discovering governing differential equations from observational data is a fundamental challenge in scientific machine learning. Existing symbolic regres

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators

Model ReleasesDGX agent

arXiv:2605.06997v1 Announce Type: new Abstract: Long chain-of-thought reasoning and agentic tool-calling produce traces spanning tens of thousands of tokens, yet Transformer KV caches grow linearly wi

Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning

SafetyDGX agent

arXiv:2605.06156v2 Announce Type: replace-cross Abstract: Integrating expressive generative policies, such as flow-matching models, into offline reinforcement learning (RL) allows agents to capture co

From observability to context: What’s next for Arize Phoenix

ToolsDGX agent

As agents start changing software, they need a way to verify their work that includes traces, evals, feedback, and APIs. This is where Phoenix goes next — not the next release, but what this product b

Multi-Environment POMDPs with Finite-Horizon Objectives

SafetyDGX agent

arXiv:2605.07537v1 Announce Type: new Abstract: Partially Observable Markov Decision Processes (POMDPs) are systems in which one agent interacts with a stochastic environment, and receives only partia

Multi-Objective Constraint Inference using Inverse reinforcement learning

Model ReleasesDGX agent

arXiv:2605.06951v1 Announce Type: new Abstract: Constraint inference is widely considered essential to align reinforcement learning agents with safety boundaries and operational guidelines by observin

OpenAI just released its answer to Claude Mythos

Model ReleasesDGX agent

OpenAI is launching Daybreak, an AI initiative focused on detecting and patching vulnerabilities before attackers find them. Daybreak uses the Codex Security AI agent that launched in March to create

SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints

SafetyDGX agent

arXiv:2512.23770v3 Announce Type: replace-cross Abstract: In safety-critical domains, reinforcement learning (RL) agents must often satisfy strict, zero-cost safety constraints while accomplishing tas

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

SafetyDGX agent

arXiv:2605.07447v1 Announce Type: cross Abstract: Vision-language models (VLMs) have advanced rapidly and are increasingly deployed in real-world applications, especially with the rise of agent-based

10 May 2026

Local AI is having its moment! Below is the number of new GGUF models created each month over the past 8 months & insights from our HF inter…

Model ReleasesDGX agent

Local AI is having its moment! Below is the number of new GGUF models created each month over the past 8 months & insights from our HF internal agent (May is partial): - 176,000 total public GGUF mode

8 May 2026

AI Ascent 2026

IndustryDGX agent

AI Ascent 2026 is Sequoia Capital's event that convened leading AI researchers and founders including Greg Brockman, Andrej Karpathy, and Demis Hassabis. The event featured discussions on AI agents, s

Anthropic gave these out yesterday at code with claude. Added personalized memory and Claude to it. You can just build things. @bcherny @trq…

Model ReleasesDGX agent

Anthropic gave these out yesterday at code with claude. Added personalized memory and Claude to it. You can just build things. @bcherny @trq212 Time to add managed agents to this. That’s going to be s

7 May 2026

9 themes defining the future of AI-driven customer engagement: Insights from Twilio’s Signal event

TutorialsDGX agent

Customer engagement is shifting from disconnected interactions to continuous, AI-driven experiences that span channels and adapt in real time — and unified data, orchestration and AI agents are rapidl

Distilling Bayesian Belief States into Language Models for Auditable Negotiation

SafetyDGX agent

arXiv:2605.04507v1 Announce Type: new Abstract: Negotiation agents must infer what their counterpart values, update those beliefs over dialogue turns, and choose actions under uncertainty. End-to-end

🚀 Introducing the genmedia CLI, generative media directly from the command line. Generate images, video, 3D and audio from your terminal, a…

Model ReleasesDGX agent

🚀 Introducing the genmedia CLI, generative media directly from the command line. Generate images, video, 3D and audio from your terminal, alongside Claude and other AI agents. • Native terminal workfl

Mozilla says 271 vulnerabilities found by Mythos have 'almost no false positives'

IndustryDGX agent

Mozilla ran an agentic harness powered by Claude Mythos Preview across Firefox's source code, identifying 271 security bugs fixed in Firefox 150 . The breakthrough was achieved through improvements in

New #YAAP episode out now 🎙️ @yuvalinthedeep sits down with @mikegchambers from @awsdevelopers to unpack harness engineering and why it's t…

ApplicationsDGX agent

New #YAAP episode out now 🎙️ @yuvalinthedeep sits down with @mikegchambers from @awsdevelopers to unpack harness engineering and why it's the reason most agents never make it to production. 🎧 Listen/W

OpenClaw and Claude can put your AI-generated podcasts in Spotify

Model ReleasesDGX agent

Save to Spotify is a new command-line tool designed specifically for AI agents like OpenClaw, Claude Code, or OpenAI Codex. If you're the kind of person who collects research on a topic, then feeds it

6 May 2026

A Benchmark for Interactive World Models with a Unified Action Generation Framework

Model ReleasesDGX agent

arXiv:2605.03941v1 Announce Type: new Abstract: Achieving Artificial General Intelligence (AGI) requires agents that learn and interact adaptively, with interactive world models providing scalable env

A Sentence Relation-Based Approach to Sanitizing Malicious Instructions

ResearchDGX agent

arXiv:2605.01078v1 Announce Type: cross Abstract: Retrieval-augmented generation and tool-integrated LLM agents increasingly depend on external textual sources. This reliance broadens the available at

Code with Claude is happening now! ▪︎ 9:00AM - Keynote ▪︎ 10:30AM - What's new in Claude Code ▪︎ 11:15AM - Building on Claude at GitHub scal…

Model ReleasesDGX agent

Code with Claude is happening now! ▪︎ 9:00AM - Keynote ▪︎ 10:30AM - What's new in Claude Code ▪︎ 11:15AM - Building on Claude at GitHub scale ▪︎ 12:00PM - Get to production faster with Managed Agents

Devs are spending only 16% of their time coding. Atlassian is engineering AI to reclaim the rest

IndustryDGX agent

The rise of AI coding agents has commoditized code generation, exposing a deeper challenge for software teams: the non-coding friction that consumes the vast majority of a developer’s day. Fixing that

From Where Things Are to What They’re For: Benchmarking Spatial–Functional Intelligence for Multimodal LLMs

Model ReleasesDGX agent

True spatial intelligence for multimodal agents transcends low-level geometric perception, evolving from knowing where things are to understanding what they are for. While existing benchmarks, such as

https://x.com/walden_yan/status/2052070983083942322?s=20

ToolsDGX agent

https://x.com/walden_yan/status/2052070983083942322?s=20 Cool to see failure modes of different coding agents in new report from @greptile - seems like Devin is better than humans in almost all catego

Intervention Complexity as a Canonical Reward and a Measure of Intelligence

SafetyDGX agent

arXiv:2605.02175v1 Announce Type: new Abstract: The Legg--Hutter universal intelligence measure provides a rigorous scalar assessment of general intelligence as expected reward across all computable e

MEMAUDIT: An Exact Package-Oracle Evaluation Protocol for Budgeted Long-Term LLM Memory Writing

ResearchDGX agent

arXiv:2605.02199v1 Announce Type: new Abstract: Long-term LLM agents must compress streams of past interactions into persistent memory before future queries are known. Existing evaluations usually mea

On-Device Fine-Tuning via Backprop-Free Zeroth-Order Optimization

Local AiDGX agent

arXiv:2511.11362v2 Announce Type: replace-cross Abstract: On-device fine-tuning is a critical capability for edge AI systems, which must support adaptation to different agentic tasks under stringent m

Poly-EPO: Training Exploratory Reasoning Models

SafetyDGX agent

arXiv:2604.17654v3 Announce Type: replace Abstract: Exploration is a cornerstone of learning from experience: it enables agents to find solutions to complex problems, generalize to novel ones, and sca

Safety and accuracy follow different scaling laws in clinical large language models

Model ReleasesDGX agent

arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation

The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence

SafetyDGX agent

arXiv:2408.12622v3 Announce Type: replace-cross Abstract: Artificial intelligence (AI) is reshaping society, from video generation to medical diagnosis, coding agents to autonomous vehicles. Yet resea

Towards Understanding Specification Gaming in Reasoning Models

Model ReleasesDGX agent

arXiv:2605.02269v1 Announce Type: new Abstract: Specification gaming is a critical failure mode of LLM agents. Despite this, there has been little systematic research into when it arises and what driv

ValueBlindBench: Agreement-Gated Stress Testing of LLM-Judged Investment Rationales Before Returns Are Observable

TutorialsDGX agent

arXiv:2604.25224v2 Announce Type: replace Abstract: LLM-based financial agents increasingly produce investment rationales before the outcomes needed to evaluate them are observable. This creates a del

We're winding back our peak hours limit reduction and doubling 5 hour limits. Excited to partner with SpaceX to bring you more compute and w…

Model ReleasesDGX agent

We're winding back our peak hours limit reduction and doubling 5 hour limits. Excited to partner with SpaceX to bring you more compute and we'll keep pushing to bring you the best coding agent in the

5 May 2026

AFFormer: Adaptive Feature Fusion Transformer for V2X Cooperative Perception under Channel Impairments

SafetyDGX agent

arXiv:2605.01888v1 Announce Type: new Abstract: Accurate 3D object detection is essential for ensuring the safety of autonomous vehicles. Cooperative perception, which leverages vehicle-to-everything

← Previous
1…234235236237238…297
Next →