AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
27 May 2026

EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation

SafetyDGX agent

arXiv:2605.26785v1 Announce Type: cross Abstract: Post-trained LLMs are often optimized to align responses with human preferences, making them safe, polite, and conversationally appropriate. In advers

Fast, faster, Qwen. 🚀 Thrilled to see Qwen3.5 reaching a record-breaking 580 tps for agentic workloads on the TokenSpeed engine! This miles…

Model ReleasesDGX agent

Fast, faster, Qwen. 🚀 Thrilled to see Qwen3.5 reaching a record-breaking 580 tps for agentic workloads on the TokenSpeed engine! This milestone wouldn't be possible without our incredible partners. Hu

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

ITBench-AA is a new benchmark developed by Artificial Analysis and IBM that evaluates frontier AI models on agentic enterprise IT tasks, with results showing that current leading models score below 50

it's boston tech week! come join @masondrxy and i with the @blitzyai team to learn about building deep agents! https://luma.com/9ob847de

TutorialsDGX agent

Boston Tech Week featured a session hosted by Mason Dry and Sydney Runkle with the Blitzy AI team focused on building deep agents. The event was promoted via Luma Events and appears to be an education

Role-Based Access Control for Humans and Agents

ApplicationsDGX agent

This article discusses implementing role-based access control (RBAC) systems that work for both human users and AI agents, likely addressing how to manage permissions and authentication in environment

Shopping Companion: A Memory-Augmented LLM Agent for Real-World E-Commerce Tasks

Model ReleasesDGX agent

arXiv:2603.14864v2 Announce Type: replace Abstract: In e-commerce, LLM agents show promise for shopping tasks such as recommendations, budget management, and bundle deals, where accurately capturing u

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets

Model ReleasesDGX agent

arXiv:2605.26165v1 Announce Type: cross Abstract: Agentic RAG systems that equip language models with dozens to hundreds of tool definitions face a critical resource conflict: tool schemas consume the

TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents

Model ReleasesDGX agent

arXiv:2601.05899v2 Announce Type: replace Abstract: Recent breakthroughs in Large Language Models (LLMs) have positioned them as a promising paradigm for agents, with long-term planning and decision-m

26 May 2026

Agent-Facing Information Design in LLM Tool Registries

SafetyDGX agent

arXiv:2605.23916v1 Announce Type: cross Abstract: LLM tool registries function as unregulated advertising platforms: providers write free-text descriptions that agents use for selection, yet no measur

Agents and AI responsibility; nice clip from @thsottiaux and @siliconvalleymm

SafetyDGX agent

Gary Marcus shares a video clip discussing the intersection of AI agents and questions of responsibility, featuring contributors Thierry Souttiaux and Silicon Valley commentators. The post highlights

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning

ResearchDGX agent

arXiv:2605.23939v1 Announce Type: new Abstract: Web agents require both high-level reasoning (for task decomposition) and low-level interactions (for page elements manipulation) to conduct different t

Dynamic Dual-Granularity Skill Bank for Agentic RL

SafetyDGX agent

arXiv:2603.28716v2 Announce Type: replace Abstract: Agentic RL can benefit substantially from reusable experience, yet existing skill-based methods mainly extract trajectory-level guidance and often l

ECHO: Terminal Agents Learn World Models for Free

SafetyDGX agent

arXiv:2605.24517v1 Announce Type: cross Abstract: CLI agents are the closest thing language models have to an embodied setting: the model emits commands, the terminal executes them, and the returned s

Memory-Induced Tool-Drift in LLM Agents

Model ReleasesDGX agent

arXiv:2605.24941v1 Announce Type: cross Abstract: Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinn

Mitigating Provenance-Role Collapse in Long-Term Agents via Typed Memory Representation

ResearchDGX agent

arXiv:2605.25869v1 Announce Type: new Abstract: Long-term memory is essential for persistent LLM agents, yet prevailing architectures store historical interactions as unstructured, flat text. This unc

MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

AgentsDGX agent

arXiv:2605.26114v1 Announce Type: new Abstract: We present MobileGym, a browser-hosted, lightweight, fully controllable environment for everyday mobile use, targeting interaction fidelity without repl

Rambus targets agentic AI workloads with faster client memory chipset

HardwareDGX agent

Rambus Inc. today announced a complete DDR5 9600 client memory module chipset designed to push PC memory speeds to 9,600 megatransfers per second, targeting the bandwidth and capacity demands of agent

The OpenAI insider @thsottiaux has a warning for everyone offloading their thinking to agents.

SafetyDGX agent

An OpenAI insider (@thsottiaux) raises concerns about the risks of over-relying on AI agents to handle cognitive tasks, warning against wholesale delegation of thinking to autonomous systems. The warn

VeriTrace: Evolving Mental Models for Deep Research Agents

Model ReleasesDGX agent

arXiv:2605.26081v1 Announce Type: new Abstract: Deep research agents face vast, interdependent, and pervasively uncertain information. Existing systems explore what evolving intermediate representatio

When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

Model ReleasesDGX agent

arXiv:2605.25981v1 Announce Type: new Abstract: We document an empirical phenomenon in chain-of-thought and ReAct agents driven by ten large language models from seven architecture families: meaning-b

25 May 2026

Ax-Prover: A Deep Reasoning Agentic Framework for Theorem Proving in Mathematics and Quantum Physics

Model ReleasesDGX agent

arXiv:2510.12787v4 Announce Type: replace Abstract: We present Ax-Prover, a multi-agent system for automated theorem proving in Lean that can solve problems across diverse scientific domains and opera

Goal-Conditioned Agents that Learn Everything All at Once

SafetyDGX agent

arXiv:2605.23551v1 Announce Type: cross Abstract: A goal-conditioned reinforcement learning agent exploring an environment will see a wealth of information throughout a trajectory, most of which is di

HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation

Local AiDGX agent

arXiv:2605.23043v1 Announce Type: new Abstract: Agentic text-simulation systems write in sequence, with each item becoming possible context for later steps. That makes uncertainty path-dependent: an e

LLM-driven design of physics-constrained constitutive models: two agents are better than one

Model ReleasesDGX agent

arXiv:2605.23754v1 Announce Type: new Abstract: Developing constitutive models that capture how materials deform under load traditionally requires years of specialized expertise in continuum mechanics

What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA

Model ReleasesDGX agent

arXiv:2605.23067v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a viable recipe for training LLM agents to reason over external memory banks in multi-session dialogue. Exist

WMAttack: Automated Attack Search for Adversarial Evaluation of World-Model Agents

ResearchDGX agent

arXiv:2605.23220v1 Announce Type: new Abstract: Despite the growing use of world models as decision-making agents, their adversarial robustness remains underexplored due to the lack of dedicated autom

22 May 2026

Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2605.22001v1 Announce Type: cross Abstract: Injection detectors deployed to protect LLM agents are calibrated on static, template-based payloads that announce themselves as override directives.

Diverse Yet Consistent: Context-Guided Diffusion with Energy-Based Joint Refinement for Multi-Agent Motion Prediction

Model ReleasesDGX agent

arXiv:2605.22017v1 Announce Type: new Abstract: Deepgenerative models havebecomeapromisingapproach for human motion prediction due to their ability to capture multimodal distributions and represent di

SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents

AgentsDGX agent

arXiv:2605.21965v1 Announce Type: new Abstract: Large language models increasingly use external tools such as web search and document retrieval to solve information-intensive tasks. However, multi-hop

VBFDD-Agent for Electric Vehicle Battery Fault Detection and Diagnosis: Descriptive Text Modeling of Battery Digital Signals

Local AiDGX agent

arXiv:2605.20742v1 Announce Type: new Abstract: With the rapid proliferation of electric vehicles, the safety and reliability of lithium-ion batteries have become critical concerns. Effective anomaly

21 May 2026

Google just revealed Omni, personalized cross-device intelligence, and Spark agents at I/O 2025. I sat down with CEO Sundar Pichai to figure…

Model ReleasesDGX agent

Google just revealed Omni, personalized cross-device intelligence, and Spark agents at I/O 2025. I sat down with CEO Sundar Pichai to figure out what comes next: 1:46 Omni: 'Nano Banana for video' 4:5

Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search

Model ReleasesDGX agent

arXiv:2605.20244v1 Announce Type: cross Abstract: We present Lean Refactor, a plug-and-play retrieval-augmented agentic framework for multi-objective, controllable, and version-robust refactoring of L

Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory

Model ReleasesDGX agent

arXiv:2602.06025v2 Announce Type: replace Abstract: Memory is increasingly central to Large Language Model (LLM) agents operating beyond a single context window, yet most existing systems rely on offl

Spotify Studio’s AI agent creates a daily podcast just for you

AgentsDGX agent

Studio by Spotify Labs is a new standalone AI app that generates a daily briefing, podcasts, and playlists on your PC using chatbot prompts. The AI-generated content draws from your Spotify listening

TRAM: Test-Time Risk Adaptation with Mixture of Agents

Model ReleasesDGX agent

arXiv:2408.08812v2 Announce Type: replace Abstract: Deployed reinforcement learning agents often face safety requirements that are specified only after training, such as new hazard maps, revised risk

20 May 2026

[AINews] Google I/O 2026: Gemini 3.5 Flash, Omni (NanoBanana for Video), Spark (background agents), and Antigravity 2.0

Model ReleasesDGX agent

Google I/O 2026 featured several new AI model releases including Gemini 3.5 Flash, an Omni model codenamed NanoBanana for video processing, Spark for background agent tasks, and Antigravity 2.0. These

Formal Skill: Programmable Runtime Skills for Efficient and Accurate LLM Agents

SafetyDGX agent

arXiv:2605.19604v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly act inside real workspaces, where tools and skills determine whether model reasoning becomes reliable act

KadiAssistant: A conversational AI Agent for information retrieval in Kadi4Mat

AgentsDGX agent

arXiv:2605.18850v1 Announce Type: cross Abstract: We introduce KadiAssistant, a privacy-by-design AI assistant integrated into the Kadi research data ecosystem, enabling researchers to efficiently acc

LLM agents & memory systems operate in continuously updated environments (Git repos, evolving docs). They must process long contexts, recove…

SafetyDGX agent

LLM agents & memory systems operate in continuously updated environments (Git repos, evolving docs). They must process long contexts, recover earlier information, and reason over many updates that cre

OpenComputer: Verifiable Software Worlds for Computer-Use Agents

ResearchDGX agent

arXiv:2605.19769v1 Announce Type: new Abstract: We present OpenComputer, a verifier-grounded framework for constructing verifiable software worlds for computer-use agents. OpenComputer integrates four

Physics-in-the-Loop: A Hybrid Agentic Architecture for Validated CAD Engineering Design

Model ReleasesDGX agent

arXiv:2605.19717v1 Announce Type: new Abstract: Large Language Models (LLMs) can generate Computer-Aided Design (CAD), yet lack physical comprehension required for reliable engineering design. Instead

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents

SafetyDGX agent

arXiv:2605.20061v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) is a promising paradigm for improving large language model (LLM) agents on long-horizon interactiv

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking

AgentsDGX agent

arXiv:2605.18852v1 Announce Type: cross Abstract: Checkpoint selection for multimodal large language models (MLLMs) presents significant challenges when performance differentials are marginal and eval

STAR-PolyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision

Model ReleasesDGX agent

arXiv:2605.19338v1 Announce Type: cross Abstract: Frontier AI models and multi-agent systems have led to significant improvements in mathematical reasoning. However, for problems requiring extended, l

Towards LLM-Assisted Architecture Recovery for Real-World ROS~2 Systems: An Agent-Based Multi-Level Approach to Hierarchical Structural Architecture Reconstruction

AgentsDGX agent

arXiv:2605.20055v1 Announce Type: cross Abstract: Explicit software architecture models are essential artifacts for communicating, analyzing, and evolving complex software-intensive systems. In ROS~2-

TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents

SafetyDGX agent

arXiv:2602.11767v3 Announce Type: replace Abstract: Advances in large language models (LLMs) are driving a shift toward using reinforcement learning (RL) to train agents from iterative, multi-turn int

19 May 2026

Agentic Chunking and Bayesian De-chunking of AI Generated Fuzzy Cognitive Maps: A Model of the Thucydides Trap

Model ReleasesDGX agent

arXiv:2605.17903v1 Announce Type: new Abstract: We automatically generate feedback causal fuzzy cognitive maps (FCMs) from text by teaching large-language-model agents to break the text into overlappi

ANNEAL: Adapting LLM Agents via Governed Symbolic Patch Learning

Local AiDGX agent

arXiv:2605.16309v1 Announce Type: new Abstract: LLM-based agents can recover from individual execution errors, yet they repeatedly fail on the same fault when the underlying process knowledge--operato

Benchmarking inference at scale: coding agents

Model ReleasesDGX agent

This article presents benchmarking results for AI coding agents evaluated at scale, likely comparing performance metrics such as code generation accuracy, execution success rates, and inference effici

Body-Grounded Perspective Formation and Conative Attunement in Artificial Agents

SafetyDGX agent

arXiv:2605.16728v1 Announce Type: new Abstract: This paper proposes a minimal architecture for body-grounded perspective formation in artificial agents. Extending prior work, the model introduces an i

Can I get my agents on the phone?

IndustryDGX agent

This article likely discusses the feasibility and methods of contacting AI agents or customer service representatives by phone, exploring whether voice communication is available as an interaction opt

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis

Model ReleasesDGX agent

arXiv:2605.18451v1 Announce Type: new Abstract: Designing realistic and functional 3D indoor rooms is essential for a wide range of applications, including interior design, virtual reality, gaming, an

ContractBench: Can LLM Agents Preserve Observation Contracts?

Model ReleasesDGX agent

arXiv:2605.17281v1 Announce Type: cross Abstract: Tool-augmented LLM agents call APIs whose intermediate outputs, such as presigned URLs, session tokens, and OAuth state parameters, are observation co

Ethical Hyper-Velocity (EHV): A Provably Deterministic Governance-Aware JIT Compiler Architecture for Agentic Systems

SafetyDGX agent

arXiv:2605.17909v1 Announce Type: new Abstract: As autonomous agentic systems scale across regulated critical infrastructures, the lack of mechanistic, hardware-rooted enforcement for high-frequency p

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis

Model ReleasesDGX agent

arXiv:2507.21035v3 Announce Type: replace Abstract: Gene expression analysis holds the key to many biomedical discoveries, yet extracting insights from raw transcriptomic data remains formidable due t

Google just announced agentic coding in search. Based on your search, Gemini processes your ask, decides whether it should build a new inter…

Model ReleasesDGX agent

Google just announced agentic coding in search. Based on your search, Gemini processes your ask, decides whether it should build a new interface, reasons through the steps, and creates a personalized

LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injectio

Model ReleasesDGX agent

arXiv:2605.17986v1 Announce Type: cross Abstract: AI agents such as OpenClaw are increasingly deployed in local workflows with access to external tools. This creates indirect prompt-injection (IPI) ri

Multi-agent AI systems outperform human teams in creativity

Local AiDGX agent

arXiv:2605.17885v1 Announce Type: cross Abstract: Although artificial intelligence (AI) now matches or exceeds human performance across numerous cognitive tasks, creativity remains a highly contested

NeuSymMS: A Hybrid Neuro-Symbolic Memory System for Persistent, Self-Curating LLM Agents

Model ReleasesDGX agent

arXiv:2605.17596v1 Announce Type: new Abstract: We present NeuSymMS, an adaptive memory system that enables large language model (LLM) agents to learn, remember, and reason about users across sessions

New upgrades to the @GeminiApp are you helping you get more done: ✨Gemini Spark is your 24/7 personal AI agent that can take action on your …

Model ReleasesDGX agent

New upgrades to the @GeminiApp are you helping you get more done: ✨Gemini Spark is your 24/7 personal AI agent that can take action on your behalf, under your direction. It seamlessly integrates with

← Previous
1…118119120121122…300
Next →