AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
12 May 2026

MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents

Model ReleasesDGX agent

arXiv:2605.09530v1 Announce Type: cross Abstract: As LLM-powered agents are increasingly deployed in edge-cloud environments, personalized memory has become a key enabler of long-term adaptation and u

MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents

ResearchDGX agent

arXiv:2603.13131v3 Announce Type: replace Abstract: Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A centra

​[PoC] Building a Local Multi-Agent AI Dev Studio alpha version (Architect/Senior/Junior) on a 10-year-old Haswell & GTX 1050 Ti (No APIs, Full AirLLM + Ollama)

Local Ai
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

This post describes a proof-of-concept implementation of a multi-agent AI development studio with architect, senior, and junior role personas, built entirely locally using AirLLM and Ollama without re

Priority-Driven Control and Communication in Decentralized Multi-Agent Systems via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.10482v1 Announce Type: cross Abstract: Event-triggered control provides a mechanism for avoiding excessive use of constrained communication bandwidth in networked multi-agent systems. Howev

SceneFactory: GPU-Accelerated Multi-Agent Driving Simulation with Physics-Based Vehicle Dynamics

SafetyDGX agent

arXiv:2605.08528v1 Announce Type: cross Abstract: Autonomous-driving simulators typically trade physical fidelity for scalable parallelism. Physics-based platforms such as CARLA and MetaDrive provide

The Invisible Handshake: Persistent Overpricing by Adaptive Market Agents

ResearchDGX agent

arXiv:2510.15995v3 Announce Type: replace-cross Abstract: We study overpricing in a repeated game between two representative agents: a market maker, who controls market liquidity, and a market taker,

TimeClaw: A Time-Series AI Agent with Exploratory Execution Learning

AgentsDGX agent

arXiv:2605.10038v1 Announce Type: new Abstract: Time series analysis underpins forecasting, monitoring, and decision making in domains such as finance and weather, where solving a task often requires

TodyComm: Task-Oriented Dynamic Communication for Multi-Round LLM-based Multi-Agent System

SafetyDGX agent

arXiv:2602.03688v2 Announce Type: replace Abstract: Multi-round LLM-based multi-agent systems rely on effective communication structures to support collaboration across rounds. However, most existing

Unpredictability dissociates from structured control in language agents

Model ReleasesDGX agent

arXiv:2605.09692v1 Announce Type: new Abstract: Unpredictable behavior is often taken as evidence of control, yet stochastic dispersion and structured action control need not coincide. This paper test

Verifiable Process Rewards for Agentic Reasoning

Local AiDGX agent

arXiv:2605.10325v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has improved the reasoning abilities of large language models (LLMs), but most existing approaches

11 May 2026

+1 to this. I was recently on a cross-continental flight without wifi, so I brought up Qwen3.6 & Gemma 4 (via @ollama) in Deep Agents on my …

Model ReleasesDGX agent

+1 to this. I was recently on a cross-continental flight without wifi, so I brought up Qwen3.6 & Gemma 4 (via @ollama) in Deep Agents on my laptop. admittedly, they fell over on some more involved/com

3 weeks since ml-intern launched and we just hit 1M messages exchanged. that's 3.3 agent-years of ML research in 21 days. 2 months worth of …

Model ReleasesDGX agent

3 weeks since ml-intern launched and we just hit 1M messages exchanged. that's 3.3 agent-years of ML research in 21 days. 2 months worth of research every day. 17,383 training jobs total. talk about A

Agent view is the best Claude Code native way to manage multiple sessions, kind of like tmux built for CC. We spent a lot of time getting th…

Model ReleasesDGX agent

Agent view is the best Claude Code native way to manage multiple sessions, kind of like tmux built for CC. We spent a lot of time getting the details right, I hope you enjoy it. New in Claude Code: ag

AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents

Model ReleasesDGX agent

arXiv:2605.06607v2 Announce Type: replace-cross Abstract: Recent LLM-based agents have closed substantial portions of the scientific discovery loop in software-only machine-learning research, in chemi

Excellent explainer video by @FryRsquared on the risks of AI agents. She also raises a crucial point: we shouldn’t make the mistake of think…

SafetyDGX agent

Excellent explainer video by @FryRsquared on the risks of AI agents. She also raises a crucial point: we shouldn’t make the mistake of thinking current limitations will necessarily persist. As we’ve s

FlightSense: An End-to-End MLOps Platform for Real-Time Flight Delay Prediction via Rotation-Chain Propagation Features and Agentic Conversational AI

AgentsDGX agent

arXiv:2605.07364v1 Announce Type: new Abstract: Flight delays impose cascading operational and financial burdens across the aviation network, costing the U.S. economy billions of dollars annually by d

GTIG AI Threat Tracker: Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access

Model ReleasesDGX agent

Executive Summary Since our February 2026 report on AI-related threat activity, Google Threat Intelligence Group (GTIG) has continued to track a maturing transition from nascent AI-enabled operations

HMACE: Heterogeneous Multi-Agent Collaborative Evolution for Combinatorial Optimization

Local AiDGX agent

arXiv:2605.07214v1 Announce Type: new Abstract: Large Language Models have recently emerged as a promising paradigm for automated heuristic design for NP-hard combinatorial optimization problems. Desp

Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents

Model ReleasesDGX agent

arXiv:2605.06957v1 Announce Type: new Abstract: We present a dynamic policy-learning approach that combines generalized planning and hierarchical task decomposition for LLM-based agents. Our method, H

// LLMs Improving LLMs // Interesting progress the past of couple of weeks around self-improving AI agents. If autoresearch was interesting,…

TutorialsDGX agent

// LLMs Improving LLMs // Interesting progress the past of couple of weeks around self-improving AI agents. If autoresearch was interesting, you will like this read. (bookmark it) We've been hand-tuni

Reachy Mini ready to go! Audio on 🔉 Cleary I'll connect it to Local AI services and to my Hermes Agent really soon 💪

Local AiDGX agent

Clem Delangue announced the Reachy Mini robot is operational with audio capabilities enabled, with plans to integrate it with local AI services and a Hermes Agent in the near future. This indicates pr

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair

Local AiDGX agent

arXiv:2605.07276v1 Announce Type: new Abstract: Code-agent RL often receives weak feedback: rollout-time signals are reliable and executable, but capture only necessary or surface conditions for task

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning

SafetyDGX agent

arXiv:2605.06130v2 Announce Type: replace Abstract: A persistent skill library allows language model agents to reuse successful strategies across tasks. Maintaining such a library requires three coupl

SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests

ResearchDGX agent

Using SocialReasoning Bench, we observed a stable pattern across models—agents execute competently, but fail to consistently improve the user’s position, even with explicit instructions to optimize fo

The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

Model ReleasesDGX agent

arXiv:2605.08060v1 Announce Type: cross Abstract: Context window expansion is often treated as a straightforward capability upgrade for LLMs, but we find it systematically fails in multi-agent social

When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

Model ReleasesDGX agent

arXiv:2605.06772v1 Announce Type: new Abstract: As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical questi

10 May 2026

> shipped the agent > opened the dashboard > latency: fine > error rate: fine > users: unhappy > checked the responses > technically correct…

SafetyDGX agent

> shipped the agent > opened the dashboard > latency: fine > error rate: fine > users: unhappy > checked the responses > technically correct > wrong tool called 3 steps earlier > no trace to follow >

9 May 2026

Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers — Last year, we released a case study on

8 May 2026

Agentic AI Is Going to the Edge – But Not the Way You Think

Local AiDGX agent

This article examines how agentic AI systems are being deployed to edge computing environments, likely challenging common assumptions about what this deployment actually entails. It discusses practica

7 May 2026

Agentic Vulnerability Reasoning on Windows COM Binaries

Model ReleasesDGX agent

arXiv:2605.05000v1 Announce Type: cross Abstract: Windows Component Object Model (COM) services run with elevated privileges and are widely accessible to authenticated users, making race conditions in

AWS unveils Amazon Bedrock AgentCore Payments and partners with Coinbase and Stripe to enable AI agents to execute transactions using stablecoins (RT Watson/The Block)

IndustryDGX agent

RT Watson / The Block: AWS unveils Amazon Bedrock AgentCore Payments and partners with Coinbase and Stripe to enable AI agents to execute transactions using stablecoins — Quick Take — Amazon Web Servi

CodeWords raises $9M to develop AI agents that automate before you ask

IndustryDGX agent

CodeWords, operated by Agemo AI Ltd., today announced it raised 9 million in seed funding led by Visionaries to build artificial intelligence agents that don’t wait to build automations. Firstminute C

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live

Model ReleasesDGX agent

arXiv:2511.02230v4 Announce Type: replace-cross Abstract: KV cache management is essential for efficient LLM inference. To maximize utilization, existing inference engines evict finished requests' KV

FaSTA^*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing

Local AiDGX agent

arXiv:2506.20911v2 Announce Type: replace Abstract: We develop a cost-efficient neurosymbolic agent to address challenging multi-turn image editing tasks such as ``Detect the bench in the image while

Graph-SND: Sparse Aggregation for Behavioral Diversity in Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2605.05020v1 Announce Type: new Abstract: System Neural Diversity (SND) measures behavioral heterogeneity in multi-agent reinforcement learning by averaging pairwise distances over all inom{n}{2

.@huggingface's agentic robotics app store for Reachy Mini is a big step toward more accessible physical AI. 🙌 Excited to see NVIDIA Isaac …

HardwareDGX agent

.@huggingface's agentic robotics app store for Reachy Mini is a big step toward more accessible physical AI. 🙌 Excited to see NVIDIA Isaac GR00T N integrated with Hugging Face LeRobot, helping develop

MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents

Model ReleasesDGX agent

arXiv:2605.03675v1 Announce Type: new Abstract: Long-running autonomous AI agents suffer from a well-documented memory coherence problem: tool-execution success rates degrade 14 percentage points over

Open Models x Headless Agent Execution 🔥

Model ReleasesDGX agent

Open Models x Headless Agent Execution 🔥 your daily reminder that open models are plenty capable for a lot of coding work. easiest place to feel that out is deepagents! swap the model and go. i've bee

QKVShare: Quantized KV-Cache Handoff for Multi-Agent On-Device LLMs

Model ReleasesDGX agent

arXiv:2605.03884v1 Announce Type: new Abstract: Multi-agent LLM systems on edge devices need to hand off latent context efficiently, but the practical choices today are expensive re-prefill or full-pr

ScrapMem: A Bio-inspired Framework for On-device Personalized Agent Memory via Optical Forgetting

Local AiDGX agent

arXiv:2605.03804v1 Announce Type: new Abstract: Long-term personalized memory for LLM agents is challenging on resource-limited edge devices due to high storage costs and multimodal complexity. To add

Spotify launches Save to Spotify, a command-line tool that allows AI agents to upload AI-generated audio summaries and personal podcasts to a user's account (Terrence O'Brien/The Verge)

Model ReleasesDGX agent

Terrence O'Brien / The Verge: Spotify launches Save to Spotify, a command-line tool that allows AI agents to upload AI-generated audio summaries and personal podcasts to a user's account — A new comma

Tessera Labs, which uses AI agents to automate enterprise IT migrations and ERP transformations, raised a 60M Series A led by a16z at a 320M valuation (Anna Tong/Forbes)

ApplicationsDGX agent

Anna Tong / Forbes: Tessera Labs, which uses AI agents to automate enterprise IT migrations and ERP transformations, raised a 60M Series A led by a16z at a 320M valuation — Kabir Nagrecha is tackling

TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments

Model ReleasesDGX agent

arXiv:2605.04107v1 Announce Type: cross Abstract: Production agent frameworks (OpenAI Function Calling, Anthropic Tool Use, MCP) transmit tool schemas as JSON, a format designed for machine parsing, n

6 May 2026

ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptoms

SafetyDGX agent

arXiv:2605.03212v1 Announce Type: cross Abstract: Modeling latent clinical constructs from unconstrained clinical interactions is a unique challenge in affective computing. We present ADAPTS (Agentic

Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework

Model ReleasesDGX agent

arXiv:2605.01604v1 Announce Type: new Abstract: Existing evaluation frameworks for large language models -- including HELM, MT-Bench, AgentBench, and BIG-bench -- are designed for controlled, single-s

From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level

Model ReleasesDGX agent

arXiv:2601.03731v3 Announce Type: replace-cross Abstract: As large language models (LLMs) evolve into autonomous agents, evaluating repository-level reasoning, the ability to maintain logical consiste

From Synthesis to Clinical Assistance: A Strategy-Aware Agent Framework for Autism Intervention based on Real Clinical Dataset

AgentsDGX agent

arXiv:2605.02916v1 Announce Type: new Abstract: The development of AI-assisted Early Intensive Behavioral Intervention (EIBI) for Autism Spectrum Disorder (ASD) is severely constrained by data scarcit

GISclaw: A Comprehensive Open-Source LLM Agent System for Realistic Multi-Step Geospatial Analysis

Local AiDGX agent

arXiv:2603.26845v2 Announce Type: replace-cross Abstract: Most LLM-driven GIS assistants solve narrow single-step tasks tightly coupled to proprietary platforms such as ArcGIS or QGIS, limiting their

GOAT: A Training Framework for Goal-Oriented Agent with Tools

Model ReleasesDGX agent

arXiv:2510.12218v2 Announce Type: replace Abstract: Current approaches rely on zero-shot evaluation due to the absence of training data; while proprietary models such as GPT-4 exhibit strong reasoning

It feels like agent harness evolution runs on two axes that usually get conflated. There’s the temporal axis: simplify as models improve, st…

Model ReleasesDGX agent

It feels like agent harness evolution runs on two axes that usually get conflated. There’s the temporal axis: simplify as models improve, stripping components that compensated for limitations the new

MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents

ResearchDGX agent

arXiv:2605.03482v1 Announce Type: cross Abstract: Persistent external memory enables LLM agents to maintain context across sessions, yet its security properties remain formally uncharacterized. We for

Monday.com relaunches as an AI work platform with native agents

IndustryDGX agent

Cloud project management provider monday.com Ltd. today relaunched itself as an “AI work platform,” repositioning itself around context-aware artificial intelligence agents that execute tasks alongsid

really cool work by the Harvey team, excited to partner with them to push forward research on designing + understanding agents across Long H…

ApplicationsDGX agent

really cool work by the Harvey team, excited to partner with them to push forward research on designing + understanding agents across Long Horizon Legal work the first peak I got at LAB was 'woah this

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use

Model ReleasesDGX agent

arXiv:2605.02964v1 Announce Type: new Abstract: Reinforcement learning (RL) trained language model agents with tool access are increasingly deployed in coding assistants, research tools, and autonomou

very fun to collab with @harvey on their Long Horizon Legal Agent Benchmark. We need more industry specific benchmarks, and Harvey is paving…

Model ReleasesDGX agent

Harrison Chase expresses enthusiasm about collaborating with Harvey on their Long Horizon Legal Agent Benchmark, highlighting the value of developing industry-specific benchmarks for AI evaluation. Th

5 May 2026

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.00425v1 Announce Type: new Abstract: Reinforcement learning (RL) has significantly advanced the ability of large language model (LLM) agents to interact with environments and solve multi-tu

Agentopic: A Generative AI Agent Workflow for Explainable Topic Modeling

Model ReleasesDGX agent

arXiv:2605.00833v1 Announce Type: new Abstract: Agentopic is a novel agent-based workflow for explainable topic modeling that leverages the reasoning capabilities of Large Language Models (LLMs). Exis

Automated Interpretability and Feature Discovery in Language Models with Agents

Model ReleasesDGX agent

arXiv:2605.01555v1 Announce Type: new Abstract: We introduce an autonomous multiagent framework for mechanistic interpretability that automates both explaining and finding internal features in large l

Beating the Style Detector: Three Hours of Agentic Research on the AI-Text Arms Race

Model ReleasesDGX agent

arXiv:2605.02620v1 Announce Type: new Abstract: Reproducing an empirical NLP study used to take weeks. Given the released data and a modern agentic-research harness, we redo every experiment of a rece

Cursor can now automatically fix CI failures. Set up always-on agents that monitor GitHub, investigate root causes, and open PRs with fixes.

ToolsDGX agent

Cursor has introduced an automated CI failure resolution feature that uses always-on agents to monitor GitHub repositories, analyze build failures, and automatically generate pull requests with fixes.

← Previous
1…120121122123124…300
Next →