AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,737 results
13 Apr 2026

SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support

AgentsDGX agent

arXiv:2604.08618v1 Announce Type: cross Abstract: Deploying LLM-powered agents in enterprise scenarios such as cloud technical support demands high-quality, domain-specific skills. However, existing s

The model is not the agent. The harness is. You need to read this recent study, and a blog post from @hwchase17 ... (links below). It will r…

Model ReleasesDGX agent

The model is not the agent. The harness is. You need to read this recent study, and a blog post from @hwchase17 ... (links below). It will resonate deeply. This diagram from a recent paper captures so

Towards Context-Aware Image Anonymization with Multi-Agent Reasoning

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2603.27817v3 Announce Type: replace-cross Abstract: Street-level imagery contains personally identifiable information (PII), some of which is context-dependent. Existing anonymization methods ei

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning

Model ReleasesDGX agent

arXiv:2510.07517v5 Announce Type: replace Abstract: Multi-agent debate (MAD) aims to improve large language model (LLM) reasoning by letting multiple agents exchange answers and then aggregate their o

Zoom Perspectives: Why ‘agentic’ work is the new enterprise standard

AgentsDGX agent

I had been waiting for the 2026 edition of Zoom Communications Inc.‘s Perspectives, its recently held annual get-together for industry analysts, because I find Zoom to be the most interesting vendor i

12 Apr 2026

I built a free, open-source CLI coding agent for 8k-context LLMs — v0.2 now shows diffs before touching your files

AgentsDGX agent

A community-built, free, open-source CLI coding agent shared on r/ollama, specifically optimized for local LLMs with 8k context windows to help developers work within the more constrained token limits

Model providers don’t lock you in with the API. They lock you in with your own data. Memory is the moat. If you don’t own your agent’s harne…

AgentsDGX agent

Model providers don’t lock you in with the API. They lock you in with your own data. Memory is the moat. If you don’t own your agent’s harness, you don’t own your agent’s memory. And switching means s

11 Apr 2026

Great piece. The lock-in point is the one nobody talks about enough. If your agent’s memory lives behind someone else’s API, you don’t have …

AgentsDGX agent

Great piece. The lock-in point is the one nobody talks about enough. If your agent’s memory lives behind someone else’s API, you don’t have a product. You have a dependency. Learned this early buildin

This is the right frame. We’re currently designing our agent memory platform, and the hardest part isn’t storage — it’s deciding what to rem…

AgentsDGX agent

This is the right frame. We’re currently designing our agent memory platform, and the hardest part isn’t storage — it’s deciding what to remember, when to retrieve, and how to keep context clean. All

10 Apr 2026

Big lineup today! Agent 4 Buildathon Week 2 winners are joining us to talk about building fast, getting attention, and shipping their ideas.…

AgentsDGX agent

Big lineup today! Agent 4 Buildathon Week 2 winners are joining us to talk about building fast, getting attention, and shipping their ideas. Plus @_talaawwad and Content Challenge winner Marcos Leal o

ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents

AgentsDGX agent

arXiv:2604.07789v1 Announce Type: cross Abstract: Recent advances in language model (LM) agents have significantly improved automated software engineering (SWE). Prior work has proposed various agenti

RAGEN-2: Reasoning Collapse in Agentic RL

AgentsDGX agent

arXiv:2604.06268v1 Announce Type: new Abstract: RL training of multi-turn LLM agents is inherently unstable, and reasoning quality directly determines task performance. Entropy is widely used to track

Robust Multi-Agent Target Tracking in Intermittent Communication Environments via Analytical Belief Merging

AgentsDGX agent

arXiv:2604.07575v1 Announce Type: new Abstract: Autonomous multi-agent target tracking in GPS-denied and communication-restricted environments (e.g., underwater exploration, subterranean search and re

9 Apr 2026

And that is what Thoth uses for its agents in the backend. Thanks for your contributions to oss @LangChain @hwchase17 https://github.com/sid…

AgentsDGX agent

And that is what Thoth uses for its agents in the backend. Thanks for your contributions to oss @LangChain @hwchase17 https://github.com/siddsachar/Thoth AI that doesn't just work for you, it knows yo

Deep Agents or @LangChain? Just use the right tool for the right job

Model ReleasesDGX agent

Deep Agents or @LangChain? Just use the right tool for the right job great q! deepagents has more 'batteries included', which langchain v1 is a very minimalistic agent harness if you are doing more co

Studying Sutton and Barto's RL book and its connections to RL for LLMs (e.g., tool use, math reasoning, agents, and so on)? [D]

AgentsDGX agent

A Reddit discussion thread on r/MachineLearning in which practitioners explore how foundational concepts from Sutton and Barto's *Reinforcement Learning: An Introduction* — including MDPs, policy g...

8 Apr 2026

Built a disaster relief app with Agent 4 on Replit. Went viral on social media in less than 24 hours. Here's the story. Floods hit Dagestan.…

AgentsDGX agent

Built a disaster relief app with Agent 4 on Replit. Went viral on social media in less than 24 hours. Here's the story. Floods hit Dagestan. 400,000 people evacuated. Thousands of homes destroyed. Peo

13 Aug 2026

Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations

Model ReleasesDGX agent

arXiv:2608.11323v1 Announce Type: new Abstract: Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (TheAgentCompany,

Designing Agentic AI-Based Screening for Portfolio Investment

AgentsDGX agent

arXiv:2603.23300v2 Announce Type: replace-cross Abstract: We introduce a new agentic artificial intelligence (AI) platform for portfolio management. Our architecture consists of three layers. First, t

12 Aug 2026

According to AMD, Arm, and Microsoft, agentic AI could push CPU-to-GPU ratios from 1:4 to even1:1

Local AiDGX agent

In OCP APAC 2026, Tai AMD SVP of compute and enterprise AI said agents don't cut GPU demand but they just pile on a whole extra layer of orchestration, retrieval, and tool-calling work that runs on CP

// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fi…

SafetyDGX agent

// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fixes cost, latency, failure mode, and auditability. New resea

MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows

Model ReleasesDGX agent

arXiv:2608.10509v1 Announce Type: new Abstract: Shared memory helps language-model agents reuse information across long workflows, yet relevant evidence may not be admissible for a particular agent or

11 Aug 2026

ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization

AgentsDGX agent

arXiv:2608.09577v1 Announce Type: new Abstract: Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisoned skill can p

FriskAI launches with $3.6M to show enterprises what their AI agents are doing

Model ReleasesDGX agent

Runtime intelligence startup FriskAI Inc. launched today with 3.6 million in pre-seed funding to give enterprises a record of what their artificial intelligence agents actually do once they go into pr

From Product Search to Preference Articulation: The Economics of Agentic Commerce

Model ReleasesDGX agent

arXiv:2608.08395v1 Announce Type: cross Abstract: Generative AI is shifting digital commerce from browsing toward agentic search, in which consumers delegate product discovery to AI agents. We compare

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

Model ReleasesDGX agent

arXiv:2608.09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and imp

Wix launches Symphony, a new standalone multi-agent system built for business operations

Model ReleasesDGX agent

Cloud-based website builder Wix Ltd. today announced the launch of Symphony, a new standalone agentic artificial intelligence platform that proactively learns business values, interests, needs, practi

10 Aug 2026

StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection

Model ReleasesDGX agent

arXiv:2608.06477v1 Announce Type: cross Abstract: Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as

Very cool idea to have agents design complex systems by searching over the model structure itself. New research from Sakana AI introduces CE…

Model ReleasesDGX agent

Very cool idea to have agents design complex systems by searching over the model structure itself. New research from Sakana AI introduces CEDAR, which uses LLM agents to write, simulate, and refine sy

9 Aug 2026

AI agentic Internet traffic will obviously VASTLY exceed human usage. Not a close call at all. Cloudflare’s forecast is accurate.

AgentsDGX agent

AI agentic Internet traffic will obviously VASTLY exceed human usage. Not a close call at all. Cloudflare’s forecast is accurate. For context, Global bandwidth is somewhere between 2-8 Pbps (2,000-8,0

7 Aug 2026

How cheap models changed multi-agent economics

AgentsDGX agent

Orchestrator-executor just became the smart default for production agents: an expensive model plans, cheap models execute, and cost per completed task decides the roster. The post How cheap models cha

6 Aug 2026

Building an agentic app deployer with Amazon Bedrock and AWS Lambda

AgentsDGX agent

PDI Technologies built PDI Brew, an agentic platform on AWS where non-technical employees describe a tool in plain English and receive a fully provisioned, multi-tenant web application in seconds. See

EASy: Towards Efficient LLM-Based Agentic System

AgentsDGX agent

arXiv:2608.04588v1 Announce Type: cross Abstract: Agentic systems have emerged as a promising paradigm for solving complex tasks by coordinating specialized LLM-based agents. However, most existing sy

Meta takes on Anthropic and OpenAI with its first AI coding agent, Muse Code

AgentsDGX agent

Meta Platforms Inc. is getting more serious in its efforts to challenge leading artificial intelligence labs Anthropic PBC and OpenAI PBC with the release of its first AI coding agent, called Muse Cod

MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding

AgentsDGX agent

arXiv:2608.04587v1 Announce Type: new Abstract: Long-form video understanding requires locating sparse, question-relevant evidence in long, multimodal videos. Real-world video distributions differ in

5 Aug 2026

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

AgentsDGX agent

arXiv:2608.00155v1 Announce Type: cross Abstract: Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominan

BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL

SafetyDGX agent

arXiv:2608.02876v1 Announce Type: new Abstract: Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context

Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG

AgentsDGX agent

arXiv:2608.02011v2 Announce Type: replace Abstract: Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snipp

Harness choice is a big deal. So much room to advance and improve results across the board with agent harnesses. Great paper highlighting th…

Model ReleasesDGX agent

Harness choice is a big deal. So much room to advance and improve results across the board with agent harnesses. Great paper highlighting this. New research releases DataSpace, a benchmark where data

How we built an MCP bridge to give our AgentCore-hosted AI agent access to local MCP tools

Local AiDGX agent

AI agents on Amazon Bedrock AgentCore run in the cloud, but users' tools and files live on their laptops. Learn how to build a secure MCP bridge that lets a cloud-hosted agent call local MCP servers b

TraceCAD: Trace-Guided Repair for Agentic CAD Generation

AgentsDGX agent

arXiv:2608.03062v1 Announce Type: new Abstract: LLM-based CAD agents produce executable parametric programs, but their correction loops may lose evidence about satisfied requirements, faulty operation

4 Aug 2026

CRISP: Critical Step Perception for Training Efficient Deep Search Agents

SafetyDGX agent

arXiv:2608.01867v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external

DiffuseAgent-MI: Distributionally-Grounded,Tool-Integrated Self-Evolving Agents for Faithful Visual Reasoning

SafetyDGX agent

arXiv:2608.00540v1 Announce Type: new Abstract: Tool-integrated vision-language agents have made remarkable progress on compositional and multi-step visual reasoning. Yet their outputs frequently exhi

MedTextWeaver: Procedural Knowledge Evolution in Agentic Medical Text Editing

AgentsDGX agent

arXiv:2602.00740v2 Announce Type: replace Abstract: Medical text editing is essential for improving communication among diverse stakeholders in clinical settings. However, adapting LLM agents to this

When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems

AgentsDGX agent

arXiv:2608.00747v1 Announce Type: new Abstract: Large language models are increasingly integrated into autonomous robotic systems for task planning and control, but this integration exposes them to pr

3 Aug 2026

Agentic Harness for Real-World Compilers

Model ReleasesDGX agent

arXiv:2603.20075v2 Announce Type: replace-cross Abstract: Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements enable

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

SafetyDGX agent

arXiv:2607.29190v1 Announce Type: new Abstract: Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates gener

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

Local AiDGX agent

arXiv:2512.03438v3 Announce Type: replace Abstract: Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally opt

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability

AgentsDGX agent

arXiv:2607.28942v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented

31 Jul 2026

datasette-agent 0.4a0

AgentsDGX agent

Release: datasette-agent 0.4a0 New await context.browser_task() mechanism allowing agent tools to run code directly in the user's browser. #33 This is an exciting new capability: it makes it easy for

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

AgentsDGX agent

arXiv:2607.28225v1 Announce Type: new Abstract: Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, h

ORCA-bench: How Ready Are Language Model Agents for Oncall?

Model ReleasesDGX agent

arXiv:2607.28545v1 Announce Type: new Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics,

SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation

Model ReleasesDGX agent

arXiv:2607.26313v1 Announce Type: cross Abstract: Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise defects are

30 Jul 2026

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

Model ReleasesDGX agent

arXiv:2607.26977v1 Announce Type: new Abstract: Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once

Voice Memory for Agentic Speech Recognition

AgentsDGX agent

arXiv:2607.26410v1 Announce Type: new Abstract: We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md

29 Jul 2026

Authoring Agent Skills: A Software-Engineering Approach

Model ReleasesDGX agent

arXiv:2607.25032v1 Announce Type: cross Abstract: Agent Skills are an emerging way to extend large language model agents with reusable procedural knowledge that the agent loads on demand. Anthropic in

Super interesting new work from NVIDIA. (bookmark it) They suggest building agents as Python objects. Very cool idea and I think it could a …

HardwareDGX agent

Super interesting new work from NVIDIA. (bookmark it) They suggest building agents as Python objects. Very cool idea and I think it could a lot with agent reliability. More below: Agent development to

28 Jul 2026

A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility

AgentsDGX agent

arXiv:2607.24663v1 Announce Type: cross Abstract: Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, i

Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents

Model ReleasesDGX agent

arXiv:2607.22689v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions th

Moral Hazard in Multi-Agent Language Models

SafetyDGX agent

arXiv:2607.23982v1 Announce Type: cross Abstract: Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmstrom's team moral-hazard

← Previous
1…3334353637…296
Next →