AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,951 results
10 Jun 2026

How do you support full-text search JSON filtering over agent traces that span up to hundreds of MBs, while keeping a median (P50) latency o…

AgentsDGX agent

How do you support full-text search JSON filtering over agent traces that span up to hundreds of MBs, while keeping a median (P50) latency of 400ms? Here’s an inside look at how we built a custom inve

in arxiv paper #2, i tackle the last topic from paper #1: @activegraphai as an architectural affordance for self-improving agents 'Regimes: …

AgentsDGX agent

in arxiv paper #2, i tackle the last topic from paper #1: @activegraphai as an architectural affordance for self-improving agents 'Regimes: An Auditable, Held-Out Gated Improvement Loop Demonstrated o

Introducing the Hermes Agent Profile Builder You can now build a complete profile in the dashboard with full control over identity/name/desc…

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents
DGX agent

Introducing the Hermes Agent Profile Builder You can now build a complete profile in the dashboard with full control over identity/name/description, model/provider, built-in + optional skills, skills-

less novel, but still very interesting impo is the gated approach to self-modification the agent basically forks itself, propose a patch, ru…

AgentsDGX agent

less novel, but still very interesting impo is the gated approach to self-modification the agent basically forks itself, propose a patch, run through multiple tests (static/sandbox/diff), and somethin

MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents

Model ReleasesDGX agent

arXiv:2606.10742v1 Announce Type: cross Abstract: External memory has become a core component of modern web agents, enabling long-horizon reasoning through the retrieval of past experiences. However,

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

Model ReleasesDGX agent

arXiv:2606.11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However,

TabClaw: An Interactive and Self-Evolving Agent for Spreadsheet Manipulation and Table Reasoning

AgentsDGX agent

arXiv:2606.10316v1 Announce Type: new Abstract: Spreadsheets and tables are widely used representations for structured data analysis, but effective analysis still requires substantial manual effort an

two fun surprises from using activegraph: - the coding agent i was using would query the trace db to debug instead of looking at the logs li…

AgentsDGX agent

two fun surprises from using activegraph: - the coding agent i was using would query the trace db to debug instead of looking at the logs like they normally would (i didn't ask it to) - when long eval

What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory

AgentsDGX agent

arXiv:2606.10299v1 Announce Type: new Abstract: Language-agent 'memory palace' systems anchor each memory to a world coordinate, on the intuition that geometry adds something text cannot. We make that

9 Jun 2026

A multi-agent system for spine MRI report generation from multi-sequence imaging

AgentsDGX agent

arXiv:2606.08897v1 Announce Type: cross Abstract: Spinal pathology is a leading cause of pain and disability worldwide. Spine MRI is central to clinical evaluation, yet its interpretation remains comp

Agentic multi-fidelity learning of quasiparticle and excitonic properties

AgentsDGX agent

arXiv:2606.07836v1 Announce Type: cross Abstract: Many-body GW-Bethe-Salpeter equation calculations are essential for accurate simulations of electronic structure and optical properties in modern low-

AgentTrust: A Self-Improving Trust Layer for AI-Agent Actions

AgentsDGX agent

arXiv:2606.08539v1 Announce Type: new Abstract: AI agents increasingly take consequential actions -- shell commands, cloud operations, and arbitrary tool-calls -- so a trust layer must decide, per act

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2606.07805v1 Announce Type: new Abstract: The rapid evolution of Large Language Models (LLMs) from passive assistants to autonomous, execution-capable agents has introduced critical operational

Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps

SafetyDGX agent

arXiv:2606.09084v1 Announce Type: cross Abstract: Tool-using LLM agents interact with the world through actions that persist state in artifacts (e.g., workspace files or logs). Consequently, jailbreak

Cost-Aware Speculative Execution for LLM-Agent Workflows: An Integrated Five-Dimension Method

AgentsDGX agent

arXiv:2606.07846v1 Announce Type: cross Abstract: LLM-agent workflows chain model calls and tool invocations, and spend most of their wall-clock time waiting on upstream operations before downstream o

Introducing Cohere's first open-source coding model: North Mini Code Small & efficient, designed for agentic performance and built for commu…

AgentsDGX agent

Cohere released North Mini Code, an open-source coding model designed to be small and efficient while optimizing for agentic performance and community use. The model represents Cohere's initial offeri

man its kinda wild to write “agent lab” on my random blog and a few months later its adopted by people like @breeves08 and @ScottWu46 🫡 htt…

AgentsDGX agent

man its kinda wild to write “agent lab” on my random blog and a few months later its adopted by people like @breeves08 and @ScottWu46 🫡 https://x.com/cognition/status/2062923088778367314?s=46 More tha

MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance

Model ReleasesDGX agent

arXiv:2605.22664v2 Announce Type: replace Abstract: LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet ente

Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human

SafetyDGX agent

arXiv:2606.08919v1 Announce Type: new Abstract: As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approva

POISE: Position-Aware Undetectable Skill Injection on LLM Agents

Model ReleasesDGX agent

arXiv:2606.07943v1 Announce Type: cross Abstract: Agent skills provide a lightweight mechanism for extending general-purpose agents, but their open format exposes them to skill-poisoning attacks. A pr

REFLECT: Intervention-Supported Error Attribution for Silent Failures in LLM Agent Traces

AgentsDGX agent

arXiv:2606.09071v1 Announce Type: new Abstract: Large language model (LLM) agents now solve complex tasks through long plan-and-execution traces, yet the ability to locate errors in a completed traces

Reminder: The only place to download the Hermes Agent Desktop App is https://hermes-agent.nousresearch.com/desktop Any other website or sour…

AgentsDGX agent

Reminder: The only place to download the Hermes Agent Desktop App is https://hermes-agent.nousresearch.com/desktop Any other website or source is dangerous and could contain old, or worse, dangerous a

SAGE: An LLM-driven Self Reflective Agentic Framework for Fraud Detection

AgentsDGX agent

arXiv:2606.08146v1 Announce Type: new Abstract: Fraud detection in payment, e-commerce, and telecommunications systems requires accuracy at the individual level, robustness under severe class imbalanc

WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

Model ReleasesDGX agent

arXiv:2606.09426v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and ext

8 Jun 2026

Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests

AgentsDGX agent

arXiv:2606.07379v1 Announce Type: cross Abstract: A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead of solving t

Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup​ 🔹Drag in videos as cod…

AgentsDGX agent

Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup​ 🔹Drag in videos as coding context: reference-to-LUT, long-video-to-short, screen-rec

LLM Agent-Assisted Reverse Engineering with Quantitative Readability Metrics

AgentsDGX agent

arXiv:2606.06838v1 Announce Type: cross Abstract: Automatic decompilers produce functionally correct but often unreadable C code. This paper addresses one stage of the reverse engineering workflow: im

Queen-Bee Agents: A BeeSpec-Centered Architecture for Governed Enterprise MCP Orchestration

Local AiDGX agent

arXiv:2606.06545v1 Announce Type: cross Abstract: Enterprise agent systems increasingly need to connect large language models to private tools, internal knowledge, and Model Context Protocol (MCP) int

SCALE: Scalable Cross-Attention Learning with Extrapolation for Agentic Workflow Scheduling

AgentsDGX agent

arXiv:2606.06820v1 Announce Type: cross Abstract: Agentic Large Language Model (LLM) systems decompose complex tasks into workflow Directed Acyclic Graphs (DAGs) whose primitives must be scheduled on

🔗Try it now: https://www.kimi.com/products/kimi-work We're just getting started. More data sources, more tools, more agent capabilities are…

AgentsDGX agent

Kimi is launching a new product called Kimi Work, an AI work platform with expandable capabilities including additional data sources, tools, and agent functionalities. The announcement indicates the p

We found that more autonomy with autonomous agents like Computer tracks with higher quality and satisfaction.

AgentsDGX agent

Research indicates that autonomous agents with greater operational autonomy, particularly in computer-based tasks, demonstrate improved performance quality and user satisfaction outcomes. The study su

We got into Y Combinator! Agnost AI (YC S26) is the infra for self-improving AI agents. We plug into conversational AI companies, find what'…

AgentsDGX agent

We got into Y Combinator! Agnost AI (YC S26) is the infra for self-improving AI agents. We plug into conversational AI companies, find what's broken, and ship the fix as a PR. You just merge. DM if th

We published new research with Harvard on the shift from chat interfaces to autonomous agents like Computer. Over 3 months, findings show wo…

AgentsDGX agent

We published new research with Harvard on the shift from chat interfaces to autonomous agents like Computer. Over 3 months, findings show workers using Computer finish tasks in 87% less time at 94% lo

7 Jun 2026

Agree with everything except the Markdown part There's got to be a better agent-native format for representing unstructured docs Not convinc…

AgentsDGX agent

Agree with everything except the Markdown part There's got to be a better agent-native format for representing unstructured docs Not convinced it's markdown or html I need Google Docs but just for mar

6 Jun 2026

Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning

AgentsDGX agent

arXiv:2503.01734v3 Announce Type: replace-cross Abstract: Attacks on machine learning models have been extensively studied through stateless optimization. In this paper, we demonstrate how a reinforce

Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2606.06036v1 Announce Type: new Abstract: Despite recent progress, LLM agents still struggle with reasoning over long interaction histories. While current memory-augmented agents rely on a stati

Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation

Model ReleasesDGX agent

arXiv:2606.05241v1 Announce Type: cross Abstract: Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the

The End of Software Engineering: How AI Agents Are Fundamentally Restructuring the Software Paradigm

Model ReleasesDGX agent

arXiv:2606.05608v1 Announce Type: cross Abstract: For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into s

This chart from Anthropic is useful, since Agent Teams and Workflows are both very new and very powerful (and token hungry). On the other ha…

AgentsDGX agent

This chart from Anthropic is useful, since Agent Teams and Workflows are both very new and very powerful (and token hungry). On the other hand, maybe it doesn't matter as a lot of the decisions about

Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems

SafetyDGX agent

arXiv:2606.06114v1 Announce Type: new Abstract: Self-evolving agents improve through continual self-play and self-generated learning signals, but autonomous evolution can also cause capability degrada

5 Jun 2026

Arena AI Agentic User Benchmark Ranking

Model ReleasesDGX agent

Arena AI's agentic benchmark ranks AI models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability. The leaderboard

EGTR-Review: Efficient Evidence-Grounded Scientific Peer Review Generation via Multi-Agent Teacher Distillation

AgentsDGX agent

arXiv:2606.06025v1 Announce Type: new Abstract: Scientific peer review generation has attracted increasing attention for reducing reviewing burdens and providing timely feedback. However, existing Lar

Harnessing Generalist Agents for Contextualized Time Series

AgentsDGX agent

arXiv:2606.05404v1 Announce Type: cross Abstract: Time series are often embedded in rich contexts that are essential for holistic modeling. Moreover, real-world practitioners often require end-to-end

Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

Model ReleasesDGX agent

arXiv:2606.06388v1 Announce Type: cross Abstract: Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly pos

In an internal message, Satya Nadella rebuked an internal memo that said Microsoft needs to 'make people addicted' to its new AI agent product called Scout (Aaron Holmes/The Information)

AgentsDGX agent

Aaron Holmes / The Information: In an internal message, Satya Nadella rebuked an internal memo that said Microsoft needs to “make people addicted” to its new AI agent product called Scout — Microsoft

Need advice on building/training an AI Agent for fully automated blog generation

AgentsDGX agent

A discussion on building an AI-powered article generator using CrewAI and Ollama with specialized AI agents for research and writing to generate comprehensive articles on any topic. The solution runs

Shopify on Replit + the new SEO Agent https://x.com/i/broadcasts/1kJzDDopENZKv

AgentsDGX agent

This post likely covers a live broadcast or announcement discussing the integration of Shopify with Replit, along with information about a newly released SEO Agent tool. The content probably demonstra

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework

Model ReleasesDGX agent

arXiv:2606.05570v1 Announce Type: new Abstract: Repository-level coding benchmarks face a trade-off between task difficulty and evaluation reliability: tasks that challenge frontier models often invol

4 Jun 2026

At @WitanLabs we're building the headless Office stack for AI agents. Here's what we worked on this week — and what we're building next ↓

AgentsDGX agent

Witan Labs is developing a headless Office stack designed specifically for AI agents, combining productivity tools and capabilities without a traditional user interface. This post appears to be a week

Be Fair! Can Machine Learning Engineering Agents Adhere to Fairness Constraints?

SafetyDGX agent

arXiv:2606.04971v1 Announce Type: new Abstract: Machine learning engineering (MLE) agents promise to automate end-to-end ML pipeline development from raw data and natural language instructions, potent

Cloudflare CEO Matthew Prince says agentic traffic is 'growing so fast that bots have now passed human traffic online for the first time' (Mark Tyson/Tom's Hardware)

AgentsDGX agent

Mark Tyson / Tom's Hardware: Cloudflare CEO Matthew Prince says agentic traffic is “growing so fast that bots have now passed human traffic online for the first time” — Bot (automated) vs. human HTTP

Cursor can now show your agent's context usage as an interactive report in a canvas. The context explorer breaks down where tokens go across…

AgentsDGX agent

Cursor can now show your agent's context usage as an interactive report in a canvas. The context explorer breaks down where tokens go across the system prompt, tool definitions, rules, skills, and mor

DAR: Deontic Reasoning with Agentic Harnesses

AgentsDGX agent

arXiv:2606.05009v1 Announce Type: cross Abstract: Deontic reasoning is the task of answering questions by applying explicit rules and policies to case-specific facts, for example computing tax liabili

Fog of Love: Engineering Virtuous Agent Behavior with Affinity-based Reinforcement Learning in a Game Environment

SafetyDGX agent

arXiv:2606.04750v1 Announce Type: new Abstract: Instilling virtuous behavior in artificial intelligence has seen increasing interest. One of the techniques proposed is known as affinity-based reinforc

Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning

AgentsDGX agent

arXiv:2507.21892v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) mitigates hallucination in LLMs by incorporating external knowledge, but relies on chunk-based retrieval that l

Learning While Acting: A Skill-Enhanced Test-Time Co-Evolution Framework for Online Lifelong Learning Agents

SafetyDGX agent

arXiv:2606.04815v1 Announce Type: cross Abstract: Lifelong learning is essential for Large Language Model (LLM) agents operating in dynamic, interactive environments. However, existing lifelong learni

On the latest episode of Max Agency, @hwchase17 sat down with @nlarusstone, Head of AI at @benchling for a conversation on building agents f…

AgentsDGX agent

On the latest episode of Max Agency, @hwchase17 sat down with @nlarusstone, Head of AI at @benchling for a conversation on building agents for scientific work. ⏯️ YouTube: https://www.youtube.com/watc

Optimizing the Cost-Quality Tradeoff of Agentic Theorem Provers in Lean

AgentsDGX agent

arXiv:2606.04883v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in workflows for generating formal proofs in Lean. These workflows often decompose problems into smal

Today I'm launching a new project called SynthTraces 🔥 It is a minimal codebase to generate synthetic coding agent session traces using Pi …

Model ReleasesDGX agent

Today I'm launching a new project called SynthTraces 🔥 It is a minimal codebase to generate synthetic coding agent session traces using Pi (from @badlogicgames) I wanted a large number of coding-agent

We partnered with Shopify so you can go from idea to live store in minutes Just tell Replit Agent what you want to sell. It will: - Build a …

AgentsDGX agent

We partnered with Shopify so you can go from idea to live store in minutes Just tell Replit Agent what you want to sell. It will: - Build a custom storefront - Create your Shopify store - Help you add

← Previous
1…8182838485…300
Next →