AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,920 results
9 Jul 2026

For coding agents, trustworthiness has to be tested in the harness where the model actually writes code. In one surveillance scenario, Kimi …

AgentsDGX agent

For coding agents, trustworthiness has to be tested in the harness where the model actually writes code. In one surveillance scenario, Kimi K2.7 complied with the request in 8/8 samples. SWE-1.7 refus

Meta prices Muse Spark 1.1 at 1.25/1M input tokens and 4.25/1M output tokens; Alexandr Wang says improving coding and agentic performance was a key focus (Ina Fried/Axios)

AgentsDGX agent

Ina Fried / Axios: Meta prices Muse Spark 1.1 at 1.25/1M input tokens and 4.25/1M output tokens; Alexandr Wang says improving coding and agentic performance was a key focus — Facebook's parent company

Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in Production

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety
DGX agent

arXiv:2607.07052v1 Announce Type: cross Abstract: AI agents deployed for IT operations are typically permanent cost centers because every execution requires full LLM inference, even for previously sol

Security and Privacy in Agentic AI: Grand Challenges and Future Directions

AgentsDGX agent

arXiv:2607.06608v1 Announce Type: cross Abstract: We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise that brought

SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis

AgentsDGX agent

arXiv:2607.07467v1 Announce Type: new Abstract: Spatial and Single-cell transcriptomics are transformative in deciphering cellular dynamics. As the fundamental paradigm for reconstructing cell develop

We're hosting a meetup on agent memory and wikis July 28th at our SF office! Come hear @jacobtpl and myself talk about the frontier research…

AgentsDGX agent

We're hosting a meetup on agent memory and wikis July 28th at our SF office! Come hear @jacobtpl and myself talk about the frontier research going on in these areas right now. https://luma.com/mylwoab

8 Jul 2026

aiAuthZ: Off-Host, Identity-Bound Authorization for AI Agents

Model ReleasesDGX agent

arXiv:2607.05518v1 Announce Type: cross Abstract: AI agents issue tool calls on the basis of text they cannot verify, so any party who controls part of the context can forge the appearance of authorit

Grok will be able to call Imagine as a tool in agentic mode for image/video generation. As Imagine keeps improving, this will be amazing for…

AgentsDGX agent

Grok will be able to call Imagine as a tool in agentic mode for image/video generation. As Imagine keeps improving, this will be amazing for game developers! I got to try Grok 4.5 in early access in C

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

SafetyDGX agent

arXiv:2607.06223v1 Announce Type: new Abstract: Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where the agent mus

MASCA: LLM based-Multi Agents System for Credit Assessment

SafetyDGX agent

arXiv:2507.22758v2 Announce Type: replace Abstract: Recent advancements in financial problem-solving have leveraged LLMs and agent-based systems, with a primary focus on trading and financial modeling

MCP-Enabled Agentic AI for Autonomous IPoDWDM Network Lifecycle Automation

AgentsDGX agent

arXiv:2607.05975v1 Announce Type: cross Abstract: This demo presents an MCP-enabled agentic AI architecture for autonomous control of vendor-agnostic IPoDWDM networks. We demonstrate live end-to-end l

“Model orchestration is in many ways the natural outgrowth of agentic engineering” Huge thanks to @AndrewYNg and @DeepLearningAI for the dee…

AgentsDGX agent

“Model orchestration is in many ways the natural outgrowth of agentic engineering” Huge thanks to @AndrewYNg and @DeepLearningAI for the deep dive into Sakana AI’s Fugu and Fugu-Ultra. The article hig

most agent labs are shy about acknowledging chinese model use because they need to sell to gov/defense cog team did the hard part to product…

AgentsDGX agent

most agent labs are shy about acknowledging chinese model use because they need to sell to gov/defense cog team did the hard part to productionize: 1. build a multilingual propaganda & censorship eval

New this week 🚀 Replit Community Profiles — Proof of Work for vibe coders. Your profile, your flex. Get an activity graph of your agent usa…

AgentsDGX agent

New this week 🚀 Replit Community Profiles — Proof of Work for vibe coders. Your profile, your flex. Get an activity graph of your agent usage and checkpoints, plus a Replit Power Ranking for pro users

Reward as An Agent for Embodied World Models

AgentsDGX agent

arXiv:2606.19990v2 Announce Type: replace Abstract: While RL has become a promising tool for refining world models, existing methods largely rely on conservative rollouts near the training distributio

Task Decomposition-Guided Reranking for Adaptive Agent Skill Retrieval

AgentsDGX agent

arXiv:2607.06283v1 Announce Type: new Abstract: Skill usage can significantly enhance the ability of modern agent systems to complete complex tasks. However, the growing scale of skill libraries makes

The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities

Model ReleasesDGX agent

arXiv:2607.05743v1 Announce Type: cross Abstract: AI coding agents now read repositories, call tools, and execute shell commands with limited human oversight, and a fast-growing body of work studies w

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

SafetyDGX agent

arXiv:2606.20023v2 Announce Type: replace-cross Abstract: As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, pri

7 Jul 2026

Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions

AgentsDGX agent

arXiv:2607.03233v1 Announce Type: cross Abstract: The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intel

Build a serverless image editing agent with Amazon Bedrock AgentCore harness

AgentsDGX agent

This post walks through building a serverless image editor where users upload a photo, describe an edit in plain English, and receive the result in seconds. The agent runs on AgentCore harness without

Even before the agentic revolution, prompting tricks stopped being very valuable, as our research has shown. The best approach to AI right n…

AgentsDGX agent

Even before the agentic revolution, prompting tricks stopped being very valuable, as our research has shown. The best approach to AI right now is to clearly specify your goals, your output, what 'good

EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

Model ReleasesDGX agent

arXiv:2607.05202v1 Announce Type: new Abstract: Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for

Hierarchical Multi-Agent Reinforcement Learning for Carbon-Aware AI Data Centers in Power Distribution Systems

Local AiDGX agent

arXiv:2607.03324v1 Announce Type: cross Abstract: Eco-friendly energy management for artificial intelligence data centers (AIDCs) is crucial because of the significant increase in energy consumption-i

llm wikis are a glimpse of the future of what agent memory looks like this blog i wrote resonated with a lot of folks will be discussing thi…

AgentsDGX agent

llm wikis are a glimpse of the future of what agent memory looks like this blog i wrote resonated with a lot of folks will be discussing this (as well as some updates ive made to my beliefs since this

Memory-Orchestrated Semantic System (MOSS): An Auditable Agentic Memory Architecture

Local AiDGX agent

arXiv:2607.04391v1 Announce Type: new Abstract: Long-term memory remains a structural weakness of AI agents. The dominant approach, retrieval-augmented generation (RAG), relies on embedding-based simi

New in Hermes Agent: pull secrets from multiple vaults at once. Run @Bitwarden and newly added vault provider @1Password side by side, or ad…

AgentsDGX agent

New in Hermes Agent: pull secrets from multiple vaults at once. Run @Bitwarden and newly added vault provider @1Password side by side, or add any other secret source as a plugin with the new vault plu

P^3: Toward Versatile Embodied Agents

ApplicationsDGX agent

arXiv:2508.07033v2 Announce Type: replace Abstract: Embodied agents have shown promising generalization capabilities across diverse physical environments, making them essential for a wide range of rea

Since this is a slightly backwards-incompatible change here's a detailed upgrade guide - which you can read, or feed into your coding agent …

AgentsDGX agent

Since this is a slightly backwards-incompatible change here's a detailed upgrade guide - which you can read, or feed into your coding agent and have it apply the upgrades for you https://sqlite-utils.

SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints

Model ReleasesDGX agent

arXiv:2607.05363v1 Announce Type: new Abstract: Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negot

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training

SafetyDGX agent

arXiv:2607.04963v1 Announce Type: new Abstract: Reinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse and delayed r

TACTIC-KG: Toward Small Agent Teams for Cyber Threat Intelligence Knowledge Graph Construction

AgentsDGX agent

arXiv:2607.05001v1 Announce Type: cross Abstract: Cyber Threat Intelligence (CTI) reports are predominantly unstructured, heterogeneous, and noisy, which limits their direct usability for automated an

TRACE: Capability-Targeted Agentic Training

Model ReleasesDGX agent

arXiv:2604.05336v2 Announce Type: replace Abstract: Models often fail to complete agentic tasks because they lack core capabilities required by the target environment. However, mainstream approaches f

6 Jul 2026

Big update to managing your past sessions in Hermes Agent. Now you can prune or archive with a huge array of filters to clean up your sessio…

AgentsDGX agent

Big update to managing your past sessions in Hermes Agent. Now you can prune or archive with a huge array of filters to clean up your session DB without losing anything you don't want to lose. Get rid

Built something great with II-Agent? Submit it to Showcase. earn credits. get featured. let others explore the result and replay the session…

AgentsDGX agent

II-Agent is a platform where users can submit their creations or projects to a showcase feature, earning credits and potential visibility while allowing other users to explore and replay their work se

🚨Emergency webinar: LLM Wikis and how to give your agent memory LLM Wikis are so hot right now - OpenWiki by @BraceSproul up to nearly 7k G…

AgentsDGX agent

🚨Emergency webinar: LLM Wikis and how to give your agent memory LLM Wikis are so hot right now - OpenWiki by @BraceSproul up to nearly 7k GitHub stars in less than a week I'll be chatting with Brace a

Hy3, the new 295B MoE model from @TencentHunyuan, is now free in Nous Portal for the next two weeks! It is focused on cost-effective agentic…

AgentsDGX agent

Hy3, the new 295B MoE model from @TencentHunyuan, is now free in Nous Portal for the next two weeks! It is focused on cost-effective agentic use, and particularly strong on coding, tool-calling reliab

Researchers document JadePuffer, the first known 'agentic ransomware', which adapts in real time and retries steps to execute an end-to-end extortion operation (Bill Toulas/BleepingComputer)

AgentsDGX agent

Bill Toulas / BleepingComputer: Researchers document JadePuffer, the first known “agentic ransomware”, which adapts in real time and retries steps to execute an end-to-end extortion operation — Resear

5 Jul 2026

I made a Hermes Agent slash command cheat sheet a while back. @NousResearch has shipped a ton of new stuff since then, so I tonight, I rebui…

AgentsDGX agent

I made a Hermes Agent slash command cheat sheet a while back. @NousResearch has shipped a ton of new stuff since then, so I tonight, I rebuilt it from scratch. Every official slash command as of 7.4.2

4 Jul 2026

Happy Fourth of July! We want to expand the plugin interface of Hermes Agent so that many developers who have PRs waiting for very long peri…

AgentsDGX agent

Happy Fourth of July! We want to expand the plugin interface of Hermes Agent so that many developers who have PRs waiting for very long periods can implement stable changes that they can share and pub

Hermes Agent is built for sovereignty and constructing your AI stack how you want and need it to be. No vendor lockins, no model limitations…

AgentsDGX agent

Hermes Agent is built for sovereignty and constructing your AI stack how you want and need it to be. No vendor lockins, no model limitations, and most importantly, your IP is built through the self im

3 Jul 2026

Atomic Task Graph: A Unified Framework for Agentic Planning and Execution

AgentsDGX agent

arXiv:2607.01942v1 Announce Type: new Abstract: LLM-based agents have shown strong potential for solving complex multi-step tasks, yet existing performance improvements often rely on either scaling to

BuilderBench: The Building Blocks of Intelligent Agents

Model ReleasesDGX agent

arXiv:2510.06288v4 Announce Type: replace Abstract: Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the limits set b

Controllable Sim Agents with Behavior Latents

Model ReleasesDGX agent

arXiv:2607.02496v1 Announce Type: cross Abstract: Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enabl

Criticality-Based Guard Rail Validation for AI Agent Decisions in Autonomous Telecom Networks

SafetyDGX agent

arXiv:2607.02210v1 Announce Type: new Abstract: The evolution toward fully autonomous telecommunications networks (Autonomous Network Levels 4-5) requires AI/ML agents to make real-time network decisi

ElephantAgent: Contextual State Continuity in Agentic Systems

SafetyDGX agent

arXiv:2607.01919v1 Announce Type: new Abstract: Agentic systems enhance their capabilities by invoking external tools and maintaining persistent memory. However, these external dependencies introduce

This is actually useful. LangChain just released OpenWiki. It's an open-source agent that creates a wiki for your codebase, connects it to y…

Model ReleasesDGX agent

This is actually useful. LangChain just released OpenWiki. It's an open-source agent that creates a wiki for your codebase, connects it to your coding agent, and keeps it updated as your repo changes.

When brainstorming new AI agent uses, I always remind myself what used to be expensive and is now nearly zero… 1) cost of reading everything…

AgentsDGX agent

When brainstorming new AI agent uses, I always remind myself what used to be expensive and is now nearly zero… 1) cost of reading everything fell → you can now watch 100% instead of a sample (every ex

Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents

Local AiDGX agent

arXiv:2511.10687v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) in multi-agent systems (MAS) have shown promise for complex tasks, yet current training methods lack principled w

2 Jul 2026

3/ Harbor integration Harbor is a framework for running long running, stateful agent evals We integrate with it in multiple different ways. …

AgentsDGX agent

3/ Harbor integration Harbor is a framework for running long running, stateful agent evals We integrate with it in multiple different ways. We wrote a blog on those integrations: https://www.langchain

// AutoMem // I quite like this idea of metamemory. (bookmark it) This new research from Stanford treats agent's memory management as a trai…

Model ReleasesDGX agent

// AutoMem // I quite like this idea of metamemory. (bookmark it) This new research from Stanford treats agent's memory management as a trainable skill instead of a fixed module. The model decides wha

Fable 5 is back on Replit! Especially great for longer, harder projects. Toggle on High effort mode in Replit Agent and try it today on your…

AgentsDGX agent

Replit has re-introduced Fable 5, an AI coding assistant designed to handle longer and more complex projects. Users can access this feature by enabling 'High effort mode' in the Replit Agent tool. Thi

From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives

AgentsDGX agent

arXiv:2607.00918v1 Announce Type: cross Abstract: Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consistency and co

I've been blogging a few of my Fable experiments today, most recently I had it more-or-less one-shot a CLI coding agent on top of my LLM Pyt…

AgentsDGX agent

I've been blogging a few of my Fable experiments today, most recently I had it more-or-less one-shot a CLI coding agent on top of my LLM Python library https://simonwillison.net/2026/Jul/2/llm-coding-

last day at @aiDotEngineer and i'll be at the Expo's poster area explaining the year's best survey paper on Agent Memory (Hu et al), as we d…

AgentsDGX agent

last day at @aiDotEngineer and i'll be at the Expo's poster area explaining the year's best survey paper on Agent Memory (Hu et al), as we did for @latentspacepod's Paper Club live from the floor with

PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents

Model ReleasesDGX agent

arXiv:2607.00436v1 Announce Type: new Abstract: Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more

.@Qualcomm is expanding its collaboration with @huggingface to scale open, developer-driven AI. From model onboarding to agentic workflows a…

AgentsDGX agent

.@Qualcomm is expanding its collaboration with @huggingface to scale open, developer-driven AI. From model onboarding to agentic workflows across edge and data center, this simplifies how developers b

Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains

Model ReleasesDGX agent

arXiv:2607.01136v1 Announce Type: cross Abstract: Agent skills package reusable operational knowledge for Large Language Model (LLM) agents, yet as they grow in scope, they become dependency-bearing a

The question is no longer “what can the agent do” because the list of things it can do keeps growing. The real question we should be asking …

AgentsDGX agent

The question is no longer “what can the agent do” because the list of things it can do keeps growing. The real question we should be asking is: what can only a human be accountable for? @addyosmani at

Using DSPy to evaluate and improve Datasette Agent's SQL system prompts

Model ReleasesDGX agent

Research: Using DSPy to evaluate and improve Datasette Agent's SQL system prompts One of this morning's AIE keynotes covered dspy, which reminded me I've been meaning to see if it could help me improv

You shouldn’t need to think about your memory or agent docs! It should *just work* which was the thesis we went into this with

AgentsDGX agent

You shouldn’t need to think about your memory or agent docs! It should *just work* which was the thesis we went into this with gonna try this out, it seems really cool. banking on this being a better

← Previous
1…7879808182…299
Next →