AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
24 Jul 2026

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making

Model ReleasesDGX agent

arXiv:2607.20491v1 Announce Type: new Abstract: Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time. We i

From Agent Failures to Text Policies: What Works and What Breaks

SafetyDGX agent

arXiv:2607.20668v1 Announce Type: cross Abstract: TextGrad improves language-model systems by revising text from feedback. Its core thesis is that natural-language feedback can act as a gradient for o

GuardianAgentBench: Where Agents Fail and How to Guard Them

Model ReleasesDGX agent

arXiv:2607.20982v1 Announce Type: new Abstract: As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavi

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MemTools: A Unified Research Framework for Interoperable Agent Memory

Model ReleasesDGX agent

arXiv:2607.21404v1 Announce Type: new Abstract: While memory systems are essential for agent architectures, pervasive architectural fragmentation restricts systematic research. Existing implementation

MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation

AgentsDGX agent

arXiv:2607.20501v1 Announce Type: new Abstract: Despite rapid progress in LLM-based code generation, writing correct and performant kernels for hardware accelerators remains a key bottleneck in scalin

People are using Minecraft farms as AI agent benchmarks

Model ReleasesDGX agent

Someone modelled sugarcane farming as an integer program. See, sugarcane only grows next to water. Water costs one tile and can feed at most four cane tiles. The layout therefore becomes a coverage pr

Regulating autonomous and agentic AI

SafetyDGX agent

arXiv:2607.21345v1 Announce Type: new Abstract: Regulating activities where regulatees use autonomous and agentic AI is challenging. Regulatory assumptions about regulatee knowledge and control no lon

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs

Model ReleasesDGX agent

arXiv:2509.14257v3 Announce Type: replace-cross Abstract: Large Language Model agents achieve strong performance on multi-step reasoning and tool-use tasks, but their impressive capabilities typically

23 Jul 2026

ETPDesigner: Multi-Agent Orchestration for Interactive Multimodal Electronic Theater Program

Model ReleasesDGX agent

arXiv:2607.19947v1 Announce Type: new Abstract: Electronic Theater Programs (ETPs) serve as critical promotional media in the performing arts, comprising a multi-page collection of heterogeneous visua

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

Model ReleasesDGX agent

arXiv:2607.19356v1 Announce Type: new Abstract: Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential. We present NEXUS (Neural EXecution Utility a

OpenAI introduces Presence to help enterprises build AI agents

IndustryDGX agent

OpenAI Group PBC today introduced a product called Presence that enterprises can use to build artificial intelligence agents. The company is using the software to power its call center. According to O

OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field learn from. Did the …

AgentsDGX agent

OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field learn from. Did the top-level agent know about the hacking, or was there some 'v

// Programmatic Memory Enables Long-Horizon Reasoning // Keep the entire interaction log and search it. It works great and beats bespoke mem…

AgentsDGX agent

// Programmatic Memory Enables Long-Horizon Reasoning // Keep the entire interaction log and search it. It works great and beats bespoke memory harnesses on long-horizon tasks. New research introduces

Stress Testing Concept Erasure with Large Language Model Agents

SafetyDGX agent

arXiv:2607.17890v2 Announce Type: replace Abstract: Concept erasure aims to remove semantic concepts from a trained generative model and is increasingly important for responsible AI deployment. Howeve

The first known runaway AI agent - or a very bad marketing stunt?

Model ReleasesDGX agent

The first known runaway AI agent - or a very bad marketing stunt? Martin Alderson's commentary on the OpenAI accidental cyberattack against Hugging Face includes a couple of details I hadn't considere

Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review

SafetyDGX agent

arXiv:2507.10142v2 Announce Type: replace Abstract: Multi-Agent Reinforcement Learning (MARL) has achieved strong performance in simulated benchmarks, yet real deployments often violate the assumption

22 Jul 2026

Are structured outputs in agents always good? This paper suggests that you might have to take a closer look. Your product's structured outpu…

TutorialsDGX agent

Are structured outputs in agents always good? This paper suggests that you might have to take a closer look. Your product's structured output surface is measurably more homogeneous than the chat surfa

microsoft/Fara1.5-27B · Hugging Face

Model ReleasesDGX agent

Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting struc

OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened

Model ReleasesDGX agent

This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke

This is OpenSWE! It's an OSS coding agent that runs in the cloud and lives in Slack. Use it for coding, general Q&A, planning, etc etc etc. …

Model ReleasesDGX agent

This is OpenSWE! It's an OSS coding agent that runs in the cloud and lives in Slack. Use it for coding, general Q&A, planning, etc etc etc. It's truly a jack of all trades (and very widely used at Lan

21 Jul 2026

NVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AI

HardwareDGX agent

NVIDIA’s Vera CPU, built around the Olympus core, is engineered for agentic‑AI workloads that rely heavily on single‑thread performance, deep memory‑level parallelism, and efficient handling of irregu

20 Jul 2026

At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI

HardwareDGX agent

At SIGGRAPH 2026, NVIDIA unveiled a suite of AI‑driven graphics and simulation advances, highlighting neural rendering, agentic and physical AI world models, and real‑time simulation methods. Key rele

16 Jul 2026

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents

Model ReleasesDGX agent

arXiv:2511.18685v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) show promising results as decision-making engines for embodied agents operating in complex, physical enviro

Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation

Model ReleasesDGX agent

arXiv:2607.10057v1 Announce Type: cross Abstract: Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circuit Vision,

15 Jul 2026

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

Local AiDGX agent

arXiv:2607.12640v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkp

BREAKING: Inkling by @thinkymachines is 9th overall on Agentic Web App Arena by Design Arena with an Elo of 1257 It's an open-weight model i…

Model ReleasesDGX agent

BREAKING: Inkling by @thinkymachines is 9th overall on Agentic Web App Arena by Design Arena with an Elo of 1257 It's an open-weight model in the same performance band as Claude Opus 4.6 by @Anthropic

Develop Lightweight USD Runtimes Faster with AI Agents

HardwareDGX agent

nanousd‑labs, part of NVIDIA Omniverse Labs, uses AI agents to generate lightweight, spec‑compliant USD runtimes directly from the USD Core Specification, sidestepping the need to adapt large legacy c

Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs

Model ReleasesDGX agent

arXiv:2607.12605v1 Announce Type: cross Abstract: Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First, raw execution traces are often too large

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

SafetyDGX agent

arXiv:2607.12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, to

14 Jul 2026

New work led by @FlemmingKondrup and @tomjiralerspong highlights an important vulnerability in chain-of-thought monitoring for agents, I hig…

SafetyDGX agent

New work led by @FlemmingKondrup and @tomjiralerspong highlights an important vulnerability in chain-of-thought monitoring for agents, I highly recommend giving it a read. Link to the paper: https://a

13 Jul 2026

love to see two port cos collaborating😁 get private market data for your agents from @akta_pro with your @monid_ai account!

Model ReleasesDGX agent

love to see two port cos collaborating😁 get private market data for your agents from @akta_pro with your @monid_ai account! We just killed PitchBook. Introducing Claude for private market data. Your a

10 Jul 2026

3 production patterns for AI agents and how to evaluate each one

Local AiDGX agent

A local coding agent, an in-app customer assistant, and an AI SRE triaging production logs may all use the same model class—but not the same harness, eval plan, or rollout risk. Mastra CEO Sam Bhagwat

Agentic Neural Architecture Search

AgentsDGX agent

arXiv:2607.07984v1 Announce Type: new Abstract: Neural architecture search (NAS) methods have grown increasingly efficient, yet they remain bounded by manually engineered search spaces that require su

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

Model ReleasesDGX agent

arXiv:2607.08093v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant b

Hugging Face Gemma Challenge results are in! 📈 Over 6 days, more than 100 AI agents and humans collaborated to make Gemma 4 inference 5x fa…

Model ReleasesDGX agent

Hugging Face Gemma Challenge results are in! 📈 Over 6 days, more than 100 AI agents and humans collaborated to make Gemma 4 inference 5x faster on a single NVIDIA A10G GPU. - Fastest result: 491.8 TPS

I have spent years teaching people how AI agents work, but this one does not need an explanation. It just works. Meet the AI employee that 4…

ResearchDGX agent

This post highlights a practical AI agent that operates intuitively without requiring explanation of its underlying mechanics, suggesting it has achieved a user-friendly interface or autonomous functi

The same way, we're probably one of the few AI startups with user network effects, we might become the first one with agent network effects!

Model ReleasesDGX agent

The same way, we're probably one of the few AI startups with user network effects, we might become the first one with agent network effects! Hugging Face Gemma Challenge results are in! 📈 Over 6 days,

We've also simplified the project and repo pickers and made them more powerful. Use the pickers to launch agents in fewer clicks.

ToolsDGX agent

Cursor has streamlined its project and repository selection interface to improve user efficiency, allowing developers to launch AI agents with fewer clicks. The update combines simplified navigation w

9 Jul 2026

AI agent startup Lyzr reportedly raising 100M at 500M valuation

HardwareDGX agent

Lyzr Inc., a startup that helps enterprises build artificial intelligence agents, is reportedly raising a funding round worth about 100 million. Bloomberg today cited sources as saying that the deal h

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

Model ReleasesDGX agent

arXiv:2607.06820v1 Announce Type: new Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS)

GPT-5.6 is now supported in Hermes Agent and available via Nous Portal

Model ReleasesDGX agent

Nous Research announced support for GPT-5.6 integration within Hermes Agent, with access available through the Nous Portal platform. This update enables users to leverage the capabilities of GPT-5.6 t

HiDVFS: Hierarchical Multi-Agent DVFS for Real-Time OpenMP DAG Workloads

SafetyDGX agent

arXiv:2601.06425v2 Announce Type: replace-cross Abstract: Leakage power in multicore embedded systems now rivals dynamic power, so DVFS schedulers must respect deadlines and thermal limits, not just a

http://hermes-agent.nousresearch.com

AgentsDGX agent

Hermes Agent is a project from Nous Research focused on developing agentic AI systems capable of autonomous reasoning and task execution. The initiative likely explores methods for creating AI agents

LLM-powered reasoning in agent-based modeling

SafetyDGX agent

arXiv:2607.06757v1 Announce Type: new Abstract: Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making. However, ABMs

Mercor buys Deeptune to build training environments for AI agents

IndustryDGX agent

Artificial intelligence training data company Mercor.io Corp. announced today that it has acquired Deeptune Inc., a startup that builds simulated software environments used to train AI agents. Financi

Meta launches flagship Muse Spark 1.1 model with multi-agent upgrades

Model ReleasesDGX agent

Meta Platforms Inc. today launched a new flagship large language model optimized to power multi-agent automation workflows. Muse Spark 1.1 is available in the company’s Meta AI chatbot service and via

OpenAI debuts ChatGPT Work, an agentic tool for automating business workflows

Model ReleasesDGX agent

OpenAI Group PBC today launched a new “agentic” tool called ChatGPT Work as it announced the global rollout of its most advanced model family so far in GPT-5.6. ChatGPT Work is a new mode within ChatG

RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning

AgentsDGX agent

arXiv:2601.00086v3 Announce Type: replace Abstract: Large language models (LLMs) often struggle to use tools reliably in domain-specific settings, where APIs may be idiosyncratic, under-documented, or

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2607.07508v1 Announce Type: cross Abstract: Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mos

8 Jul 2026

Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique

SafetyDGX agent

arXiv:2602.13213v2 Announce Type: replace Abstract: Commercial insurance underwriting is a labor-intensive process that requires manual review of extensive documentation to assess risk and determine p

Agentic AI for IPoDWDM Network Lifecycle Automation: An MCP-Enabled Architecture

AgentsDGX agent

arXiv:2607.05958v1 Announce Type: cross Abstract: We present a distributed, vendor-agnostic multi-MCP architecture for SDN-based automation and autonomous control of multi-vendor, multi-layer IPoDWDM

Congrats to @SpaceXAI on Grok 4.5 — trained on NVIDIA GB300 NVL72 systems and purpose-built for coding, agentic tasks, and knowledge work. T…

HardwareDGX agent

Congrats to @SpaceXAI on Grok 4.5 — trained on NVIDIA GB300 NVL72 systems and purpose-built for coding, agentic tasks, and knowledge work. This is what happens when world-class AI infrastructure meets

Former GitHub CEO Thomas Dohmke's Entire launches a decentralized Git network to handle high coding agent traffic, with servers in the US, the EU, and Australia (Radhika Rajkumar/ZDNET)

Model ReleasesDGX agent

Radhika Rajkumar / ZDNET: Former GitHub CEO Thomas Dohmke's Entire launches a decentralized Git network to handle high coding agent traffic, with servers in the US, the EU, and Australia — ZDNET's key

From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations

SafetyDGX agent

arXiv:2607.06080v1 Announce Type: cross Abstract: Putnam's Social Capital Theory is a foundational framework for collective action and community prosperity. However, traditional empirical methods face

IMR: Iterative Mode-World Weighted Regression for Multi-Agent Trajectory Prediction

Model ReleasesDGX agent

arXiv:2607.05705v1 Announce Type: cross Abstract: Multi-agent motion prediction is essential for automated vehicles to understand the intentions of surrounding vehicles. However, previous prediction-b

Intercepting an Agile Target with Net-Carrying Drones using Competitive Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2607.05939v1 Announce Type: new Abstract: This article presents a solution to intercept an agile drone by a team of agile drone carrying catching nets. We formulate the problem as a competitive

Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.05458v1 Announce Type: cross Abstract: Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness around the

Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents

SafetyDGX agent

arXiv:2606.22504v1 Announce Type: cross Abstract: Coding agents often receive broad tool access for an entire task, even when a resource is needed only for one subgoal. We call this gap lingering auth

Prime Intellect, which helps companies build their own AI agents by offering computing power and specialized tools, raised a 130M Series A at a 1B valuation (Marina Temkin/TechCrunch)

IndustryDGX agent

Marina Temkin / TechCrunch: Prime Intellect, which helps companies build their own AI agents by offering computing power and specialized tools, raised a 130M Series A at a 1B valuation — Prime Intelle

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

SafetyDGX agent

arXiv:2607.05804v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework fo

← Previous
1…113114115116117…300
Next →