AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
23 Apr 2026

In the enterprise AI race, who is leading and who is just reacting?

AgentsDGX agent

Enterprise AI scaling is accelerating as organizations shift from experimentation to full deployment, embedding intelligence into core workflows. At the same time, agentic systems are driving a broade

Information Aggregation with AI Agents

ResearchDGX agent

arXiv:2604.20050v1 Announce Type: cross Abstract: Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing pr

OpenAI subscribers get new ‘workspace agents’ to automate complex tasks across teams

IndustryDGX agent

OpenAI Group PBC said today it’s pushing ChatGPT outside its usual chat interface with the launch of “workspace agents,” which is a new feature that allows business users to automate recurring tasks,

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ParseBench is now live on @Kaggle. The first document OCR benchmark built for AI agents — 2,000 enterprise pages, 167K+ test rules, 5 dimens…

Model ReleasesDGX agent

ParseBench is now live on @Kaggle. The first document OCR benchmark built for AI agents — 2,000 enterprise pages, 167K+ test rules, 5 dimensions that actually break downstream agents. Benchmark your p

Rilian raises $17.5M to automate security software procurement and deployment in the defense sector

AgentsDGX agent

A startup called Rilian that’s building agentic systems integration tools for companies operating in the defense and national security industries, said today it has raised 17.5 million in seed funding

Stateless Decision Memory for Enterprise AI Agents

ApplicationsDGX agent

arXiv:2604.20158v1 Announce Type: new Abstract: Enterprise deployment of long-horizon decision agents in regulated domains (underwriting, claims adjudication, tax examination) is dominated by retrieva

Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents

Model ReleasesDGX agent

arXiv:2601.20144v3 Announce Type: replace Abstract: Tool-calling agents are increasingly deployed in real-world customer-facing workflows. Yet most studies on tool-calling agents focus on idealized se

22 Apr 2026

AI scientists produce results without reasoning scientifically

AgentsDGX agent

arXiv:2604.18805v1 Announce Type: new Abstract: Large language model (LLM)-based systems are increasingly deployed to conduct scientific research autonomously, yet whether their reasoning adheres to t

Hugging Face Releases ml-intern: An Open-Source AI Agent that Automates the LLM Post-Training Workflow [The 'AI Intern' that actually ships …

Model ReleasesDGX agent

Hugging Face Releases ml-intern: An Open-Source AI Agent that Automates the LLM Post-Training Workflow [The 'AI Intern' that actually ships SOTA models ] This isn't just another ML Research Loop wrapp

Iterable launches Nova agent to assist markets in scaling customer personalization

Model ReleasesDGX agent

Iterable Inc., a customer engagement platform, today announced that it launched an artificial intelligence agent designed to help marketers keep customer interactions relevant as campaigns scale. The

NVIDIA and Google Cloud Collaborate to Advance Agentic and Physical AI

HardwareDGX agent

NVIDIA and Google Cloud have collaborated for more than a decade, co‑engineering a full‑stack AI platform that spans every technology layer — from performance‑optimized libraries and frameworks to ent

Owner-Harm: A Missing Threat Model for AI Agent Safety

Model ReleasesDGX agent

arXiv:2604.18658v1 Announce Type: cross Abstract: Existing AI agent safety benchmarks focus on generic criminal harm (cybercrime, harassment, weapon synthesis), leaving a systematic blind spot for a d

Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment

Model ReleasesDGX agent

arXiv:2604.19548v1 Announce Type: cross Abstract: Large Language Model agents have rapidly evolved from static text generators into dynamic systems capable of executing complex autonomous workflows. T

We are excited to launch VideoGameBench on Antim Labs, created by @a1zhang, Thomas L. Griffiths (@cocosci_lab), @karthik_r_n, and @OfirPress…

AgentsDGX agent

VideoGameBench is a new benchmark launched on Antim Labs, created by a1zhang, Thomas L. Griffiths, Karthik R. N, and Ofir Press. The benchmark likely evaluates AI model performance on video game-relat

What’s new in Cloud Run at Next ‘26

Model ReleasesDGX agent

From vibe-coded and large-scale apps to AI models and agents, Cloud Run delivers on-demand compute with zero overhead and pay-per-use pricing for all of your workloads. Last year, the number of extern

21 Apr 2026

BOIL: Learning Environment Personalized Information

AgentsDGX agent

arXiv:2604.17137v1 Announce Type: new Abstract: Navigating complex environments poses challenges for multi-agent systems, requiring efficient extraction of insights from limited information. In this p

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

Model ReleasesDGX agent

arXiv:2604.18543v1 Announce Type: cross Abstract: Constructing environments for training and evaluating claw-like agents remains a manual, human-intensive process that does not scale. We argue that wh

ComPASS: Towards Personalized Agentic Social Support via Tool-Augmented Companionship

Model ReleasesDGX agent

arXiv:2604.18356v1 Announce Type: new Abstract: Developing compassionate interactive systems requires agents to not only understand user emotions but also provide diverse, substantive support. While r

Ivanti extends Neurons platform with autonomous IT and security capabilities

AgentsDGX agent

Information technology security software company Ivanti Inc. today announced new capabilities that are focused on allowing autonomous IT operations and organizations to secure their environments more

On-Orbit Space AI: Federated, Multi-Agent, and Collaborative Algorithms for Satellite Constellations

SafetyDGX agent

arXiv:2604.16518v1 Announce Type: new Abstract: Satellite constellations are transforming space systems from isolated spacecraft into networked, software-defined platforms capable of on-orbit percepti

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning

Local AiDGX agent

arXiv:2601.15724v2 Announce Type: replace Abstract: Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning

What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations

AgentsDGX agent

arXiv:2510.17795v3 Announce Type: replace Abstract: Replicating AI research is a crucial yet challenging task for large language model (LLM) agents. Existing approaches often struggle to generate exec

20 Apr 2026

ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence

Model ReleasesDGX agent

arXiv:2603.24621v2 Announce Type: replace Abstract: We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents

ChemGraph-XANES: An Agentic Framework for XANES Simulation and Analysis

Model ReleasesDGX agent

arXiv:2604.16205v1 Announce Type: cross Abstract: Computational X-ray absorption near-edge structure (XANES) is widely used to probe local coordination environments, oxidation states, and electronic s

Classic study gave 146 economist teams the same dataset & got wildly different answers New paper reruns it with agentic AI. Claude Code & Co…

Model ReleasesDGX agent

Classic study gave 146 economist teams the same dataset & got wildly different answers New paper reruns it with agentic AI. Claude Code & Codex land near the human median, but with far tighter dispers

HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?

Model ReleasesDGX agent

arXiv:2604.15415v1 Announce Type: cross Abstract: Large language models (LLMs) have evolved into autonomous agents that rely on open skill ecosystems (e.g., ClawHub and Skills.Rest), hosting numerous

Preference Estimation via Opponent Modeling in Multi-Agent Negotiation

Model ReleasesDGX agent

arXiv:2604.15687v1 Announce Type: new Abstract: Automated negotiation in complex, multi-party and multi-issue settings critically depends on accurate opponent modeling. However, conventional numerical

19 Apr 2026

Headless everything for personal AI

AgentsDGX agent

Headless everything for personal AI Matt Webb thinks headless services are about to become much more common: Why? Because using personal AIs is a better experience for users than using services direct

18 Apr 2026

Hermes Agent is model & tool backend agnostic for a reason, everyone should have access to AI. We don't dictate the rules of use for your ag…

Model ReleasesDGX agent

Hermes Agent is model & tool backend agnostic for a reason, everyone should have access to AI. We don't dictate the rules of use for your agent, YOU do Anthropic shut down an entire company's Claude a

17 Apr 2026

@grok you go first

AgentsDGX agent

This post likely discusses Grok, an AI assistant developed by xAI, possibly exploring its capabilities, performance, or a specific interaction with the system. Given the casual phrasing and that it's

Through the end of this weekend, we are doubling Composer 2 usage limits inside of Cursor's new agents window. Enjoy!

ToolsDGX agent

Cursor is temporarily doubling the usage limits for Composer 2 through the end of the weekend, specifically within Cursor's new agents window feature. This promotion allows users to access increased c

16 Apr 2026

A real issue with the current state of our knowledge on the work implications of AI is that there was a genuine discontinuity in AI ability …

AgentsDGX agent

A real issue with the current state of our knowledge on the work implications of AI is that there was a genuine discontinuity in AI ability with the rise of practical agentic systems in 2026. We were

Coding agents learn from experience, but that knowledge stays locked in silos. Solve a thousand SWE tasks, and none of that wisdom helps wit…

TutorialsDGX agent

Coding agents learn from experience, but that knowledge stays locked in silos. Solve a thousand SWE tasks, and none of that wisdom helps with competitive coding. What if memories could transfer across

RiskWebWorld: A Realistic Interactive Benchmark for GUI Agents in E-commerce Risk Management

Model ReleasesDGX agent

arXiv:2604.13531v1 Announce Type: cross Abstract: Graphical User Interface (GUI) agents show strong capabilities for automating web tasks, but existing interactive benchmarks primarily target benign,

Why Your Agents Can’t Read Enterprise Documents — and How to Fix It

TutorialsDGX agent

AI agents struggle to effectively process and extract information from complex enterprise documents due to limitations in context windows, reasoning capabilities, and handling of unstructured data for

15 Apr 2026

AlphaEval: Evaluating Agents in Production

Model ReleasesDGX agent

arXiv:2604.12162v1 Announce Type: new Abstract: The rapid deployment of AI agents in commercial settings has outpaced the development of evaluation methodologies that reflect production realities. Exi

How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm

Model ReleasesDGX agent

arXiv:2604.12250v1 Announce Type: new Abstract: This study examines how model-specific characteristics of Large Language Model (LLM) agents, including internal alignment, shape the effect of memory on

Register domains wherever you build: Cloudflare Registrar API now in beta

AgentsDGX agent

The Cloudflare Registrar API is now in beta. Developers and AI agents can search, check availability, and register domains at cost directly from their editor, their terminal, or their agent — without

SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents

Model ReleasesDGX agent

arXiv:2604.12040v1 Announce Type: cross Abstract: We present SIR-Bench, a benchmark of 794 test cases for evaluating autonomous security incident response agents that distinguishes genuine forensic in

🆕 The Full Story of Notion AI https://latent.space/p/notion We're so excited to chat with @simonlast and @sarahmsachs about Notion's 'Token…

HardwareDGX agent

🆕 The Full Story of Notion AI https://latent.space/p/notion We're so excited to chat with @simonlast and @sarahmsachs about Notion's 'Token Town' - the crack team of AI Engineers and Model Behavior En

Transferable Expertise for Autonomous Agents via Real-World Case-Based Learning

Model ReleasesDGX agent

arXiv:2604.12717v1 Announce Type: new Abstract: LLM-based autonomous agents perform well on general reasoning tasks but still struggle to reliably use task structure, key constraints, and prior experi

VULCAN: Vision-Language-Model Enhanced Multi-Agent Cooperative Navigation for Indoor Fire-Disaster Response

Model ReleasesDGX agent

arXiv:2604.12831v1 Announce Type: new Abstract: Indoor fire disasters pose severe challenges to autonomous search and rescue due to dense smoke, high temperatures, and dynamically evolving indoor envi

When to Forget: A Memory Governance Primitive

AgentsDGX agent

arXiv:2604.12007v1 Announce Type: new Abstract: Agent memory systems accumulate experience but currently lack a principled operational metric for memory quality governance -- deciding which memories t

14 Apr 2026

AIRA_2: Overcoming Bottlenecks in AI Research Agents

HardwareDGX agent

arXiv:2603.26499v2 Announce Type: replace Abstract: Existing research has identified three structural performance bottlenecks in AI research agents: (1) synchronous single-GPU execution constrains sam

BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows

Model ReleasesDGX agent

arXiv:2604.11304v1 Announce Type: new Abstract: Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-

ClawVM: Harness-Managed Virtual Memory for Stateful Tool-Using LLM Agents

Model ReleasesDGX agent

arXiv:2604.10352v1 Announce Type: new Abstract: Stateful tool-using LLM agents treat the context window as working memory, yet today's agent harnesses manage residency and durability as best-effort, c

CodeComp: Structural KV Cache Compression for Agentic Coding

Local AiDGX agent

arXiv:2604.10235v1 Announce Type: new Abstract: Agentic code tasks such as fault localization and patch generation require processing long codebases under tight memory constraints, where the Key-Value

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

Model ReleasesDGX agent

arXiv:2604.01687v2 Announce Type: replace Abstract: Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool

CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance

AgentsDGX agent

arXiv:2508.16644v4 Announce Type: replace Abstract: Diffusion models excel at photorealistic synthesis but struggle with precise object counts, especially in high-density settings. We introduce COUNTL

Detecting Safety Violations Across Many Agent Traces

Model ReleasesDGX agent

arXiv:2604.11806v1 Announce Type: new Abstract: To identify safety violations, auditors often search over large sets of agent traces. This search is difficult because failures are often rare, complex,

EE-MCP: Self-Evolving MCP-GUI Agents via Automated Environment Generation and Experience Learning

SafetyDGX agent

arXiv:2604.09815v1 Announce Type: new Abstract: Computer-use agents that combine GUI interaction with structured API calls via the Model Context Protocol (MCP) show promise for automating software tas

HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks

Model ReleasesDGX agent

arXiv:2604.09937v1 Announce Type: new Abstract: Healthcare administration accounts for over $1 trillion in annual spending, making it a promising target for LLM-based computer-use agents (CUAs). While

NetAgentBench: A State-Centric Benchmark for Evaluating Agentic Network Configuration

Model ReleasesDGX agent

arXiv:2604.09678v1 Announce Type: cross Abstract: As agentic network management gains popularity, there is a critical need for evaluation frameworks that transcend static, one-shot testing. To address

Scaling MCP adoption: Our reference architecture for simpler, safer and cheaper enterprise deployments of MCP

AgentsDGX agent

We share Cloudflare's internal strategy for governing MCP using Access, AI Gateway, and MCP server portals. We also launch Code Mode to slash token costs and recommend new rules for detecting Shadow M

The AI-first workday is here, but is the data layer ready to handle it?

AgentsDGX agent

As “deploy fast” meets enterprise reality, data governance has quickly emerged as the bottleneck between AI ambition and outcomes. The disconnect is measurable. In their Agentic AI Study, Qlik Technol

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse

HardwareDGX agent

arXiv:2511.00413v4 Announce Type: replace Abstract: Agentic large language model (LLM) training often involves multi-turn interaction trajectories that branch into multiple execution paths due to conc

Working Paper: Towards Schema-based Learning from a Category-Theoretic Perspective

AgentsDGX agent

arXiv:2604.10589v1 Announce Type: new Abstract: We introduce a hierarchical categorical framework for Schema-Based Learning (SBL) structured across four interconnected levels. At the schema level, a f

13 Apr 2026

ActionNex: A Virtual Outage Manager for Cloud Computing

AgentsDGX agent

arXiv:2604.03512v2 Announce Type: replace Abstract: Outage management in large-scale cloud operations remains heavily manual, requiring rapid triage, cross-team coordination, and experience-driven dec

Adaptive Tuning of Parameterized Traffic Controllers via Multi-Agent Reinforcement Learning

Model ReleasesDGX agent

arXiv:2512.07417v2 Announce Type: replace Abstract: Effective traffic control is essential for mitigating congestion in transportation networks. Conventional traffic management strategies, including r

AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society

SafetyDGX agent

arXiv:2502.08691v2 Announce Type: replace-cross Abstract: Understanding human behavior and society is a central focus in social sciences, with the rise of generative social science marking a significa

← Previous
1…99100101102103…300
Next →