AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

agents

GridTimelineEvolution
7,214 results
29 May 2026

feel very aligned with our vision & Ronak + the awesome Trajectory team’s on practically tackling Continual Learning at scale 🚀 there’s a v…

AgentsDGX agent

feel very aligned with our vision & Ronak + the awesome Trajectory team’s on practically tackling Continual Learning at scale 🚀 there’s a very good reason why teams are partly building “Observability

FLIP: Real-Time and Resilient Formation Planning for Large-Scale DIstributed Swarms via Point Cloud Registration

AgentsDGX agent

arXiv:2605.29704v1 Announce Type: new Abstract: Traditional large-scale formation planning either oversimplify the formation representation which leads to poor performance, or they employ complete col

Formalizing Mathematics at Scale

AgentsDGX agent

arXiv:2605.29955v1 Announce Type: new Abstract: We present AutoformBot, a multi-agent system for building an Autoformalized Textbook Library At Scale (Atlas) in Lean 4. AutoformBot orchestrates thousa


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

From Blind Guess to Informed Judgment: Teaching LLMs to Evaluate Materials by Building Knowledge-Augmented Preference Signals

AgentsDGX agent

arXiv:2605.29555v1 Announce Type: new Abstract: As candidate generation and high-throughput experimentation advance, the primary bottleneck in materials discovery is shifting from property prediction

From Prompts to Context: An Ontology-Driven Framework for Human-Generative AI Collaboration

AgentsDGX agent

arXiv:2605.29675v1 Announce Type: cross Abstract: Collaborations with Generative AI often begin with a short prompt and end with an opaque output, leaving implicit who was involved, what task was bein

FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views

AgentsDGX agent

arXiv:2605.29997v1 Announce Type: new Abstract: We present FRUC, a feed-forward 3D Gaussian splatting framework for dynamic scene reconstruction from uncalibrated collaborative driving views. Existing

GenClaw: Code-Driven Agentic Image Generation

AgentsDGX agent

arXiv:2605.30248v1 Announce Type: new Abstract: Image generation models have evolved from text-conditioned pixel synthesis toward multimodal agents endowed with visual comprehension and tool invocatio

GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling

AgentsDGX agent

arXiv:2605.28835v1 Announce Type: cross Abstract: Large Language Models (LLMs) extend their capabilities through function-calling (FC), which relies on training data with high quality, diversity, and

Governing Technical Debt in Agentic AI Systems

AgentsDGX agent

arXiv:2605.29129v1 Announce Type: new Abstract: Agentic AI systems are increasingly being explored as production infrastructure: they reason over multiple steps, call tools, act through workflows, and

grok-build-0.1 is now available via the xAI API in public beta. This is the same model that powers the Grok Build CLI and excels at agentic …

AgentsDGX agent

grok-build-0.1 is now available via the xAI API in public beta. This is the same model that powers the Grok Build CLI and excels at agentic coding. Priced at 1/m input and 2/m output, it’s extremely c

Had a blast yesterday attending at @techeurope_'s Applied AI Conference in Berlin! I had a talk about building document agents and agentic d…

AgentsDGX agent

Had a blast yesterday attending at @techeurope_'s Applied AI Conference in Berlin! I had a talk about building document agents and agentic development in general, that you can find here: https://astra

Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills

AgentsDGX agent

arXiv:2605.29354v1 Announce Type: cross Abstract: LLM-powered coding agents increasingly participate in software development workflows by generating code, selecting dependencies, and producing package

Hermes Agent now has Tool Search, so your agent only loads what it needs

AgentsDGX agent

Hermes Agent has been updated with a Tool Search feature that enables agents to dynamically identify and load only the tools necessary for a given task, rather than loading all available tools upfront

Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction

AgentsDGX agent

arXiv:2605.29960v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly leverage long term memory to support persistent and autonomous task execution. However, this capability

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

AgentsDGX agent

arXiv:2605.29963v1 Announce Type: cross Abstract: Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation

How Consistent Are LLM Agents? Measuring Behavioral Reproducibility in Multi-Step Tool-Calling Pipelines

AgentsDGX agent

arXiv:2605.28840v1 Announce Type: cross Abstract: Large language model (LLM) agents with tool-calling capabilities are increasingly deployed in production systems, yet a fundamental reliability questi

How to build a better agent harness with traces and evals

AgentsDGX agent

Agents are easy to prototype and hard to improve. A repeatable loop of traces, evals, failed-span inspection, and targeted harness changes makes agent behavior easier to debug and improve. The post Ho

How to Relieve Distribution Shifts in Semantic Segmentation for Off-Road Environments

AgentsDGX agent

arXiv:2605.29599v1 Announce Type: cross Abstract: Semantic segmentation is crucial for autonomous navigation in off-road environments, enabling precise classification of surroundings to identify trave

https://x.com/huntlovell/status/2060399973506924612

AgentsDGX agent

I cannot provide an accurate summary because the URL appears to be invalid or the tweet is no longer accessible (the status ID seems implausible for the current date). To create a reliable knowledge b

Human-in-the-Loop Swarms: A Bionic Swarm Approach to Real-World Soil Mapping

AgentsDGX agent

arXiv:2605.29091v1 Announce Type: new Abstract: Swarm and field robotics face significant barriers to real-world validation due to the high cost and development time to deploy hardware. This paper int

I had the experience of playing against Sony AI’s “Project Ace,” the most advanced high-speed autonomous table tennis robot system, which ha…

AgentsDGX agent

I had the experience of playing against Sony AI’s “Project Ace,” the most advanced high-speed autonomous table tennis robot system, which has defeated elite human athletes. I managed to win a point. F

i had to pack away my coding agents on monday cuz I knew I’d stay up too late if I played w activegraph during the week (yay, it’s Friday!) …

AgentsDGX agent

i had to pack away my coding agents on monday cuz I knew I’d stay up too late if I played w activegraph during the week (yay, it’s Friday!) gautham kept playing with it and just showed me a custom UI

I literally haven’t typed anything in weeks Since I started using Grok’s speech-to-text in Hermes Agent... I just talk It transcribes everyt…

AgentsDGX agent

I literally haven’t typed anything in weeks Since I started using Grok’s speech-to-text in Hermes Agent... I just talk It transcribes everything perfectly Every word. Every single time My thoughts flo

Improving agents The old way: Manually reading traces, looking for patterns, writing evals, and creating fixes. The better way: Letting Lang…

AgentsDGX agent

This post from LangChain's Harrison Chase contrasts traditional manual methods of improving AI agents (tracing execution, identifying patterns, writing evaluations, and implementing fixes) with a more

Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models

AgentsDGX agent

arXiv:2605.29625v1 Announce Type: new Abstract: The topic of Co-creation, i.e., AI agents interacting with humans to generate outputs (e.g., art), has gained significant attention recently. However, m

I've been using state-of-the-art models to teach small models running on my computer how I work. The result : a personal agent that runs my …

AgentsDGX agent

I've been using state-of-the-art models to teach small models running on my computer how I work. The result : a personal agent that runs my inbox, my deal pipeline, my blog, my calendar, & my research

@jerryjliu0 We also automatically update the ParseBench leaderboard on Kaggle :) https://www.kaggle.com/benchmarks/llamaindex-org/parsebench

AgentsDGX agent

Jerry Liu announces that the ParseBench leaderboard is automatically updated on Kaggle, providing a continuously maintained benchmark for parsing performance metrics. The leaderboard is hosted under t

KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning

AgentsDGX agent

arXiv:2605.30002v1 Announce Type: new Abstract: Cross-domain multimodal time series forecasting is a challenging task, requiring models to integrate precise numerical comprehension, cross-domain seman

LangSmith LLM Gateway lets you enforce spend limits and redacts PII before requests reach the model. Not after the fact.

AgentsDGX agent

LangSmith LLM Gateway is a feature that enables proactive cost and privacy controls by enforcing spending limits and redacting personally identifiable information (PII) before API requests are sent to

llm spend is starting to get really high... a key part of our LLM gateway is spend visibility and spend control sign up for the private beta…

AgentsDGX agent

llm spend is starting to get really high... a key part of our LLM gateway is spend visibility and spend control sign up for the private beta to try it out today! https://www.langchain.com/langsmith-ll

make something agents want

AgentsDGX agent

make something agents want studying ActiveGraph by @yoheinakajima and apart from the concept the implementation itself is brilliant on the site it shares a prompt that will make my agent study docs, i

mcp-proto-okn: Natural-language access to open scientific knowledge graphs through the Model Context Protocol

AgentsDGX agent

arXiv:2605.30283v1 Announce Type: new Abstract: MCP Server Proto-OKN (mcp-proto-okn) is a Python-based Model Context Protocol server that enables AI assistants to discover, inspect, query and integrat

MemCollab: Cross-Model Memory Collaboration via Contrastive Trajectory Distillation

AgentsDGX agent

arXiv:2603.23234v2 Announce Type: replace Abstract: LLM agents increasingly rely on memory mechanisms to reuse knowledge from past problem-solving experiences. However, existing methods typically cons

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs

AgentsDGX agent

arXiv:2605.29512v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as interactive agents, yet their capacity for social and strategic reasoning over extended intera

Molecular Lead Optimization via Agentic Tool Planning

AgentsDGX agent

arXiv:2605.28862v1 Announce Type: new Abstract: Drug discovery is a lengthy and resource-intensive process composed of multiple stages. Among these stages, lead optimization plays a critical role in t

MOOSE-Copilot: A Web-Based Interactive Assistant for Unified Exploratory and Fine-Grained Scientific Hypothesis Discovery

AgentsDGX agent

arXiv:2605.29475v1 Announce Type: cross Abstract: Large language models (LLMs) show remarkable potential in scientific hypothesis discovery. However, existing approaches face two critical limitations:

Most people training agentic LLMs with RL right now have a silently broken training loop and have no idea. Here's the trap: single-turn RL w…

AgentsDGX agent

Most people training agentic LLMs with RL right now have a silently broken training loop and have no idea. Here's the trap: single-turn RL works beautifully. Clean curves, sane rewards, everything con

Network Optimization Aspects of Autonomous Vehicles: Challenges and Future Directions

AgentsDGX agent

arXiv:2605.29518v1 Announce Type: cross Abstract: Global megatrends, such as urbanization, population growth, and emerging network solutions are accelerating the development of the Connected and Auton

No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand

AgentsDGX agent

arXiv:2605.28836v1 Announce Type: cross Abstract: The Plain Writing Act in the United States requires government documents to be accessible in clear and simple language that the general public can eas

nobody told me Hermes Agent could just... join your Discord VC and talk back for those using Discord w/ Hermes Agent theres a feature where …

AgentsDGX agent

nobody told me Hermes Agent could just... join your Discord VC and talk back for those using Discord w/ Hermes Agent theres a feature where you can just have your Hermes agent jump in on a vc call wit

Offloading Score: Measuring AI Reliance Through Counterfactual Workflows

AgentsDGX agent

arXiv:2605.29392v1 Announce Type: cross Abstract: AI tools are increasingly integrated into real-world workflows. However, existing measures of reliance on these tools focus on AI output adoption or o

On Distributional Reinforcement Learning in Chaotic Dynamical Systems

AgentsDGX agent

arXiv:2605.30160v1 Announce Type: cross Abstract: Chaotic dynamical systems pose a fundamental challenge for Reinforcement Learning (RL): exponential sensitivity to initial conditions induces high-var

open models are having a moment!

AgentsDGX agent

open models are having a moment! The latest finding in the LangSmith Signal: Open Models are having a moment. 1 in 3 AI teams ran an open-weights model in April 2026, up from 1 in 5 nine months ago. T

open models will become a huge force for coding.

AgentsDGX agent

open models will become a huge force for coding. The latest finding in the LangSmith Signal: Open Models are having a moment. 1 in 3 AI teams ran an open-weights model in April 2026, up from 1 in 5 ni

Open Source Browser Agent That Learns and Repeats Workflows

AgentsDGX agent

An open-source browser with built-in AI agents that emphasizes privacy and automation, enabling task automation through natural language without coding. The browser supports multiple AI providers incl

Opus 4.8 dropped today. ParseBench results are out. ✅ Slight gains: tables, semantic formatting, layout ⚠️ Slight regressions: charts, conte…

AgentsDGX agent

Opus 4.8 dropped today. ParseBench results are out. ✅ Slight gains: tables, semantic formatting, layout ⚠️ Slight regressions: charts, content faithfulness 💰 Slight price/page increase Lots of alpha l

Our users love @StepFun_ai models and this new release packs a punch at a small size. Looking forward to seeing how well it works with Herme…

AgentsDGX agent

Our users love @StepFun_ai models and this new release packs a punch at a small size. Looking forward to seeing how well it works with Hermes Agent! ⚡️ Step 3.7 Flash is here: The new frontier is agen

Parse PDFs at lightspeed (this video is at 1x) Absolute cinema

AgentsDGX agent

Parse PDFs at lightspeed (this video is at 1x) Absolute cinema Media We've created the world's fastest PDF parser ⚡️ And it's more accurate than any other open-source, model-free PDF parser out there

PatchBoard: Schema-Grounded State Mutation for Reliable and Auditable LLM Multi-Agent Collaboration

AgentsDGX agent

arXiv:2605.29313v1 Announce Type: new Abstract: LLM multi-agent systems often coordinate through natural-language dialogue or loosely structured shared memory, making intermediate state difficult to v

PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions

AgentsDGX agent

arXiv:2605.30268v1 Announce Type: cross Abstract: We address the task of generating physically accurate and visually faithful 4D Human-Object Interaction (HOI). Given a static 3D human and target obje

Production traffic from frontier models is a golden data asset. If you can efficiently mine the traces, filter for quality, and fine-tune sm…

AgentsDGX agent

Production traffic from frontier models is a golden data asset. If you can efficiently mine the traces, filter for quality, and fine-tune smaller models on them, you get specialized performance at a f

Quality went up alongside output. Even with more PRs shipping, total incidents dropped 5%. They built security guardrails and quality standa…

AgentsDGX agent

Quality went up alongside output. Even with more PRs shipping, total incidents dropped 5%. They built security guardrails and quality standards into the agentic workflow itself. Productivity vs qualit

Real-rootedness of the Poincare polynomials of overline{mathcal M}_{0,n}: an AI-assisted proof

AgentsDGX agent

arXiv:2605.29151v1 Announce Type: cross Abstract: We prove real-rootedness for the Poincare polynomial [ P_n(t)=sum_{i=0}^{n-3} im H^{2i}(overline{mathcal M}_{0,n};Q)t^i ] of the Deligne--Mumford modu

Revisiting Observation Reduction for Web Agents: Comprehensive Evaluation with a Lightweight Framework

AgentsDGX agent

arXiv:2605.29397v1 Announce Type: new Abstract: HTML observations in LLM-based web agents are extremely long, and while many reduction methods have been proposed, it remains unclear which methods redu

RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models

AgentsDGX agent

arXiv:2603.18859v2 Announce Type: replace Abstract: Reinforcement learning (RL) shows promise for enhancing LLM agentic reasoning, yet sparse terminal rewards hinder fine-grained optimization. Process

// Scaling Laws for Agent Harnesses // If you build agent harnesses, this one is worth your time. (bookmark it) Most harness tuning treats e…

AgentsDGX agent

// Scaling Laws for Agent Harnesses // If you build agent harnesses, this one is worth your time. (bookmark it) Most harness tuning treats every token and tool call as if volume is all that counts. Ne

Scaling Small Agents Through Strategy Auctions

AgentsDGX agent

arXiv:2602.02751v2 Announce Type: replace-cross Abstract: Small language models are increasingly viewed as a promising, cost-effective approach to agentic AI, with proponents claiming they are suffici

SchGen: PCB Schematic Generation with Semantic-Grounded Code Representations

AgentsDGX agent

arXiv:2605.30345v1 Announce Type: new Abstract: Printed circuit board (PCB) schematic design defines nearly all electronic hardware, but it remains manual and expertise-intensive. While generative AI

SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?

AgentsDGX agent

arXiv:2605.30104v1 Announce Type: new Abstract: Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot re

self-modification

AgentsDGX agent

Self-modification in AI systems refers to the capability of an artificial intelligence to alter its own code, parameters, or behavior patterns without external intervention. This concept, discussed by

← Previous
1…5556575859…121
Next →