AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

agents

GridTimelineEvolution
7,214 results
1 Jun 2026

Fighting Numerical Hallucinations via Data-centric Compilation for Online Financial QA

AgentsDGX agent

arXiv:2605.31064v1 Announce Type: cross Abstract: Large Language Models (LLMs) have significantly advanced online data services, particularly in the domain of financial question answering (FinQA). How

Fleet computer use is now available in LangSmith's APAC instance! You can now give your Fleet agents access to a virtual computer if you're …

AgentsDGX agent

LangSmith's APAC instance now supports Fleet computer use, enabling Fleet agents to access virtual computers for enhanced functionality. This feature allows users in the Asia-Pacific region to leverag

Generalized Intention Modeling in Multi-Agent Reinforcement Learning

AgentsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.31318v1 Announce Type: new Abstract: Modeling an opponent's intent is critical for effective decision-making in non-cooperative, competitive, and general-sum multi-agent reinforcement learn

great look into how Rippling built RipplingAI

AgentsDGX agent

great look into how Rippling built RipplingAI .@Rippling AI runs on Deep Agents and LangSmith. Here’s how they shipped to millions of users in 6 months. https://www.langchain.com/blog/how-rippling-wen

HADT: A Heterogeneous Multi-Agent Differential Transformer for Autonomous Earth Observation Satellite Cluster

AgentsDGX agent

arXiv:2605.31023v1 Announce Type: new Abstract: This work addresses the problem of autonomous resource management in heterogeneous satellite cluster conducting Earth Observation (EO) missions includin

have manually read 1000s traces since joining LangChain! it’s a great way to learn and understand your agent but completely infeasible to do…

AgentsDGX agent

have manually read 1000s traces since joining LangChain! it’s a great way to learn and understand your agent but completely infeasible to do at agent scale 🙃 engine helps us automate that process so h

Hermes Agent is now natively supported on @Windows

AgentsDGX agent

Nous Research announced native support for Hermes Agent on Windows, expanding the availability of their Hermes model to Windows-based systems. This development enables Windows users to run Hermes Agen

How Hermes implements an open source agent harness architecture

AgentsDGX agent

Hermes from NousResearch is a strong open-source agent harness. This post examines how its runtime loop, context management, tool scoping, session infrastructure, and orchestration patterns map to a m

HypoAgent: An Agentic Framework for Interactive Abductive Hypothesis Generation over Knowledge Graphs

AgentsDGX agent

arXiv:2605.31370v1 Announce Type: new Abstract: Abductive reasoning over knowledge graphs aims to generate logical hypotheses that explain observed entities or facts. Existing controllable hypothesis

i know it's backwards order, but experiment #2: https://x.com/ActiveGraphAI/status/2061106299748909090?s=20

AgentsDGX agent

i know it's backwards order, but experiment #2: https://x.com/ActiveGraphAI/status/2061106299748909090?s=20 [New Technical Blog Post] In our second longmemeval experiment, we introduce semantic ingest

IAF-Net: Illumination-Adaptive Fusion for Low-Light Urban Road Segmentation

AgentsDGX agent

arXiv:2605.30939v1 Announce Type: new Abstract: Semantic road segmentation is important for autonomous driving, but existing methods suffer severe performance degradation under low-light conditions. M

IDOL: Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous Driving

AgentsDGX agent

arXiv:2605.31476v1 Announce Type: new Abstract: End-to-end autonomous driving has emerged as a compelling paradigm for learning planning directly from sensor observations, while recent world-model-bas

If LLMs Have Human-Like Attributes, Then So Does Age of Empires II

AgentsDGX agent

arXiv:2605.31514v1 Announce Type: cross Abstract: Much research has been carried out on large language models (LLMs) and LLM-powered agentic workflows. However, many works within the field state emerg

if you build any agent on activegraph, the trace is automatic and first-class, not bolted on

AgentsDGX agent

if you build any agent on activegraph, the trace is automatic and first-class, not bolted on a parallel experiment building a coding agent on top of @activegraphai. you can see everything flattened do

In @latentspacepod podcast, I shared my view on video generation, world models, LLMs, agents, continual learning and where the next frontier…

AgentsDGX agent

In @latentspacepod podcast, I shared my view on video generation, world models, LLMs, agents, continual learning and where the next frontier is. 1. Video models get most of their intelligence from lan

Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation

AgentsDGX agent

arXiv:2605.31278v1 Announce Type: new Abstract: Reliable evaluation of agentic systems requires unbiased estimates with valid uncertainty, but standard practice navigates between costly human annotati

Introducing Search as Code, our new search architecture for AI agents. It writes Python that calls our search stack directly, instead of loo…

AgentsDGX agent

Introducing Search as Code, our new search architecture for AI agents. It writes Python that calls our search stack directly, instead of looping through function calls one at a time. Available in the

Investigating Detection and Obfuscation of Prompt Injection Attacks Against Software Reverse Engineering AI Agents

AgentsDGX agent

arXiv:2605.30677v1 Announce Type: cross Abstract: Agentic software reverse engineering systems are vulnerable to prompt injection attacks placed into the source code of executable binary files. This r

Join us, @mercor_ai, @Etched, and @AnthropicAI for a one-day hackathon in SF with a $50k top prize. Registrations close on 6/12. We can't wa…

AgentsDGX agent

Join us, @mercor_ai, @Etched, and @AnthropicAI for a one-day hackathon in SF with a 50k top prize. Registrations close on 6/12. We can't wait what to see you build! We're running a 24-hour hackathon J

Langsmith engine is the future

AgentsDGX agent

LangSmith is a development platform designed to build, test, and monitor LLM applications, offering tools for debugging and evaluation to improve AI application quality and reliability. The statement

Learning Agent-Compatible Context Management for Long-Horizon Tasks

AgentsDGX agent

arXiv:2605.30785v1 Announce Type: new Abstract: LLM agents increasingly face long-horizon tasks such as web search and deep research in real-world applications, where accumulated context can cause lon

Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration

AgentsDGX agent

arXiv:2605.31365v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to promising progress in web agents. However, existing web agents often rely on han

Learning to Perceive the World Through Control: Empowerment-Based Representation Learning

AgentsDGX agent

arXiv:2605.30656v1 Announce Type: new Abstract: In many practical reinforcement learning environments, observations are far higher-dimensional than the variables that matter for control. In this work,

Let engine help you build better agents

AgentsDGX agent

This post likely discusses how the LangChain framework (which Harrison Chase co-founded) can assist developers in constructing more effective AI agents by providing tools, abstractions, and patterns f

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

AgentsDGX agent

arXiv:2603.22744v2 Announce Type: replace Abstract: Large language models excel on objectively verifiable tasks such as math and programming, where evaluation reduces to unit tests or a single correct

LLM Anonymization Against Agentic Re-Identificatio

AgentsDGX agent

arXiv:2605.30848v1 Announce Type: cross Abstract: Agentic LLMs with web search change the threat model for text anonymization: weak contextual cues can become cross-referenceable evidence for re-ident

longmemeval experiment arch: 1) deterministic ingestion/extraction (85.6% accuracy, 86.2% retrieval) 2) semantic ingestion/extraction (84.8%…

AgentsDGX agent

longmemeval experiment arch: 1) deterministic ingestion/extraction (85.6% accuracy, 86.2% retrieval) 2) semantic ingestion/extraction (84.8% accuracy, 94.9% retention) 3) semantic ingestion/determinis

LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards

AgentsDGX agent

arXiv:2605.31584v1 Announce Type: cross Abstract: Long-context reasoning remains a central challenge for large language models, which often fail to locate and integrate key information in extensive di

Managed Deep Agents keeps the project shape you already know: ↳ AGENTS.md, skills/, subagents/, + tools.json Context Hub gives your agent a …

AgentsDGX agent

Managed Deep Agents keeps the project shape you already know: ↳ AGENTS.md, skills/, subagents/, + tools.json Context Hub gives your agent a managed place to retain and update this context across sessi

MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks

AgentsDGX agent

arXiv:2603.02630v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved great success in many real-world applications, especially the one serving as the cognitive backbone

MatchFixAgent: Language-Agnostic Autonomous Repository-Level Code Translation Validation and Repair

AgentsDGX agent

arXiv:2509.16187v3 Announce Type: replace-cross Abstract: Code translation transforms source code from one programming language (PL) to another. Validating the functional equivalence of translation an

May 2026 newsletter

AgentsDGX agent

I just sent out the May edition of my sponsors-only monthly newsletter. If you are a sponsor (or if you start a sponsorship now) you can access it here. This month: Al got expensive, and Anthropic had

MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive Regulation

AgentsDGX agent

arXiv:2602.07905v2 Announce Type: replace Abstract: Large Language Models (LLMs) have shown strong potential in complex medical reasoning yet face diminishing gains under inference scaling laws. While

MiniMax M3 imminent. Will be doing deep testing with it on my own coding agent and harness. Review coming soon.

AgentsDGX agent

MiniMax M3, an upcoming AI model, is expected to be released soon and will undergo comprehensive testing within a custom coding agent framework. A detailed technical review of the model's performance

MiniMax-M3 will by arrive on HuggingFace openweight at next week!

AgentsDGX agent

MiniMax-M3 will by arrive on HuggingFace openweight at next week! Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Ben

More info about Search as Code in the Perplexity Agent API docs: https://docs.perplexity.ai/docs/agent-api/tools/sandbox

AgentsDGX agent

The Perplexity Agent API documentation includes a 'Search as Code' feature accessible through the sandbox tools section, enabling developers to integrate search functionality programmatically within a

More on the Hermes Skills Hub: https://hermes-agent.nousresearch.com/docs/guides/work-with-skills#the-skills-hub

AgentsDGX agent

The Hermes Skills Hub is a feature that allows users to discover, manage, and integrate skills within the Hermes agent framework, enabling extended functionality and customization of agent capabilitie

.@MukilLoganathan’s Interrupt keynote on Sandboxes. https://youtu.be/IIchUA5T3gs In 20 minutes, you’ll learn how to run agent code safely. I…

AgentsDGX agent

.@MukilLoganathan’s Interrupt keynote on Sandboxes. https://youtu.be/IIchUA5T3gs In 20 minutes, you’ll learn how to run agent code safely. Isolated from your runtime, with network controls, persistent

Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely

AgentsDGX agent

arXiv:2605.31387v1 Announce Type: new Abstract: Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected

NEMO: Execution-Aware Optimization Modeling via Autonomous Coding Agents

AgentsDGX agent

arXiv:2601.21372v2 Announce Type: replace Abstract: We present NEMO, a system that translates Natural-language descriptions of decision problems into formal Executable Mathematical Optimization implem

NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving

AgentsDGX agent

arXiv:2605.31116v1 Announce Type: new Abstract: Recent perception-free end-to-end (E2E) autonomous driving methods bypass explicit perception outputs by compressing dense image patch tokens into compa

Open models!

AgentsDGX agent

Open models! Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency

PictSure: Pretraining Embeddings Matters for In-Context Learning Image Classifiers

AgentsDGX agent

arXiv:2506.14842v2 Announce Type: replace-cross Abstract: Building image classification models remains cumbersome in data-scarce domains, where collecting large labeled datasets is impractical. In-con

Provably Convergent Actor-Critic for MARL through Risk-aversion

AgentsDGX agent

arXiv:2602.12386v2 Announce Type: replace-cross Abstract: Learning stationary policies in infinite-horizon general-sum Markov games (MGs) remains a fundamental open problem in Multi-Agent Reinforcemen

Pull Requests as a Training Signal for Repo-Level Code Editing

AgentsDGX agent

arXiv:2602.07457v2 Announce Type: replace-cross Abstract: Repository-level code editing requires models to understand complex dependencies and execute precise multi-file modifications across a large c

.@Rippling AI runs on Deep Agents and LangSmith. Here’s how they shipped to millions of users in 6 months. https://www.langchain.com/blog/ho…

AgentsDGX agent

.@Rippling AI runs on Deep Agents and LangSmith. Here’s how they shipped to millions of users in 6 months. https://www.langchain.com/blog/how-rippling-went-ai-native-across-every-product-in-6-months-w

RT @jayfarei: I think this idea a lot!

AgentsDGX agent

I cannot provide an accurate summary of this post as the tweet content itself is not provided, only a reference indicating Jay Farei's engagement with an idea shared by Harrison Chase. Without access

S^3LDBO: A Snapshot Single-Loop Algorithm for Decentralized Bilevel Optimization

AgentsDGX agent

arXiv:2605.31311v1 Announce Type: cross Abstract: Networked AI systems increasingly rely on multiple agents that collaboratively learn and adapt models over communication networks. In such systems, bi

SAGE: A Novelty Gate for Efficient Memory Evolution in Agentic LLMs

AgentsDGX agent

arXiv:2605.30711v1 Announce Type: cross Abstract: Agentic LLMs must continuously decide whether newly extracted facts should be added, merged with existing memories, or ignored, yet prior work has foc

Sent out the May edition of my sponsors-only newsletter, for people who don't have time to read my blog every day and want to pay me money t…

AgentsDGX agent

Simon Willison announced the release of a May edition of his sponsors-only newsletter, which is designed for supporters who prefer a curated summary format instead of following his daily blog posts. T

Skill Reuse as Compression in Agentic RL

AgentsDGX agent

arXiv:2605.31509v1 Announce Type: cross Abstract: Large language model agents trained with reinforcement learning (RL) often learn brittle, task-specific shortcuts. We hypothesize that agents generali

slowly we're all realizing that tools should be called from code, not from within the llm api

AgentsDGX agent

slowly we're all realizing that tools should be called from code, not from within the llm api Introducing Search as Code, our new search architecture for AI agents. It writes Python that calls our sea

Social Reasoning in Machines: Investigating Collective Truth-Seeking Dynamics in Large Language Model Debate

AgentsDGX agent

arXiv:2605.30391v1 Announce Type: cross Abstract: Human reasoning has long been theorised to operate socially, not through isolated individual cognition, but through collective adversarial discourse,

Sophrosyne: Agentic Exploration of Relational Data Systems Needs Moderation

AgentsDGX agent

arXiv:2605.30862v1 Announce Type: cross Abstract: Text2SQL agents powered by LLMs translate natural language intent into SQL by exploring the data system through tool calls before formulating the quer

Sources: Tencent, which has fallen behind domestic rivals in AI models, plans to test an AI agent for WeChat with a small group of users before a phased rollout (Zijing Wu/Financial Times)

AgentsDGX agent

Zijing Wu / Financial Times: Sources: Tencent, which has fallen behind domestic rivals in AI models, plans to test an AI agent for WeChat with a small group of users before a phased rollout — Maker of

SpecDB: LLM-Generated Customized Databases via Feature-Oriented Decomposition

AgentsDGX agent

arXiv:2605.31097v1 Announce Type: cross Abstract: Mainstream relational databases ship a uniform feature set across deployments, although individual workloads exercise only a fraction of the available

Stop manually triaging agent failures. Let LangSmith Engine fix it.

AgentsDGX agent

LangSmith Engine is a tool designed to automatically diagnose and resolve agent failures, eliminating the need for manual troubleshooting and triage. The feature appears to leverage automated analysis

Subspace-Decomposed JEPAs: Disentangling Progression and Content in Latent World Models

AgentsDGX agent

arXiv:2605.31111v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) learn compact latent world models by predicting future embeddings, but no single coordinate of the late

Surprised by Attention: Predictable Query Dynamics for Time Series Anomaly Detection

AgentsDGX agent

arXiv:2603.12916v3 Announce Type: replace-cross Abstract: Multivariate time series anomalies often manifest as shifts in cross-channel dependencies rather than simple amplitude excursions. In autonomo

Survival Reinforcement Learning: Toward Scalable Self-Supervised RL

AgentsDGX agent

arXiv:2605.31273v1 Announce Type: new Abstract: While self-supervised Contrastive Reinforcement Learning (CRL) has shown remarkable depth-scaling capabilities, successfully using networks over 64 laye

← Previous
1…5354555657…121
Next →