AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,973 results
9 Jul 2026

Join us for a LangChain + Clay meetup with @palashshah, @jeffbarg, Vyshu Khota, and Soroush Khadem. https://luma.com/jqif2hti Palash will br…

AgentsDGX agent

Join us for a LangChain + Clay meetup with @palashshah, @jeffbarg, Vyshu Khota, and Soroush Khadem. https://luma.com/jqif2hti Palash will break down how he built a self-improving agent at LangChain, L

RLVP: Penalize the Path, Reward the Outcome

AgentsDGX agent

arXiv:2607.07435v1 Announce Type: cross Abstract: Agents acting on our behalf in the real world (e.g. placing phone calls) must learn online from costly, often irreversible interactions rather than ch

7 Jul 2026

AutoCedar: An Agentic Framework for Verifier-Guided Access Control Policy Synthesis

Model Releases
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2607.03656v1 Announce Type: cross Abstract: Large Language Models are increasingly used to turn natural-language requirements into code. In access control, that shortcut is dangerous: a generate

GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads

Model ReleasesDGX agent

arXiv:2607.02518v1 Announce Type: cross Abstract: OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the c

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

Model ReleasesDGX agent

arXiv:2510.11098v5 Announce Type: replace-cross Abstract: Recent advances in large audio language models (LALMs) have greatly enhanced multimodal conversational systems. However, existing benchmarks r

6 Jul 2026

Come join us on Thursday! It'll be a great session you'll want to add to your MEMORY.md 😄

AgentsDGX agent

Come join us on Thursday! It'll be a great session you'll want to add to your MEMORY.md 😄 does your agent need a wiki? should humans and agent use the same wiki? should you have one wiki or many wikis

devin has a hot take that wikis are NOT the right abstraction should be a fun webinar!!

AgentsDGX agent

devin has a hot take that wikis are NOT the right abstraction should be a fun webinar!! does your agent need a wiki? should humans and agent use the same wiki? should you have one wiki or many wikis?

5 Jul 2026

sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

Model ReleasesDGX agent

I wrote about the sqlite-utils 4.0rc1 release a couple of weeks ago. Since we only have Claude Fable on our Max subscriptions for a few more days, I decided to see if it could help me get to a 4.0 sta

3 Jul 2026

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows

SafetyDGX agent

arXiv:2607.01465v1 Announce Type: new Abstract: Large language models are trained to predict the next token, not to act inside a specific API. In niche enterprise SaaS workflows -- where success means

COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows

Model ReleasesDGX agent

arXiv:2607.01709v1 Announce Type: new Abstract: Agents are increasingly used to construct workflows and assist humans in completing recurring tasks more efficiently. As these workflows become repeated

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents

SafetyDGX agent

arXiv:2607.01557v1 Announce Type: cross Abstract: Large Language Models (LLMs) often struggle with persuasion in high-stakes scenarios. People's individual personalities and concerns require tailored

Episodic-to-Semantic Consolidation Without Identity Drift

SafetyDGX agent

arXiv:2607.01988v1 Announce Type: new Abstract: Long-running adaptive intelligent agents face a structural tension between knowledge consolidation and information integrity. Memory consolidation is co

SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplaces

AgentsDGX agent

arXiv:2607.02345v1 Announce Type: cross Abstract: Large Language Model (LLM)-based agents increasingly automate software engineering tasks through reusable skills, natural-language instruction documen

World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments

TutorialsDGX agent

arXiv:2607.01470v1 Announce Type: new Abstract: Clinical protocol-execution tasks -- checking a lab value, applying a threshold, placing a correctly structured FHIR order -- are natural candidates for

2 Jul 2026

Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions

Model ReleasesDGX agent

arXiv:2607.00304v1 Announce Type: cross Abstract: The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strat

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

SafetyDGX agent

arXiv:2510.24636v3 Announce Type: replace Abstract: Reward models (RMs) have become essential for aligning large language models (LLMs), serving as scalable proxies for human evaluation in both traini

1 Jul 2026

How Inscribe uses Amazon Bedrock to stop document fraud in seconds

AgentsDGX agent

In this post, you will learn how Inscribe developed an agentic AI system using Amazon Bedrock that reasons across documents the way an expert fraud analyst would. With this new agentic AI system, Insc

MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents

Model ReleasesDGX agent

arXiv:2606.31167v1 Announce Type: cross Abstract: VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current s

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

SafetyDGX agent

arXiv:2606.30887v1 Announce Type: cross Abstract: Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control

30 Jun 2026

A Diagnostic Framework and Multi-Evaluator Audit of Evaluator-Driven Preference Dynamics in Self-Adapting LLM Agents

Model ReleasesDGX agent

arXiv:2606.29719v1 Announce Type: cross Abstract: Measurements of proprietary LLM evaluators can become invalid within weeks -- we document one case and provide the diagnostic framework to detect it.

Agentic Safety is an Epistemic Property, Not a Behavioral One

SafetyDGX agent

arXiv:2606.28347v1 Announce Type: cross Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. These methods

Improved Multi-Dimensional Forecasting for Swap Regret

AgentsDGX agent

arXiv:2606.29533v1 Announce Type: cross Abstract: We study the problem of forecasting for an arbitrary number of downstream agents with unknown objectives, each of whom best responds to the forecaster

Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts

AgentsDGX agent

arXiv:2606.29279v1 Announce Type: cross Abstract: LLM agents carry conclusions across steps and sessions in compressed memory, and memory products (e.g., mem0, LangMem) rewrite conversation into store

NVIDIA BioNeMo Agent Toolkit Brings Accelerated AI to Life Sciences Researchers in Claude Science

Model ReleasesDGX agent

Life sciences has entered an era of computational scale, and for more than a decade, NVIDIA has built the full GPU-accelerated computing stack — spanning hardware, frameworks, libraries, models, micro

On the Necessity of a Liquid Substrate for Mesh Intelligence

AgentsDGX agent

arXiv:2606.28413v1 Announce Type: cross Abstract: A mesh of sovereign agents has no center: no shared clock, no shared model, and no coordinator to gather data or retrain. Its competence rests on each

The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance

SafetyDGX agent

arXiv:2606.28710v1 Announce Type: new Abstract: We ask under what conditions an agent with a harm-minimizing policy can displace an approval-seeking (RLHF) agent in a competitive market, and when that

29 Jun 2026

AI agents are not your “coworkers”

TutorialsDGX agent

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Imagine coming in to work to learn that a new underling will r

26 Jun 2026

Chai: Agentic Discovery of Cryptographic Misuse Vulnerabilities

SafetyDGX agent

arXiv:2606.26933v1 Announce Type: cross Abstract: AI-assisted vulnerability discovery has proven effective for bug classes like memory safety, where instrumentation confirms memory violations and effi

25 Jun 2026

Open + closed models = better together. Our previous research with @harvey showed the benefits of combining a frontier closed model as an ad…

AgentsDGX agent

Open + closed models = better together. Our previous research with @harvey showed the benefits of combining a frontier closed model as an advisor agent with fine-tuned, open-source worker agents. Thre

SoK: AI Secure Code Generation: Progress, Pitfalls, and Paths Forward

AgentsDGX agent

arXiv:2606.25195v1 Announce Type: cross Abstract: The increasing use of AI systems for code generation raises a central security question: what can today's models and coding agents actually do to prod

24 Jun 2026

Introducing Claude for Music. You can now create songs from Claude Code, Hermes, Codex, or any agent you’re using. SOTA music model @MiniMax…

Model ReleasesDGX agent

I cannot provide an accurate summary for this entry. The URL and source attribution appear inconsistent (title credits Yohei Nakajima but URL references a different user), and the post references prod

When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents

Model ReleasesDGX agent

arXiv:2606.23937v1 Announce Type: cross Abstract: Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model. We test t

23 Jun 2026

NVIDIA Brings Trusted, 24/7 AI Agents to Telecom Operations

HardwareDGX agent

Telecom operators have seen remarkable returns from using generative AI to automate network management, customer care and back-office operations. Most of that impact has been task‑based: automation th

Position: Correct Answer, Wrong Mechanism -- When AI Scientists Defend General Claims Their Own Data Contradicts

AgentsDGX agent

arXiv:2606.23175v1 Announce Type: new Abstract: AI scientist systems are described as tools, coauthors, or founders, but we evaluate them as if only the final answer matters. This position paper argue

11 Jun 2026

Fourier Features Let Agents Learn High Precision Policies with Imitation Learning

SafetyDGX agent

arXiv:2606.12334v1 Announce Type: new Abstract: High-precision robotic manipulation requires fine-grained spatial reasoning that is often difficult to achieve with RGB-only policies due to depth ambig

10 Jun 2026

BadRobot: Jailbreaking Embodied LLM Agents in the Physical World

Model ReleasesDGX agent

arXiv:2407.20242v5 Announce Type: replace-cross Abstract: Embodied AI represents systems where AI is integrated into physical entities. Large Language Model (LLM), which exhibits powerful language und

Constructing coherent spatial memory in LLM agents through graph rectification

Model ReleasesDGX agent

arXiv:2510.04195v2 Announce Type: replace Abstract: Given a map description through global traversal navigation instructions, an LLM can often infer the implicit spatial layout and answer user queries

it’s actually so cool to work at LangChain the…database company (??) yup, the cracked team that built SmithDB is doing a cool blog series on…

AgentsDGX agent

it’s actually so cool to work at LangChain the…database company (??) yup, the cracked team that built SmithDB is doing a cool blog series on “How to build the internals of a database” —> for agent sca

Regimes: An Auditable, Held-Out-Gated Improvement Loop Demonstrated on LongMemEval with ActiveGraph

AgentsDGX agent

arXiv:2606.10241v1 Announce Type: new Abstract: Autonomous improvement loops are hard to trust because the improvement process is usually external scaffolding bolted onto the agent: failures go unlogg

Trace2Policy: From Expert Behavior Traces to Self-Evolving Decision Agents

ApplicationsDGX agent

arXiv:2606.10457v1 Announce Type: new Abstract: Decision rules that enterprise experts apply tacitly -- in auditing, compliance, and contract review -- can be systematically recovered and improved thr

8 Jun 2026

On the Hardness of Optimal Motion on Trees

AgentsDGX agent

arXiv:2606.06686v1 Announce Type: new Abstract: This paper presents a simple framework that settles the complexity of Multi-Agent Path Finding (MAPF) on trees across standard objectives--distance, mak

7 Jun 2026

New MIT study. Code volume surges by 300%, but output increases by only 30%: The AI dividend meets an awkward reality Autonomous AI coding a…

AgentsDGX agent

New MIT study. Code volume surges by 300%, but output increases by only 30%: The AI dividend meets an awkward reality Autonomous AI coding agents raised commits by 180%, but releases rose only 30%. Th

6 Jun 2026

PerceptUI: LLM Agents as Human-Aligned Synthetic Users for UI/UX Evaluation

ResearchDGX agent

arXiv:2606.05697v1 Announce Type: new Abstract: User interface (UI) and user experience (UX) evaluation is central to product development, yet reliable feedback still relies on recruiting human partic

Running Python code in a sandbox with MicroPython and WASM

Model ReleasesDGX agent

I've been experimenting with different approaches to running code in a sandbox for several years now, but my latest attempt feels like it might finally have all of the characteristics I've been lookin

this looks cool

AgentsDGX agent

this looks cool An experimental programming language from Vercel Labs that is truly made for AI agents! Not just a new syntax. Zero is a graph-first language where agents can Read & edit program struc

5 Jun 2026

MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery

AgentsDGX agent

arXiv:2606.06473v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly applied to long-horizon tasks such as scientific discovery and machine learning engineering (MLE),

4 Jun 2026

AgenticDiffusion: Agentic Diffusion-based Path Planning for Vision-Based UAV Navigation

Local AiDGX agent

arXiv:2606.04111v1 Announce Type: cross Abstract: Indoor UAV navigation requires efficient exploration, scene understanding, and reliable trajectory execution under limited field-of-view observations.

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents

SafetyDGX agent

arXiv:2606.04703v1 Announce Type: new Abstract: Experience internalization converts contextual experience from past interactions into reusable parametric capability, offering a promising path toward c

3 Jun 2026

Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition

Model ReleasesDGX agent

arXiv:2606.03657v1 Announce Type: new Abstract: Large language models for code generation often need to use APIs that are absent from their pretraining data. This requires more than recalling a functi

Microsoft’s open trust stack runs on OpenInference

AgentsDGX agent

Microsoft's open trust stack for AI agents puts ASSERT and Agent Control Specification on top of OpenInference, connecting evaluation, runtime controls, and observability through a shared trace contra

NVIDIA Research Unlocks Advanced Grasping, Smarter Autonomous Driving and Agent Training at Scale

HardwareDGX agent

What makes a robot gripper useful isn’t that it can pick up one object — it’s that it can pick up the next one, and the one after that, with a tool it’s never held before. What makes an autonomous veh

2 Jun 2026

Available on all platforms at the link below: https://hermes-agent.nousresearch.com/desktop

AgentsDGX agent

Nous Research announced the availability of Hermes Agent across all platforms through a desktop application accessible via the provided link. This release makes their Hermes Agent tool universally ava

Knowing Isn't Understanding: Re-grounding Generative Proactivity with Epistemic and Behavioral Insight

AgentsDGX agent

arXiv:2602.15259v2 Announce Type: replace-cross Abstract: Generative AI agents equate understanding with resolving explicit queries, an assumption that confines interaction to what users can articulat

VideoBrain: Learning Adaptive Frame Sampling for Long Video Understanding

AgentsDGX agent

arXiv:2602.04094v2 Announce Type: replace Abstract: Long-form video understanding remains challenging for Vision-Language Models (VLMs) due to the inherent tension between computational constraints an

1 Jun 2026

Choosing the Lens: Strategic Perspective Activation in Context-Dependent Argumentation

AgentsDGX agent

arXiv:2605.31581v1 Announce Type: new Abstract: The same arguments often need to be evaluated under different external regimes. An agent with influence over the regime has a strategic lever that stand

COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation

AgentsDGX agent

arXiv:2605.31264v1 Announce Type: new Abstract: LLM agents are increasingly expected not only to complete isolated tasks, but also to carry bounded representations of human expertise, judgment, and in

S^3LDBO: A Snapshot Single-Loop Algorithm for Decentralized Bilevel Optimization

AgentsDGX agent

arXiv:2605.31311v1 Announce Type: cross Abstract: Networked AI systems increasingly rely on multiple agents that collaboratively learn and adapt models over communication networks. In such systems, bi

31 May 2026

Why ‘human in the loop’ falls short – and what to do about it

AgentsDGX agent

Agentic artificial intelligence governance depends upon humans to keep agentic AI from going off the rails. However, putting humans in the loop is woefully insufficient. Here are the problems – and pe

29 May 2026

CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization

Model ReleasesDGX agent

arXiv:2510.14150v5 Announce Type: replace Abstract: We introduce CodeEvolve, an open-source framework that couples large language models with island-based evolutionary search for end-to-end algorithmi

Evaluation of Conversational Agents: Understanding Culture, Context and Environment in Emotion Detection

ApplicationsDGX agent

arXiv:2605.30099v1 Announce Type: new Abstract: Valuable decisions and highly prioritized analysis now depend on applications such as facial biometrics, social media photo tagging, and human robots in

← Previous
1…151152153154155…300
Next →