AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,964 results
9 Aug 2026

the term “RLM” (recursive language model) got a lot of buzz this week, but this idea is not new! @a1zhang wrote the og RLM paper 10 months a…

AgentsDGX agent

the term “RLM” (recursive language model) got a lot of buzz this week, but this idea is not new! @a1zhang wrote the og RLM paper 10 months ago! thats like 5 agent-years! would highly recommend followi

8 Aug 2026

HUD mode Hermes stops being a window you switch to and becomes a layer over the app you're both working in. Or keep it around as a little bu…

AgentsDGX agent

HUD mode Hermes stops being a window you switch to and becomes a layer over the app you're both working in. Or keep it around as a little buddy agent. Ask it random things, drag it anywhere, it's your

7 Aug 2026

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Predicting Task Difficulty Without Rollouts
AgentsDGX agent

arXiv:2608.05797v1 Announce Type: cross Abstract: Task difficulty dictates an agent's likelihood of success, and estimating it without rollouts means forecasting this directly from a task description

Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforcement Learning

AgentsDGX agent

arXiv:2608.05245v1 Announce Type: new Abstract: Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path toward self-e

The backgrounds were generated with MiniMax M3. Everything else, including the full game design, was done with DeepSeek Flash. Most importan…

Model ReleasesDGX agent

The backgrounds were generated with MiniMax M3. Everything else, including the full game design, was done with DeepSeek Flash. Most importantly, all of this was done inside Hermes Agent. You won't bel

this talk by openai researchers going through hugging face incident is totally insane, so much to unpack openai only realized it was their a…

Model ReleasesDGX agent

this talk by openai researchers going through hugging face incident is totally insane, so much to unpack openai only realized it was their agent who hacked hugging face infra while asking hf to revoke

6 Aug 2026

CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications

AgentsDGX agent

arXiv:2608.04942v1 Announce Type: cross Abstract: CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological ap

4 Aug 2026

Is LM Studio abandoning their core product?

Model ReleasesDGX agent

Some of you may be aware that a few weeks ago, LM Studio announced a new agent, Bionic. This is pretty much an agentic harness for both local models and paid cloud models. But most aren't aware that L

New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

Model ReleasesDGX agent

I released LLM 0.32 this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider t

OpenAI says one of its models exploited a website after third-party AI security lab Irregular mistakenly gave it access to the internet during evaluations (Wired)

AgentsDGX agent

Wired: OpenAI says one of its models exploited a website after third-party AI security lab Irregular mistakenly gave it access to the internet during evaluations — Rogue AI agents from OpenAI and Anth

1 Aug 2026

Grok Build can do almost anything you can think of http://X.ai/cli

AgentsDGX agent

Grok Build can do almost anything you can think of http://X.ai/cli Most people seriously underestimate what Grok Build can do They assume an AI coding agent is only useful for building apps or writing

31 Jul 2026

Evidence-Ledger Adjudication for Claim-Evidence Traceability

Model ReleasesDGX agent

arXiv:2607.26512v1 Announce Type: new Abstract: AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a

GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios

AgentsDGX agent

arXiv:2607.28073v1 Announce Type: cross Abstract: In demanding professional environments and meeting review scenarios, lengthy text often imposes a high cognitive load. To facilitate efficient informa

KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation

Model ReleasesDGX agent

arXiv:2607.27231v1 Announce Type: cross Abstract: Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specia

PUDA: An AI-Native Hardware Harness for Self-Driving Laboratories

AgentsDGX agent

arXiv:2607.26464v1 Announce Type: cross Abstract: Physical Unified Device Architecture (PUDA) is an AI-native hardware harness for self-driving laboratories (SDLs). Rather than building a human-center

Representation and Invariance in Reinforcement Learning

AgentsDGX agent

arXiv:2112.07752v4 Announce Type: replace-cross Abstract: Researchers have formalized reinforcement learning (RL) in different ways. If an agent in one RL framework is to run within another RL framewo

Training Skills Like Parameters via Self-Supervised Semantic Diffusion

AgentsDGX agent

arXiv:2607.27557v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate remarkable general instruction-following capabilities, they often fall short of human experts in highly s

30 Jul 2026

Introducing the Parse Gateway There's been an explosion of interest in model routing - you don't always need the best model for every task. …

AgentsDGX agent

Introducing the Parse Gateway There's been an explosion of interest in model routing - you don't always need the best model for every task. That is especially true for document parsing 📄🔀: - Some page

Mental World Modeling

AgentsDGX agent

arXiv:2607.27201v1 Announce Type: new Abstract: World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and h

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

SafetyDGX agent

arXiv:2607.19321v2 Announce Type: replace-cross Abstract: As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be

29 Jul 2026

SAFAARI: Schema-Aware Framework for Accelerated Advertiser Response Intelligence

AgentsDGX agent

arXiv:2607.25042v1 Announce Type: new Abstract: The evolution of customer support systems is rapidly advancing with agentic chatbots, yet these systems face significant limitations when accessing ente

Specula: Scaling formal specifications for autonomous model checking of system code

AgentsDGX agent

arXiv:2607.25333v1 Announce Type: cross Abstract: Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the specifications f

28 Jul 2026

CRAFT: Learn the Schema, Execute the Plan

SafetyDGX agent

arXiv:2607.22642v1 Announce Type: new Abstract: Enterprise coding agents translate natural-language analytical requests into executable code over proprietary APIs, schemas, and metric definitions. Yet

HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System

AgentsDGX agent

arXiv:2603.14807v3 Announce Type: replace Abstract: LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) tasks. However, most zero-shot methods prima

KG2Code: Bridging Knowledge Graphs and Large Language Models via Executable Code for Question Answering

AgentsDGX agent

arXiv:2607.22652v1 Announce Type: new Abstract: Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream know

25 Jul 2026

Reuters confirms my hypothesized time line. OpenAI did not realize its AI has breached the sandbox for a week. Astounding. https://www.reute…

AgentsDGX agent

Reuters confirms my hypothesized time line. OpenAI did not realize its AI has breached the sandbox for a week. Astounding. https://www.reuters.com/business/its-ai-agent-spent-days-hacking-company-sour

23 Jul 2026

Dynamic workflows are a generalization of harnesses, automations, loops, routing, and graphs. It's the most powerful feature I have built in…

Model ReleasesDGX agent

Dynamic workflows are a generalization of harnesses, automations, loops, routing, and graphs. It's the most powerful feature I have built into my agent orchestrator. Supports all kinds of patterns tha

inclusionAI/LLaDA2.2-flash · Hugging Face

Model ReleasesDGX agent

LLaDA2.2-flash is an agent-oriented diffusion language model in the LLaDA2 series. By introducing Levenshtein Editing (with DELETE and INSERT control tokens) to diffusion language modeling, it represe

PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

AgentsDGX agent

arXiv:2607.20268v1 Announce Type: new Abstract: While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterativ

16 Jul 2026

UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

Model ReleasesDGX agent

arXiv:2607.13621v1 Announce Type: new Abstract: Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visib

15 Jul 2026

Differentiable Clone-Structured Causal Graphs for End-to-End Cognitive Map Learning from Image Sequences

AgentsDGX agent

arXiv:2607.12382v1 Announce Type: new Abstract: How can an agent build a structured map of its world from nothing but an ongoing sequence of raw sensory input and its own movements, especially when na

In-Context Reinforcement Learning under Non-Stationarity: A Survey

Model ReleasesDGX agent

arXiv:2607.11906v1 Announce Type: new Abstract: The development of decision-pretrained transformers, algorithm distillation, long-context meta-RL, and retrieval-augmented agents has renewed interest i

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

Model ReleasesDGX agent

arXiv:2607.12625v1 Announce Type: new Abstract: OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a we

Token Reduction Is Not Cost Reduction

Model ReleasesDGX agent

arXiv:2607.12161v1 Announce Type: new Abstract: Context-reduction layers for API-based coding agents, including command-output compressors, retrieval rankers, and payload-optimizing proxies, are usual

14 Jul 2026

Context harnessing is super challenging to do right: what info needs to be collected, how to accumulate it into a knowledge graph, how to ke…

AgentsDGX agent

Context harnessing is super challenging to do right: what info needs to be collected, how to accumulate it into a knowledge graph, how to keep it up to date, etc... Super excited for @QodoAI, harnessi

ScienceSoft’s HIPAA-compliant AI voice scheduler built on AWS

SafetyDGX agent

In this post, you will learn how ScienceSoft, an Amazon Web Services (AWS) Services Partner, integrated Amazon Nova 2 Sonic with Amazon Bedrock Guardrails to build a Health Insurance Portability and A

10 Jul 2026

3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse

AgentsDGX agent

arXiv:2607.07980v1 Announce Type: cross Abstract: Coding agents now author entire pull requests, and practitioners sharply disagree about what this does to code review: whether it becomes the bottlene

Computation, Condensation, and the Incompleteness Between Them: A Coupled Foundation of Intelligence

AgentsDGX agent

arXiv:2303.04203v4 Announce Type: replace-cross Abstract: The theory of computation was built to answer Turing's question: what is effectively calculable by an unbounded, immortal, disembodied agent f

9 Jul 2026

Distributed Dynamic Associative Memory via Online Convex Optimization

AgentsDGX agent

arXiv:2511.23347v2 Announce Type: replace Abstract: An associative memory (AM) enables cue-response recall, and it has recently been recognized as a key mechanism underlying modern neural architecture

Safely run AI-generated code in Cloud Run sandboxes

Model ReleasesDGX agent

Here’s a question we hear often at Google Cloud: How do you safely run AI-generated code or untrusted binaries without putting your host application, data, and cloud credentials at risk? In other word

3 Jul 2026

Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics

AgentsDGX agent

arXiv:2607.02329v1 Announce Type: new Abstract: Autonomous-research agents have demonstrated end-to-end LLM automation in machine-learning sandboxes where execution provides calibration. Frontier phys

1 Jul 2026

A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization

AgentsDGX agent

arXiv:2606.30775v1 Announce Type: cross Abstract: Enterprise AI agents route user queries to specialized skills by matching queries against natural language skill descriptions. When two skills share o

Better Understanding, Understanding Better

AgentsDGX agent

arXiv:2606.31892v1 Announce Type: cross Abstract: 'Any fool can know; the point is to understand.' A well-known remark often attributed to Einstein captures a widely shared intuition: understanding is

Emergent Culture in Minimal LLM Systems

AgentsDGX agent

arXiv:2606.30668v1 Announce Type: cross Abstract: What happens when LLM agents operate with no context outside a turn, minimal prompting, and simple tools? Inspired by swarm engineering, we give colle

Imagine if, as an engineer, your working memory of a codebase was wiped every time you started a new ticket Sounds ridiculous, but this is f…

AgentsDGX agent

Imagine if, as an engineer, your working memory of a codebase was wiped every time you started a new ticket Sounds ridiculous, but this is functionally what happens to agents whenever we start a new t

Nazrin: An Atomic Neural Proof Automation Tactic in Lean 4

AgentsDGX agent

arXiv:2602.18767v3 Announce Type: replace-cross Abstract: In Machine-Assisted Theorem Proving, a theorem proving agent searches for a sequence of expressions and tactics that can prove a statement in

30 Jun 2026

a very cool Harbor x LangSmith flow I love to help you “look at the data”: 1. you do evals or rollouts for RL 2. all reward metrics and trac…

AgentsDGX agent

a very cool Harbor x LangSmith flow I love to help you “look at the data”: 1. you do evals or rollouts for RL 2. all reward metrics and traces and rollouts get automatically propulates into Experiment

Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models

AgentsDGX agent

arXiv:2606.28524v1 Announce Type: new Abstract: Recent work suggests that Large Language Models (LLMs) are sensitive to the belief states of agents described by text, as measured by the false belief t

UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image

AgentsDGX agent

arXiv:2606.30608v1 Announce Type: new Abstract: Articulated 3D objects are essential for interactive environments in embodied AI, robotics, and virtual reality, but reconstructing their structure and

29 Jun 2026

We're all over AI Engineer World's Fair on June 29 to July 2. 🦙 📍Visit us at booth L-G47. LlamaParse demos + Fear of Docs swag 🎤 @jerryjl…

AgentsDGX agent

We're all over AI Engineer World's Fair on June 29 to July 2. 🦙 📍Visit us at booth L-G47. LlamaParse demos + Fear of Docs swag 🎤 @jerryjliu0, our Founder & CEO, on agentic document parsing and shippin

27 Jun 2026

I’m speaking at AIE SF World’s Fair on July 1st in the Sandbox & Platform Engineering track. My talk is called 'From fork() to Fleet: Design…

AgentsDGX agent

I’m speaking at AIE SF World’s Fair on July 1st in the Sandbox & Platform Engineering track. My talk is called 'From fork() to Fleet: Designing an Agent Sandbox Cloud'. I’ll cover the OS and Infrastru

25 Jun 2026

Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Based Web Penetration Testing

AgentsDGX agent

arXiv:2606.25332v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promise for automated penetration testing, yet existing end-to-end black-box evaluations are highly susceptibl

Loops, one of the most symbolic tools ever invented, are rescuing generative AI. Neurosymbolic AI is completely dominating.

AgentsDGX agent

Loops, one of the most symbolic tools ever invented, are rescuing generative AI. Neurosymbolic AI is completely dominating. Anthropic engineers just showed how to build agents that can run for days wi

Tinker Tales: A Tangible Dialogue System for Child-AI Co-Creative Storytelling

AgentsDGX agent

arXiv:2602.04109v2 Announce Type: replace-cross Abstract: Conversational AI agents are increasingly explored as creative partners, yet how conversation design shapes child-AI dialogue in co-creative s

24 Jun 2026

From Task-Guided Conversational Graphs to Goal-Oriented Dialogue Runtimes

AgentsDGX agent

arXiv:2606.23797v1 Announce Type: cross Abstract: Graph and multi-agent orchestration frameworks make production large language model (LLM) workflows practical, but they do not by themselves solve con

Sleep-time compute is the next scaling axis for intelligence

AgentsDGX agent

Sleep-time compute is the next scaling axis for intelligence 🧠LangSmith Engine as Sleep Time Compute Memory for agents is often described as “sleep time compute” or “dreaming” This involves running a

23 Jun 2026

Reinforcement Learning to Disentangle Multiqubit Quantum States from Partial Observations

AgentsDGX agent

arXiv:2406.07884v2 Announce Type: replace-cross Abstract: Using partial knowledge of a quantum state to control multiqubit entanglement is a largely unexplored paradigm in the emerging field of quantu

Sakana Fugu Technical Report

AgentsDGX agent

arXiv:2606.21228v1 Announce Type: new Abstract: The capabilities of frontier Large Language Models (LLMs) continue to advance, with different providers increasingly specializing in distinct domains. T

UECP: Uncertainty-Enhanced Collaborative Perception

AgentsDGX agent

arXiv:2606.23046v1 Announce Type: new Abstract: Collaborative perception serves as a pivotal solution to enhance the perception capability of individual agents in autonomous driving, where a core chal

22 Jun 2026

Something we’re thinking a bunch about as well Would love thoughts!

AgentsDGX agent

Something we’re thinking a bunch about as well Would love thoughts! context engineering docs for agentic engineering - plans, research, etc SHOULD NOT be stored in version control: A good docs managem

← Previous
1…155156157158159…300
Next →