AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
10 Jul 2026

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

Model ReleasesDGX agent

arXiv:2607.07985v1 Announce Type: cross Abstract: We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform,

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

SafetyDGX agent

arXiv:2607.07859v1 Announce Type: new Abstract: Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While h

The Hermes Desktop app can now discover and connect to your Hermes Cloud agents. Sign in with Nous Portal and any active Cloud instances are…

ResearchDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Hermes Desktop app can now discover and connect to your Hermes Cloud agents. Sign in with Nous Portal and any active Cloud instances are auto-discovered. http://portal.nousresearch.com/cloud Media

9 Jul 2026

As a result, we can show that open-source-derived models are not inherently untrustworthy. They can be suitable for production agents with c…

ApplicationsDGX agent

As a result, we can show that open-source-derived models are not inherently untrustworthy. They can be suitable for production agents with careful development, realistic evals, and targeted mitigation

Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation

Model ReleasesDGX agent

arXiv:2607.06957v1 Announce Type: cross Abstract: Realistic and diverse traffic simulation is essential to autonomous driving development. Yet prevailing benchmarks predominantly reward realism, and r

Got this setup in Claude Code to trace models routed through @merge_api! Being able to trace what your agents are doing is so valuable, I lo…

Model ReleasesDGX agent

Got this setup in Claude Code to trace models routed through @merge_api! Being able to trace what your agents are doing is so valuable, I love open source 🙌 We built a plugin that traces every Claude

langsmith for coding agents

Model ReleasesDGX agent

langsmith for coding agents We built a plugin that traces every Claude Code session straight into LangSmith. Three commands, one JSON block, and every message, tool call, and subagent run shows up as

Massive day for us @OpenAI: - GPT-5.6 SOTA at ~everything & by far most token efficient - Agents for everyone in the new ChatGPT app Work an…

Model ReleasesDGX agent

Massive day for us @OpenAI: - GPT-5.6 SOTA at ~everything & by far most token efficient - Agents for everyone in the new ChatGPT app Work and Codex modes - Work mode available on desktop (most powerfu

Mercor acquires Deeptune, which builds reinforcement learning environments for AI agents, three months after CEO Brendan Foody backed Deeptune's $43M Series A (Lily Mae Lazarus/Fortune)

IndustryDGX agent

Lily Mae Lazarus / Fortune: Mercor acquires Deeptune, which builds reinforcement learning environments for AI agents, three months after CEO Brendan Foody backed Deeptune's $43M Series A — Brendan Foo

Multi-Agent Robotic Control with Onboard Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.07403v1 Announce Type: cross Abstract: Vision Language Models (VLMs) and Vision Language Action (VLA) models have shown promise in robotic control. Yet, they face significant challenges reg

Notes on GPT-5.6, which includes some interesting new additions to the API (programmatic tool calling and multi-agent in particular) - plus …

Model ReleasesDGX agent

Notes on GPT-5.6, which includes some interesting new additions to the API (programmatic tool calling and multi-agent in particular) - plus 18 pelicans for the 6 reasoning levels and 3 new models: htt

OpenAI broadly releases GPT-5.6, and launches ChatGPT Work, an AI agent that can gather context across apps and files to create documents, on Mac and Windows (Axios)

Model ReleasesDGX agent

Axios: OpenAI broadly releases GPT-5.6, and launches ChatGPT Work, an AI agent that can gather context across apps and files to create documents, on Mac and Windows — - Sol is the most powerful versio

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

SafetyDGX agent

arXiv:2607.07436v1 Announce Type: new Abstract: A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the stru

8 Jul 2026

Announcing Grok 4.5, our first model trained specifically for coding and agents. It was trained with Cursor and offers frontier intelligence…

IndustryDGX agent

Announcing Grok 4.5, our first model trained specifically for coding and agents. It was trained with Cursor and offers frontier intelligence at leading speeds and cost efficiency. https://x.ai/news/gr

Benchmarking Coding Agents on Databricks’ Multi-Million Line Codebase

IndustryDGX agent

This Databricks blog post evaluates the performance and capabilities of coding agents when applied to real-world scenarios involving their own multi-million line codebase, likely assessing metrics suc

Busy couple days ahead! Today, @nvidia Nemotron 3 Ultra support for @LangChain Deep Agents at a fraction of the cost -w/ Quick Start Prompt …

Model ReleasesDGX agent

Busy couple days ahead! Today, @nvidia Nemotron 3 Ultra support for @LangChain Deep Agents at a fraction of the cost -w/ Quick Start Prompt Tomorrow is Wikimania🔥 @hwchase17 is chatting with with @Bra

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

Model ReleasesDGX agent

arXiv:2607.05682v1 Announce Type: new Abstract: LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first

Harnessing Code Agents for Automatic Software Verification

Model ReleasesDGX agent

arXiv:2607.06341v1 Announce Type: cross Abstract: Formal verification offers the strongest guarantee of software correctness, but it does not scale: the proofs demanded by interactive theorem provers

My team talks directly to *my* AI workforce. They can skip me entirely. It's the first time working with AI agents actually feels team-orien…

Model ReleasesDGX agent

My team talks directly to *my* AI workforce. They can skip me entirely. It's the first time working with AI agents actually feels team-oriented and collaborative (and not just one person becoming more

Nice stats on usage of open models across OpenCode. GLM-5.2 is still underrated, but one of the models that has really surprised me on agent…

Model ReleasesDGX agent

Nice stats on usage of open models across OpenCode. GLM-5.2 is still underrated, but one of the models that has really surprised me on agentic tasks is deepseek-v4-flash. Extremely cheap and effective

Onnes: A Physics-Grounded Multi-Agent LLM Simulator for Cryogenic Fault Diagnosis in Quantum Computing Infrastructure

Model ReleasesDGX agent

arXiv:2607.05805v1 Announce Type: new Abstract: Dilution refrigerators are the enabling infrastructure of superconducting quantum computers, yet their fault diagnosis is still dominated by threshold a

Prime Intellect raises $130M Series A to help enterprises build their own AI agents https://techcrunch.com/2026/07/08/prime-intellect-raises…

IndustryDGX agent

Prime Intellect raises $130M Series A to help enterprises build their own AI agents https://techcrunch.com/2026/07/08/prime-intellect-raises-130m-series-a-to-help-enterprises-build-their-own-ai-agents

To audit SWE-Bench Pro, we used model-based investigator agents alongside independent reviews from five independent experienced software eng…

Model ReleasesDGX agent

To audit SWE-Bench Pro, we used model-based investigator agents alongside independent reviews from five independent experienced software engineers. That helped us examine tasks at scale while keeping

7 Jul 2026

A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training

SafetyDGX agent

arXiv:2607.04574v1 Announce Type: cross Abstract: For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels condition on.

Agent-driven Long-tail Simulation for Autonomous Driving

Model ReleasesDGX agent

arXiv:2607.04331v1 Announce Type: cross Abstract: Evaluating autonomous driving systems in closed-loop settings requires realistic and interactive simulation, yet existing simulators largely rely on l

DrugAgent: Reliable Multi-Agent Integration of Conflicting Biomedical Evidence for Drug-Target Interaction Assessment

Model ReleasesDGX agent

arXiv:2408.13378v5 Announce Type: replace Abstract: Workflows in drug-target interaction (DTI) assessment require integrating heterogeneous data from predictive models, curated resources, and observat

GameEngineBench: Evaluating Coding Agents on Real C++ Runtime Environments

Model ReleasesDGX agent

arXiv:2607.03525v1 Announce Type: cross Abstract: Game engines provide real-time simulation, rendering, physics, interaction, networking, and asset pipelines, making them valuable not only for games b

PDEFlow: Autonomous Agentic PDE Pipelines for Neural Operator Learning and Solver-Free Inference

Model ReleasesDGX agent

arXiv:2607.05134v1 Announce Type: cross Abstract: We present PDEFlow, an autonomous agentic framework that turns user-level ODE and PDE descriptions into solver-backed neural-operator pipelines. The w

Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2607.04470v1 Announce Type: cross Abstract: Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent reinforcement

Sakana AI (@SakanaAILabs) is now a model vendor on Merge Gateway, and Fugu Ultra is live through them. It's a multi-agent orchestration mode…

Model ReleasesDGX agent

Sakana AI (@SakanaAILabs) is now a model vendor on Merge Gateway, and Fugu Ultra is live through them. It's a multi-agent orchestration model that routes across frontier models behind one API. You get

SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference

HardwareDGX agent

arXiv:2607.03333v1 Announce Type: cross Abstract: LLM agents are becoming a common interface for research, coding, and question answering, yet their Thought-Action-Observation loop is often serial: th

The release of Fable 5 just points to the importance of agent orchestration. You really don't need Fable 5 for most tasks. You can plan with…

Model ReleasesDGX agent

The release of Fable 5 just points to the importance of agent orchestration. You really don't need Fable 5 for most tasks. You can plan with Opus 4.8/Fable 5, execute with GPT-5.5, and design with GLM

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents

ResearchDGX agent

The web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories fo

6 Jul 2026

128 GB of memory is nice, but you can get started with local agentic AI workflows with much less. By connecting gemma 4 in @lmstudio to MATL…

Model ReleasesDGX agent

128 GB of memory is nice, but you can get started with local agentic AI workflows with much less. By connecting gemma 4 in @lmstudio to MATLAB MCP Server, you can run a local AI model that uses MATLAB

5 Jul 2026

As working with AI agents looks more like management, we may want to consider large-scale management training for the AI era. The US governm…

ApplicationsDGX agent

As working with AI agents looks more like management, we may want to consider large-scale management training for the AI era. The US government actually did this once, & the WW2 Engineering, Science,

ByteDance's Doubao and Alibaba's Qwen will disable humanlike and user-created agents before July 15, as China's anthropomorphic AI interaction rules take effect (Wency Chen/South China Morning Post)

Model ReleasesDGX agent

Wency Chen / South China Morning Post: ByteDance's Doubao and Alibaba's Qwen will disable humanlike and user-created agents before July 15, as China's anthropomorphic AI interaction rules take effect

4 Jul 2026

Learn why multimodal prompting is a big deal when working with coding agents.

TutorialsDGX agent

Multimodal prompting enhances coding agents by enabling them to process and integrate multiple types of input data—such as text, images, and code snippets—simultaneously, improving their ability to un

Multimodal prompting is clearly the future. How we interact with agents is evolving. I share a bit (including a video walkthrough) of how I …

TutorialsDGX agent

This post discusses the evolution of multimodal prompting in AI agent interactions, highlighting how users can now communicate with AI systems using multiple input types beyond text. The author provid

Q&A with Doug Brooks, senior product manager of Apple silicon, about Mac minis becoming preferred AI agent machines, future of on-device AI, and more (Jason Hiner/The Deep View)

Local AiDGX agent

Jason Hiner / The Deep View: Q&A with Doug Brooks, senior product manager of Apple silicon, about Mac minis becoming preferred AI agent machines, future of on-device AI, and more — W — alk into any of

3 Jul 2026

A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory

Model ReleasesDGX agent

arXiv:2607.01935v1 Announce Type: new Abstract: Long term memory lets LLM agents act as persistent assistants, but user facts change. A useful memory system must know what is true now, what used to be

A^{2}utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction

Model ReleasesDGX agent

arXiv:2607.02141v1 Announce Type: new Abstract: Most LP-from-text benchmarks are static datasets of word problems written and labeled by hand. Once such a dataset is released, its size is fixed, its d

Auto-FL-Research: Agentic Search for Federated Learning Algorithms

Local AiDGX agent

arXiv:2607.01366v1 Announce Type: new Abstract: Federated learning (FL) research often depends on many small but consequential algorithmic choices: optimizer variants, server aggregation rules, local

Bringing Agentic Search to Earth Observation Data Discovery

Model ReleasesDGX agent

arXiv:2607.02387v1 Announce Type: cross Abstract: NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding

Copewell: A Multi-Agent Swarm Architecture for Equitable Mental Wellness Support

SafetyDGX agent

arXiv:2607.02245v1 Announce Type: new Abstract: Mental health disorders affect nearly one billion people globally, yet 75% of individuals in low- and middle-income countries receive no treatment due t

Repair the Amplifier, Not the Symptom: Stable World-Model Correction for Agent Rollouts

Local AiDGX agent

arXiv:2607.01767v1 Announce Type: new Abstract: As agent planning moves from short tool chains toward persistent workflows with thousands or tens of thousands of steps, failures will occur inside larg

Understanding Agent-Based Patching of Compiler Missed Optimizations

Model ReleasesDGX agent

arXiv:2607.02370v1 Announce Type: cross Abstract: Compiler missed optimizations refer to cases in which compilers failed to optimize certain code. It takes many compiler developers' efforts to impleme

2 Jul 2026

big week at langchain, with a lot of launches: 1/ OpenWiki - auto generate a wiki of a github repo 2/ two different voice agent tutorials 3/…

Model ReleasesDGX agent

big week at langchain, with a lot of launches: 1/ OpenWiki - auto generate a wiki of a github repo 2/ two different voice agent tutorials 3/ Harbor integration and tutorial for long running, stateful

Cloudflare sets a September 15 deadline for AI companies to differentiate their web crawlers into search, AI training, and AI agents or face being blocked (Samantha Elkins/NBC News)

IndustryDGX agent

Samantha Elkins / NBC News: Cloudflare sets a September 15 deadline for AI companies to differentiate their web crawlers into search, AI training, and AI agents or face being blocked — Cloudflare gave

Hot take: I think it's still important to understand the code that our agents write! In this mega thread (based on my AIE talk today), I wil…

TutorialsDGX agent

Hot take: I think it's still important to understand the code that our agents write! In this mega thread (based on my AIE talk today), I will explain why that's the case, and show some ideas for how t

i finally tried hermes agent and the hype is real btw. @NousResearch cooked. been onboarding my young relatives who can't afford Claude, sho…

Model ReleasesDGX agent

i finally tried hermes agent and the hype is real btw. @NousResearch cooked. been onboarding my young relatives who can't afford Claude, showing them how to use $1-5 of tokens to bootstrap hermes and

Memo: Microsoft is merging the consumer and enterprise versions of its Copilot chatbots into a single app featuring coding tools and AI agents dubbed AutoPilot (The Information)

ApplicationsDGX agent

The Information: Memo: Microsoft is merging the consumer and enterprise versions of its Copilot chatbots into a single app featuring coding tools and AI agents dubbed AutoPilot — Microsoft is merging

Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments

SafetyDGX agent

arXiv:2607.00457v1 Announce Type: new Abstract: Embodied agents operating in the real world require multi-scale reasoning and knowledge adaptation as conditions change. We identify two challenges in a

Self-Evolving Agents with Anytime-Valid Certificates

SafetyDGX agent

arXiv:2607.00871v1 Announce Type: new Abstract: Self-evolving agents violate the assumption behind most learning-theoretic guarantees: the data, evaluator, components, and hypothesis space are produce

Your coding agent bill doubled and nobody can tell you why. Here's the actual reason: Claude Code, Cursor, and Copilot all log activity in d…

Model ReleasesDGX agent

Your coding agent bill doubled and nobody can tell you why. Here's the actual reason: Claude Code, Cursor, and Copilot all log activity in different formats. The second your team uses more than one (t

Z.ai launches ZCode, an 'Agentic Development Environment' optimized for its new GLM-5.2 model; Z.ai's GLM Coding Plan costs from 16.20 to 144 per month (Michael Nuñez/VentureBeat)

Model ReleasesDGX agent

Michael Nuñez / VentureBeat: Z.ai launches ZCode, an “Agentic Development Environment” optimized for its new GLM-5.2 model; Z.ai's GLM Coding Plan costs from 16.20 to 144 per month — The move marks th

1 Jul 2026

“Agentic kernel optimization is the future of on-device inference” @xenovacom used Fable 5 to write kernels that pushed Gemma 4 to a massive…

Model ReleasesDGX agent

“Agentic kernel optimization is the future of on-device inference” @xenovacom used Fable 5 to write kernels that pushed Gemma 4 to a massive 255 tok/s on WebGPU with M4. He shared the demo, so you can

AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance

HardwareDGX agent

arXiv:2606.30949v1 Announce Type: new Abstract: High-Level Synthesis (HLS) provides a fast path from concepts to silicon, but converting real-world software into synthesizable HLS code remains challen

AxDafny: Agentic Verified Code Generation in Dafny

Model ReleasesDGX agent

arXiv:2606.32007v1 Announce Type: new Abstract: We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny

do you know what you pay for in agentic workloads? cached tokens! session with 50+ tool calls -> prompt is billed 50 times all providers giv…

Model ReleasesDGX agent

do you know what you pay for in agentic workloads? cached tokens! session with 50+ tool calls -> prompt is billed 50 times all providers give 1/5 cached discount for GLM-5.2 we at @FireworksAI_HQ drop

ECHO: Prune to act, trace to learn with selective turn memory in agentic RL

SafetyDGX agent

arXiv:2606.31650v1 Announce Type: cross Abstract: Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Existing cont

← Previous
1…127128129130131…300
Next →