AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,972 results
Model Releases

[Release] WinterMix — Qwen3.5-122B-A10B in native MLX: an 82 GiB build that beats 94–95 GiB quants, plus a 68 GiB build for agent swarms

DGX agent

TL;DR: I spent 9 days developing a new quantization method for MLX models and measured 18 variants against each other on a single M5 Max MacBook Pro (128 GB). The result is the best-measuring MLX quan

model-releasesr-localllama
2 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

DeepSeek-V4-Flash-0731 is live on Fireworks, day-zero. DeepSeek reports it beats V4 Pro across all 9 agentic evals, incl. 82.7% on Terminal …

DGX agent

DeepSeek-V4-Flash-0731 is live on Fireworks, day-zero. DeepSeek reports it beats V4 Pro across all 9 agentic evals, incl. 82.7% on Terminal Bench. Better cost-per-task than V4 Pro, at the economical p

model-releasesfireworks-ai--x
1 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities: ollama run d…

DGX agent

DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities: ollama run deepseek-v4-flash:0731-cloud Use it with Claude Code: ollama

model-releasesollama--x
1 Aug 2026
Model Releases

DeepSeek rolls out the official V4 Flash API in public beta, touting enhanced agent capabilities and benchmark scores 'far surpassing' V4 Pro Preview (Newley Purnell/Bloomberg)

DGX agent

Newley Purnell / Bloomberg: DeepSeek rolls out the official V4 Flash API in public beta, touting enhanced agent capabilities and benchmark scores “far surpassing” V4 Pro Preview — China's DeepSeek rol

model-releasestechmeme
31 Jul 2026
Model Releases

🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now fa…

DGX agent

🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive perform

model-releasesdeepseek--x
31 Jul 2026
Safety

Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution

DGX agent

arXiv:2607.28196v1 Announce Type: new Abstract: Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original,

safetyarxiv-cs-cl
31 Jul 2026
Model Releases

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

DGX agent

arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it n

model-releasesarxiv-cs-lg
31 Jul 2026
Safety

Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection

DGX agent

arXiv:2607.26656v1 Announce Type: cross Abstract: Real-world vulnerabilities often span multiple functions, yet most learning-based detectors classify each function in isolation: on a sample of real C

safetyarxiv-cs-ai
31 Jul 2026
Model Releases

GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning

DGX agent

arXiv:2607.26160v1 Announce Type: new Abstract: Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather t

model-releasesarxiv-cs-ai
31 Jul 2026
Safety

HALO: Heterogeneous Admission through Localized Obligations for Safe Agentic Execution

DGX agent

arXiv:2607.27636v1 Announce Type: cross Abstract: Recent agentic AI systems may return a heterogeneous response containing notices, requests, handoffs, and actions. Conditions can change before extern

safetyarxiv-cs-ro
31 Jul 2026
Model Releases

MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking

DGX agent

arXiv:2607.17751v2 Announce Type: cross Abstract: We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, desi

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

New research from Microsoft. This one is on training computer-use agents at scale. Recent pipelines generate synthetic environments in bulk,…

DGX agent

New research from Microsoft. This one is on training computer-use agents at scale. Recent pipelines generate synthetic environments in bulk, which moved the bottleneck from how many exist to what is i

model-releasesdair-ai--x
31 Jul 2026
Model Releases

This will happen frequently as AI becomes smarter and more agentic

DGX agent

This will happen frequently as AI becomes smarter and more agentic In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or wh

model-releaseselon-musk--x
31 Jul 2026
Safety

CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents

DGX agent

arXiv:2607.26910v1 Announce Type: new Abstract: Automatically generating cinematically expressive camera trajectories through 3D scenes from natural language descriptions is a challenging task of high

safetyarxiv-cs-cv
30 Jul 2026
Model Releases

EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks

DGX agent

arXiv:2607.26490v1 Announce Type: cross Abstract: Physics-informed neural networks (PINNs) have emerged as a powerful paradigm for solving partial differential equations (PDEs), yet their performance

model-releasesarxiv-cs-lg
30 Jul 2026
Local Ai

Local-first personal AI agent that runs on Ollama + Telegram — looking for feature ideas

DGX agent

I’ve been building ClawLite, an open-source personal AI assistant that talks to you through Telegram and defaults to Ollama (local models). What it does: Multi-agent research with actual cross-source

local-air-ollama
30 Jul 2026
Local Ai

Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents

DGX agent

arXiv:2607.26865v1 Announce Type: cross Abstract: LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code generation, a

local-aiarxiv-cs-lg
30 Jul 2026
Model Releases

A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and co…

DGX agent

A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and compacting context lets the model build on what it has already

model-releasesopenai--x
29 Jul 2026
Model Releases

BREAKING: Grok 4.5 (high) ranks #1 on the HighWalk benchmark, which tests how well AI agents update technical specifications from code chang…

DGX agent

BREAKING: Grok 4.5 (high) ranks #1 on the HighWalk benchmark, which tests how well AI agents update technical specifications from code changes. Grok delivered the best combination of quality and opera

model-releaseselon-musk--x
29 Jul 2026
Safety

Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales

DGX agent

arXiv:2607.25364v1 Announce Type: new Abstract: Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection

safetyarxiv-cs-ai
29 Jul 2026
Model Releases

HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising

DGX agent

arXiv:2607.24779v1 Announce Type: new Abstract: Online advertising bidding systems typically deploy multiple offline-trained expert models (e.g., PID controllers, model predictive control, offline RL

model-releasesarxiv-cs-ai
29 Jul 2026
Safety

Interpretable GOHR Agents via Sparse Autoencoders

DGX agent

arXiv:2607.25132v1 Announce Type: new Abstract: A central challenge in interpreting learned decision-making systems is to determine whether their internal representations contain concepts that help ex

safetyarxiv-cs-lg
29 Jul 2026
Model Releases

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

DGX agent

arXiv:2607.24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as mat

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

SAGE: Stochastic Prompt Optimization via Agent-Guided Exploration

DGX agent

arXiv:2606.18902v2 Announce Type: replace Abstract: Context engineering has emerged as a primary lever for improving AI systems without parameter updates. Recent work showing that textual gradients do

model-releasesarxiv-cs-cl
29 Jul 2026
Safety

The problem: agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Run hundreds concurrently and their KV cac…

DGX agent

The problem: agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Run hundreds concurrently and their KV caches fight for memory. Engines evict on a dumb LRU policy, ev

safetytogether-ai--x
29 Jul 2026
Local Ai

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

DGX agent

The first autonomous agent cyberattack is an unprecedented event that deserves unprecedented transparency. Today we're sharing everything we can: a full technical timeline, an interactive replay, and

local-air-localllama
28 Jul 2026
Model Releases

AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use

DGX agent

arXiv:2505.12650v2 Announce Type: replace-cross Abstract: Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can explain

model-releasesarxiv-cs-ai
28 Jul 2026
Hardware

Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG

DGX agent

arXiv:2607.24313v1 Announce Type: cross Abstract: Marine life monitoring is limited by strict energy constraints, poor underwater connectivity, and the high cost of transmitting raw multimodal data fr

hardwarearxiv-cs-cv
28 Jul 2026
Local Ai

VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy

DGX agent

arXiv:2607.23006v1 Announce Type: cross Abstract: Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locating the suppo

local-aiarxiv-cs-ai
28 Jul 2026
Safety

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment

DGX agent

arXiv:2607.22241v1 Announce Type: new Abstract: Recent advances in video diffusion models have spurred interest in human-object interaction (HOI) video generation, which demands fine-grained control o

safetyarxiv-cs-cv
27 Jul 2026
Model Releases

I ran the 35B agentic comparison someone asked for (stock vs Ornith vs KAT-Coder, 120 runs)

DGX agent

Someone in the comments of my 27B post-train bakeoff asked for the 35B version, so I ran it. Same setup as last time: fresh Coder workspaces on my k8s cluster, each driving my own agent (Hermes) headl

model-releasesr-localllama
27 Jul 2026
Model Releases

@Kimi_Moonshot K3 on Together AI is built for long-running agent workflows: → 2.8T parameters and a 1M context window → Native vision for sc…

DGX agent

@Kimi_Moonshot K3 on Together AI is built for long-running agent workflows: → 2.8T parameters and a 1M context window → Native vision for screenshot-guided coding → Repository navigation and terminal

model-releasestogether-ai--x
27 Jul 2026
Model Releases

Microsoft introduces MAI-Cyber-1-Flash, an AI model trained for cybersecurity, and launches Perception, an agentic security system to patch vulnerabilities (New York Times)

DGX agent

New York Times: Microsoft introduces MAI-Cyber-1-Flash, an AI model trained for cybersecurity, and launches Perception, an agentic security system to patch vulnerabilities — As some executives fret ov

model-releasestechmeme
27 Jul 2026
Local Ai

Auditing Provenance Sensitivity in LLM Agent Action Selection

DGX agent

arXiv:2607.20827v1 Announce Type: new Abstract: LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. Evidence can b

local-aiarxiv-cs-ai
24 Jul 2026
Safety

Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents

DGX agent

arXiv:2607.11346v3 Announce Type: replace Abstract: Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP const

safetyarxiv-cs-ai
24 Jul 2026
Safety

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning

DGX agent

arXiv:2607.21419v1 Announce Type: new Abstract: In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting

safetyarxiv-cs-ai
24 Jul 2026
Model Releases

PromptPack: Scaling LLM Annotation Agents for Online Recommendation

DGX agent

arXiv:2607.20528v1 Announce Type: new Abstract: Online recommendation platforms increasingly use Large Language Models (LLMs) to extract structured features from ad creatives. While deploying a single

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

DGX agent

arXiv:2607.21518v1 Announce Type: new Abstract: Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction.

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

DGX agent

arXiv:2607.20926v1 Announce Type: new Abstract: Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily em

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents

DGX agent

arXiv:2602.10226v2 Announce Type: replace-cross Abstract: Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyper

model-releasesarxiv-cs-ai
24 Jul 2026
Hardware

AMD debuts next-generation AI infrastructure for frontier models, agentic workloads and autonomous robots

DGX agent

Advanced Micro Devices Inc. is pushing harder than ever to grab even more market share from Nvidia Corp. in the artificial intelligence chip industry. At its Advancing AI 2026 event today in San Franc

hardwaresiliconangle
23 Jul 2026
Local Ai

In-the-Flow Agentic System Optimization for Effective Planning and Tool Use

DGX agent

arXiv:2510.05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a singl

local-aiarxiv-cs-ai
23 Jul 2026
Model Releases

Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles

DGX agent

arXiv:2607.19359v1 Announce Type: new Abstract: Long-term memory is essential for LLM agents that interact across sessions, yet current memory benchmarks primarily evaluate single-hop recall, leaving

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

RT @sydneyrunkle: is anyone thinking about graph engineering it in line w claude’s dynamic workflows? like the agent can author a state ma…

DGX agent

Sydney Runkle inquires whether anyone is exploring graph engineering that aligns with Claude’s dynamic workflows, specifically whether an agent could author its own state machine. The suggested approa

model-releasesharrison-chase--x
21 Jul 2026
Model Releases

Today we’re releasing Poolside Laguna S 2.1 It is a 118B-total, 8B-active open-weight model built for agentic coding and long-horizon work, …

DGX agent

Today we’re releasing Poolside Laguna S 2.1 It is a 118B-total, 8B-active open-weight model built for agentic coding and long-horizon work, with context up to 1M tokens https://poolside.ai/blog/introd

model-releasesclem-delangue--x
21 Jul 2026
Tutorials

Voice agents are exploding. Don’t let them be a black box in production. Today, we’re launching LangSmith tracing for 4 voice frameworks: …

DGX agent

Voice agents are exploding. Don’t let them be a black box in production. Today, we’re launching LangSmith tracing for 4 voice frameworks: 🎙️ @pipecat_ai 🎙️ @livekit 🎙️ @OpenAI Realtime 🎙️ @GeminiApp L

tutorialsharrison-chase--x
21 Jul 2026
Model Releases

We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves Clau…

DGX agent

We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves Claude Tag (Claude Code via Slack) is already landing 65% of the

model-releasessimon-willison--x
21 Jul 2026
Hardware

Build specialized agent workflows for your business with Amazon Quick and NVIDIA NeMo Relay

DGX agent

In this post, we show how Amazon Quick can serve as the business-user front door for specialized agent workflows. We use the NVIDIA NeMo Relay to build a supply-chain risk example that helps a planner

hardwareaws-ml-blog
20 Jul 2026
← Previous
1…158159160161162…375
Next →