AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,964 results
16 Apr 2026

Alibaba unveils Qwen3.6-35B-A3B, an open-weight MoE model with 35B total and 3B active parameters, saying it rivals larger dense models in agentic coding tasks (Qwen)

Model ReleasesDGX agent

Qwen: Alibaba unveils Qwen3.6-35B-A3B, an open-weight MoE model with 35B total and 3B active parameters, saying it rivals larger dense models in agentic coding tasks — · 4355 words · QwenTeam丨Translat

🆕 Building pi in a World of Slop https://www.youtube.com/watch?v=RjfbvDXpFls @badlogicgames talks about why today's agents are still Mercha…

ToolsDGX agent

🆕 Building pi in a World of Slop https://www.youtube.com/watch?v=RjfbvDXpFls @badlogicgames talks about why today's agents are still Merchants of Learned Complexity, and gives 3 specific ways that hum

Copenhagen-based Spektr, which develops AI agents for compliance teams in financial services, raised a 20M Series A led by NEA, taking total funding to 26M (Mary Ann Azevedo/Crunchbase News)

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
IndustryDGX agent

Mary Ann Azevedo / Crunchbase News: Copenhagen-based Spektr, which develops AI agents for compliance teams in financial services, raised a 20M Series A led by NEA, taking total funding to 26M — For th

Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization

SafetyDGX agent

arXiv:2604.13533v1 Announce Type: cross Abstract: Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face li

Golden Handcuffs make safer AI agents

SafetyDGX agent

arXiv:2604.13609v1 Announce Type: new Abstract: Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand th

Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents

ApplicationsDGX agent

arXiv:2604.14004v1 Announce Type: cross Abstract: Memory-based self-evolution has emerged as a promising paradigm for coding agents. However, existing approaches typically restrict memory utilization

MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.13579v1 Announce Type: new Abstract: Conventional Retrieval-Augmented Generation (RAG) systems often struggle with complex multi-hop queries over long documents due to their single-pass ret

Opus 4.7 feels more intelligent, agentic, and precise than 4.6. It took a few days for me to learn how to work with it effectively, to fully…

Model ReleasesDGX agent

Opus 4.7 feels more intelligent, agentic, and precise than 4.6. It took a few days for me to learn how to work with it effectively, to fully take advantage of its new capabilities. Will post a few mor

Opus 4.7 is a model I’ve loved working with in Claude Code. It’s more agentic and instruction following but also incredibly smart and creati…

Model ReleasesDGX agent

Opus 4.7 is a model I’ve loved working with in Claude Code. It’s more agentic and instruction following but also incredibly smart and creative. I think it takes a slight adjustment to get used to, but

Opus 4.7 is in Claude Code today. It's more agentic, more precise, and a lot better at long-running work. It carries context across sessions…

Model ReleasesDGX agent

Opus 4.7 is in Claude Code today. It's more agentic, more precise, and a lot better at long-running work. It carries context across sessions and handles ambiguity much better. Introducing Claude Opus

Opus 4.7 is now supported in Hermes Agent 🚀🚀

Model ReleasesDGX agent

Opus 4.7 is now supported in Hermes Agent 🚀🚀 Introducing Claude Opus 4.7, our most capable Opus model yet. It handles long-running tasks with more rigor, follows instructions more precisely, and verif

Replit Agent 4 is even smarter now with Claude Opus 4.7! 50% off for a limited time. Go try it now ↓

Model ReleasesDGX agent

Replit Agent 4 has been upgraded to use Claude Opus 4.7, an advanced AI model, enhancing its code generation and problem-solving capabilities. The company is offering a 50% discount for a limited time

15 Apr 2026

5/5 The takeaway: If your agent relies on an LLM judge for selection accuracy, measuring code quality isn’t enough; you need a measure of th…

SafetyDGX agent

5/5 The takeaway: If your agent relies on an LLM judge for selection accuracy, measuring code quality isn’t enough; you need a measure of the model's inductive bias toward the 'fingerprint' of a gold

Enhancing Agentic Textual Graph Retrieval with Synthetic Stepwise Supervision

SafetyDGX agent

arXiv:2510.03323v2 Announce Type: replace Abstract: Integrating textual graphs into Large Language Models (LLMs) is promising for complex graph-based QA. However, a key bottleneck is retrieving inform

GCA Framework: A Gulf-Grounded Dataset and Agentic Pipeline for Climate Decision Support

Model ReleasesDGX agent

arXiv:2604.12306v1 Announce Type: cross Abstract: Climate decision-making in the Gulf increasingly demands systems that can translate heterogeneous scientific and policy evidence into actionable guida

Identity as Attractor: Geometric Evidence for Persistent Agent Architecture in LLM Activation Space

Model ReleasesDGX agent

arXiv:2604.12016v1 Announce Type: new Abstract: Large language models map semantically related prompts to similar internal representations -- a phenomenon interpretable as attractor-like dynamics. We

Let's talk parsing tables. Two days ago we launched ParseBench,the first document OCR benchmark built for AI agents. This deep dive breaks d…

Model ReleasesDGX agent

Let's talk parsing tables. Two days ago we launched ParseBench,the first document OCR benchmark built for AI agents. This deep dive breaks down TableRecordMatch (GTRM), our metric for evaluating compl

Self-Monitoring Benefits from Structural Integration: Lessons from Metacognition in Continuous-Time Multi-Timescale Agents

Model ReleasesDGX agent

arXiv:2604.11914v1 Announce Type: new Abstract: Self-monitoring capabilities -- metacognition, self-prediction, and subjective duration -- are often proposed as useful additions to reinforcement learn

14 Apr 2026

Agentic Application in Power Grid Static Analysis: Automatic Code Generation and Error Correction

Model ReleasesDGX agent

arXiv:2604.09995v1 Announce Type: cross Abstract: This paper introduces an LLM agent that automates power grid static analysis by converting natural language into MATPOWER scripts. The framework utili

Anthropic details using AI agents to accelerate alignment research on 'weak-to-strong supervision', where a weak model supervises the training of a stronger one (Anthropic)

SafetyDGX agent

Anthropic: Anthropic details using AI agents to accelerate alignment research on “weak-to-strong supervision”, where a weak model supervises the training of a stronger one — Large language models' eve

AutoMS: Multi-Agent Evolutionary Search for Cross-Physics Inverse Microstructure Design

Model ReleasesDGX agent

arXiv:2603.27195v2 Announce Type: replace Abstract: Designing microstructures with coupled cross-physics objectives is a fundamental challenge where traditional topology optimization is often computat

BLUEmed: Retrieval-Augmented Multi-Agent Debate for Clinical Error Detection

Model ReleasesDGX agent

arXiv:2604.10389v1 Announce Type: new Abstract: Terminology substitution errors in clinical notes, where one medical term is replaced by a linguistically valid but clinically different term, pose a pe

Claude Code Routines are here! In addition to a schedule, you can now trigger templated agents via GitHub event or API – with our infra & yo…

Model ReleasesDGX agent

Claude Code Routines are here! In addition to a schedule, you can now trigger templated agents via GitHub event or API – with our infra & your MCP+repos They've changed how we do docs, backlog mainten

Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions

TutorialsDGX agent

arXiv:2601.07516v2 Announce Type: replace-cross Abstract: Vision-language models are increasingly employed as multimodal conversational agents (MCAs) for diverse conversational tasks. Recently, reinfo

DERM-3R: A Resource-Efficient Multimodal Agents Framework for Dermatologic Diagnosis and Treatment in Real-World Clinical Settings

Model ReleasesDGX agent

arXiv:2604.09596v1 Announce Type: new Abstract: Dermatologic diseases impose a large and growing global burden, affecting billions and substantially reducing quality of life. While modern therapies ca

Dialectic-Med: Mitigating Diagnostic Hallucinations via Counterfactual Adversarial Multi-Agent Debate

SafetyDGX agent

arXiv:2604.11258v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) in healthcare suffer from severe confirmation bias, often hallucinating visual details to support initial, pote

EvoNash-MARL: A Closed-Loop Multi-Agent Reinforcement Learning Framework for Medium-Horizon Equity Allocation

SafetyDGX agent

arXiv:2604.10911v1 Announce Type: new Abstract: Medium-to-long-horizon stock allocation presents significant challenges due toveak predictive structures, non-stadonary market regimes, and the degradat

I just updated our license. For personal use, you’re free to run the software on your own servers for coding, building applications, agents,…

IndustryDGX agent

I just updated our license. For personal use, you’re free to run the software on your own servers for coding, building applications, agents, tools, or integrations, as well as for research, experiment

llama.cpp built from source + qwen3.5 27b running locally + hermes agent on top + camofox for scraping (tls spoofing) + scweet to scrape x (…

Model ReleasesDGX agent

llama.cpp built from source + qwen3.5 27b running locally + hermes agent on top + camofox for scraping (tls spoofing) + scweet to scrape x (no api keys) + tailscale to access from other devices bro Me

M2.7 w/ hermes cli is replacing ~75% of my claude code / opus usage now, but we need clarity for using it as a coding agent @ work. We're tr…

Model ReleasesDGX agent

M2.7 w/ hermes cli is replacing ~75% of my claude code / opus usage now, but we need clarity for using it as a coding agent @ work. We're truly blessed to have the weights of this one, looking forward

RCBSF: A Multi-Agent Framework for Automated Contract Revision via Stackelberg Game

Model ReleasesDGX agent

arXiv:2604.10740v1 Announce Type: new Abstract: Despite the widespread adoption of Large Language Models (LLMs) in Legal AI, their utility for automated contract revision remains impeded by hallucinat

RISK: A Framework for GUI Agents in E-commerce Risk Management

Model ReleasesDGX agent

arXiv:2509.21982v2 Announce Type: replace Abstract: E-commerce risk management requires aggregating diverse, deeply embedded web data through multi-step, stateful interactions, which traditional scrap

SCMAPR: Self-Correcting Multi-Agent Prompt Refinement for Complex-Scenario Text-to-Video Generation

Model ReleasesDGX agent

arXiv:2604.05489v3 Announce Type: replace Abstract: Text-to-Video (T2V) generation has benefited from recent advances in diffusion models, yet current systems still struggle under complex scenarios, w

The Missing Knowledge Layer in Cognitive Architectures for AI Agents

Model ReleasesDGX agent

arXiv:2604.11364v1 Announce Type: new Abstract: The two most influential cognitive architecture frameworks for AI agents, CoALA [21] and JEPA [12], both lack an explicit Knowledge layer with its own p

This is why we released liteparse :) Free, open-source, designed for agents. Natively supports OCR / screenshotting for deeper visual unders…

Model ReleasesDGX agent

This is why we released liteparse :) Free, open-source, designed for agents. Natively supports OCR / screenshotting for deeper visual understanding in a document when needed. @kepano I just tried it t

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

Model ReleasesDGX agent

arXiv:2604.11557v1 Announce Type: new Abstract: Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external systems through structured function calls. However

ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM Agents

Model ReleasesDGX agent

arXiv:2508.04038v2 Announce Type: replace Abstract: Motion sensor time-series are central to Human Activity Recognition (HAR), yet conventional approaches are constrained to fixed activity sets and ty

13 Apr 2026

AlphaLab: Autonomous Multi-Agent Research Across Optimization Domains with Frontier LLMs

Model ReleasesDGX agent

arXiv:2604.08590v1 Announce Type: cross Abstract: We present AlphaLab, an autonomous research harness that leverages frontier LLM agentic capabilities to automate the full experimental cycle in quanti

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?

Model ReleasesDGX agent

arXiv:2604.09251v1 Announce Type: new Abstract: Deep research agents increasingly interleave web browsing with multi-step computation, yet existing benchmarks evaluate these capabilities in isolation,

Ornn Compute Price Index: renting one Nvidia Blackwell GPU for an hour now costs 4.08, up 48% from 2.75 two months ago, driven by agentic AI demand (Wall Street Journal)

HardwareDGX agent

Wall Street Journal: Ornn Compute Price Index: renting one Nvidia Blackwell GPU for an hour now costs 4.08, up 48% from 2.75 two months ago, driven by agentic AI demand — AI companies are rationing of

PilotBench: A Benchmark for General Aviation Agents with Safety Constraints

Model ReleasesDGX agent

arXiv:2604.08987v1 Announce Type: new Abstract: As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained

Process Reward Agents for Steering Knowledge-Intensive Reasoning

Local AiDGX agent

arXiv:2604.09482v1 Announce Type: new Abstract: Reasoning in knowledge-intensive domains remains challenging as intermediate steps are often not locally verifiable: unlike math or code, evaluating ste

Trending well! We're glad the traces are useful to the @NousResearch Hermes Agent community. A third batch is in progress.

Model ReleasesDGX agent

Trending well! We're glad the traces are useful to the @NousResearch Hermes Agent community. A third batch is in progress. Very cool open-source traces from @TheZachMueller @LambdaAPI: https://hugging

12 Apr 2026

Hermes Agent + Ollama returns tool JSON but doesn’t actually execute anything

Local AiDGX agent

Users building agentic pipelines with Hermes models in Ollama report an issue where the model correctly generates tool call JSON in its response but the actual tool functions are never invoked or exec

@hwchase17 I think harness/managed agents is a way for Anthropic to keep its Moat. As models get mature, the need for cloud LLMs might reduc…

Local AiDGX agent

@hwchase17 I think harness/managed agents is a way for Anthropic to keep its Moat. As models get mature, the need for cloud LLMs might reduce and local LLM models might increase (save cost) and thus t

10 Apr 2026

I compared sandbox options for AI agents. Here’s my ranking.

Local AiDGX agent

The search did not return the specific Reddit post. However, result index 2 appears to be a closely related article on the same topic. Let me use what's available to craft an accurate summary based...

On Emotion-Sensitive Decision Making of Small Language Model Agents

Model ReleasesDGX agent

arXiv:2604.06562v1 Announce Type: new Abstract: Small language models (SLM) are increasingly used as interactive decision-making agents, yet most decision-oriented evaluations ignore emotion as a caus

8 Apr 2026

I care because understanding which index helps me understand things like what I need to submit content for crawling to, what search agents I…

ToolsDGX agent

I care because understanding which index helps me understand things like what I need to submit content for crawling to, what search agents I should allow in robots.txt, how frequently I can expect the

Mornings feel different lately. Before I even get organized, I already have 5+ agents running. Same thing before I go to sleep. Claude Code,…

Model ReleasesDGX agent

Mornings feel different lately. Before I even get organized, I already have 5+ agents running. Same thing before I go to sleep. Claude Code, ChatGPT, Qodo, Notion, Linear, NotebookLM. All moving in pa

14 Aug 2026

Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting

Local AiDGX agent

arXiv:2608.12590v1 Announce Type: new Abstract: Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these

Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language

SafetyDGX agent

arXiv:2608.13029v1 Announce Type: cross Abstract: The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an n

13 Aug 2026

Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System

Model ReleasesDGX agent

arXiv:2608.11738v1 Announce Type: cross Abstract: Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses distinct chal

AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection

Model ReleasesDGX agent

arXiv:2608.11679v1 Announce Type: new Abstract: Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreting anomalies

Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence

Model ReleasesDGX agent

arXiv:2608.12290v1 Announce Type: cross Abstract: Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and rel

Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning

Model ReleasesDGX agent

arXiv:2608.11260v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals. Existing approaches exhibit a 'when-what' dissoci

12 Aug 2026

A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem

AgentsDGX agent

arXiv:2608.10760v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has become the de-facto interface for connecting LLM agents to enterprise tools, and adoption has been explosive: wit

Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents

SafetyDGX agent

arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expen

Grok is now an AI ‘teammate’ you can assign work

AgentsDGX agent

SpaceXAI has introduced Grok Bot, an always-on AI agent service designed to behave like independent 'AI teammates' that can do your work for you. The bots share their own cloud-based computer environm

How OneAdvanced deployed over 50 AI agents on UK-sovereign AWS

Model ReleasesDGX agent

Learn how OneAdvanced, a UK enterprise software provider, built a UK-sovereign AI platform by self-hosting Llama 4 Maverick and Llama Guard 4 on Amazon SageMaker AI, with a RAG pipeline on pgvector an

Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization

AgentsDGX agent

arXiv:2608.10694v1 Announce Type: cross Abstract: Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs an answerin

← Previous
1…136137138139140…300
Next →