AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
27 Apr 2026

Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives

ResearchDGX agent

arXiv:2603.26041v3 Announce Type: replace Abstract: In recent years, GUI visual agents built upon Multimodal Large Language Models (MLLMs) have demonstrated strong potential in navigation tasks. Howev

SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning

AgentsDGX agent

arXiv:2604.22558v1 Announce Type: cross Abstract: As Multimodal Large Language Models (MLLMs) mature, GUI agents are evolving from static interactions to complex navigation. While Reinforcement Learni

The greatest con of the decade was calling autocomplete “AI”. The second greatest is calling autocomplete-in-a-loop an “agent”.

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Gary Marcus critiques the marketing of large language model autocomplete systems as 'artificial intelligence,' arguing this terminology misrepresents their actual capabilities. He extends this critici

The terminal hasn’t changed much since the 1970s. What you do with it has. Introducing Devin for Terminal: everything we learned building De…

AgentsDGX agent

The terminal hasn’t changed much since the 1970s. What you do with it has. Introducing Devin for Terminal: everything we learned building Devin, now as a local agent, available right in your shell. An

26 Apr 2026

7) Multi-agent design. I loved the design of Cove (they were the acquired by Microsoft). It was more of a whiteboard than tabs or chat threa…

Model ReleasesDGX agent

7) Multi-agent design. I loved the design of Cove (they were the acquired by Microsoft). It was more of a whiteboard than tabs or chat threads. I don’t think we’ve cracked the right UI for managing ag

24 Apr 2026

DeepSeek-V4: a million-token context that agents can actually use

Model ReleasesDGX agent

DeepSeek-V4 is an advanced language model featuring a million-token context window that enables practical agentic applications beyond simple retrieval. The model demonstrates improved efficiency and u

GPT-5.5 is now available in Devin as an Agent Preview! GPT-5.5 has set a new bar for what's possible with Devin. It runs longer and more aut…

Model ReleasesDGX agent

GPT-5.5 is now available in Devin as an Agent Preview! GPT-5.5 has set a new bar for what's possible with Devin. It runs longer and more autonomously than any GPT model we've tested, surfacing bugs no

Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework

AgentsDGX agent

arXiv:2604.21090v1 Announce Type: cross Abstract: AI governance programmes increasingly rely on natural language prompts to constrain and direct AI agent behaviour. These prompts function as executabl

Tool Attention Is All You Need

AgentsDGX agent

Tool Attention Is All You Need // Tool Attention Is All You Need // New research proposes a practical fix for the hidden 'MCP tax.' The work introduces a dynamic tool gating mechanism built on an Inte

23 Apr 2026

ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration

SafetyDGX agent

arXiv:2604.19856v1 Announce Type: cross Abstract: Large Language Models (LLMs) show promise for generating Register-Transfer Level (RTL) code from natural language specifications, but single-shot gene

Cyera acquires Ryft to give enterprises traceable data access for AI agents

IndustryDGX agent

Artificial intelligence and data security company Cyera Ltd. announced today that it has acquired Ryft Data Inc., an Israeli startup with an automated data lake platform designed for enterprises deplo

Environmental Understanding Vision-Language Model for Embodied Agent

SafetyDGX agent

arXiv:2604.19839v1 Announce Type: cross Abstract: Vision-language models (VLMs) have shown strong perception and reasoning abilities for instruction-following embodied agents. However, despite these a

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards

SafetyDGX agent

arXiv:2604.20486v1 Announce Type: new Abstract: Training multimodal agents via reinforcement learning for knowledge-intensive visual reasoning is fundamentally hindered by the extreme sparsity of outc

zero shot Kimi K2.6, go try it out its a good model sir! this is @Kimi_Moonshot running on @togethercompute, @opencode harness prompt below…

AgentsDGX agent

zero shot Kimi K2.6, go try it out its a good model sir! this is @Kimi_Moonshot running on @togethercompute, @opencode harness prompt below👇 Media Introducing Kimi K2.6 from @Kimi_Moonshot, a multimod

22 Apr 2026

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps

Model ReleasesDGX agent

arXiv:2604.19533v1 Announce Type: cross Abstract: We introduce the Cyber Defense Benchmark, a benchmark for measuring how well large language model (LLM) agents perform the core SOC analyst task of th

Now Meta will track what employees do on their computers to train its AI agents

IndustryDGX agent

Meta employees' activity at work is now being used to train the company's AI agents. As reported by Reuters, Meta is installing a tool it calls Model Capability Initiative (MCI) on US-based employees'

SAGE-32B: Agentic Reasoning via Iterative Distillation

Model ReleasesDGX agent

arXiv:2601.04237v2 Announce Type: replace Abstract: We demonstrate SAGE-32B, a 32 billion parameter language model that focuses on agentic reasoning and long range planning tasks. Unlike chat models t

21 Apr 2026

Boomi builds a role for agents and guardrails in the data-connected enterprise

ApplicationsDGX agent

Artificial intelligence and machine learning have been a central focus for Boomi LP over much of its 26 years as a data activation company. The organization began storing anonymized metadata from cust

Generative midtended cognition and Artificial Intelligence. Thinging with thinging things

AgentsDGX agent

arXiv:2411.06812v2 Announce Type: replace-cross Abstract: This paper introduces the concept of ``generative midtended cognition'', exploring the integration of generative AI with human cognition. The

Kimi is the current open-source SOTA on Artificial Analysis

AgentsDGX agent

Kimi is the current open-source SOTA on Artificial Analysis Moonshot’s Kimi K2.6 is the new leading open weights model. Kimi K2.6 lands at #4 on the Artificial Analysis Intelligence Index (54) behind

Let's talk parsing charts 📊📈. Last week we released ParseBench, the first document OCR benchmark for AI agents. New in ParseBench: ChartDa…

Model ReleasesDGX agent

Let's talk parsing charts 📊📈. Last week we released ParseBench, the first document OCR benchmark for AI agents. New in ParseBench: ChartDataPointMatch. Most document look at a chart and OCR the captio

OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection

SafetyDGX agent

arXiv:2511.21064v2 Announce Type: replace-cross Abstract: Open-Vocabulary Object Detection (OVOD) aims to enable detectors to generalize across categories by leveraging semantic information. Although

SUSE and Vultr’s open cloud infrastructure push goes global

AgentsDGX agent

Global AI ambitions keep colliding with the same wall: cloud infrastructure that can’t keep up with performance demands without compromising where data lives and who controls it. As organizations look

20 Apr 2026

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

Model ReleasesDGX agent

arXiv:2604.15715v1 Announce Type: cross Abstract: The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows

InfoChess: A Game of Adversarial Inference and a Laboratory for Quantifiable Information Control

AgentsDGX agent

arXiv:2604.15373v1 Announce Type: cross Abstract: We propose InfoChess, a symmetric adversarial game that elevates competitive information acquisition to the primary objective. There is no piece captu

Long-Term Memory for VLA-based Agents in Open-World Task Execution

SafetyDGX agent

arXiv:2604.15671v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated significant potential for embodied decision-making; however, their application in complex chemical

PolicyBank: Evolving Policy Understanding for LLM Agents

Model ReleasesDGX agent

arXiv:2604.15505v1 Announce Type: cross Abstract: LLM agents operating under organizational policies must comply with authorization constraints typically specified in natural language. In practice, su

We recently shipped quality-of-life improvements to the Cursor CLI to make working with agents in the terminal more delightful. Use /debug t…

ToolsDGX agent

We recently shipped quality-of-life improvements to the Cursor CLI to make working with agents in the terminal more delightful. Use /debug to find root causes and fix tricky bugs that are hard to repr

17 Apr 2026

Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems

Model ReleasesDGX agent

arXiv:2604.14228v1 Announce Type: cross Abstract: Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study describes

Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization

SafetyDGX agent

arXiv:2604.14267v1 Announce Type: new Abstract: Search agents extend Large Language Models (LLMs) beyond static parametric knowledge by enabling access to up-to-date and long-tail information unavaila

LLM agents loop, drift, and get stuck on hard reasoning tasks up to 30% of the time. Current fixes are either too blunt (hard step limits) o…

TutorialsDGX agent

LLM agents loop, drift, and get stuck on hard reasoning tasks up to 30% of the time. Current fixes are either too blunt (hard step limits) or too expensive (LLM-as-judge adding 10-15% overhead per ste

SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment

Model ReleasesDGX agent

arXiv:2604.13630v1 Announce Type: cross Abstract: The performance of large language model (LLM) agents depends critically on the execution harness, the system layer that orchestrates tool use, context

VeruSAGE: A Study of Agent-Based Verification for Rust Systems

Model ReleasesDGX agent

arXiv:2512.18436v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown impressive capability to understand and develop code. However, their capability to rigorously reason a

16 Apr 2026

AI labs are buying Slack, Jira, and email archives from defunct startups to build 'reinforcement learning gyms' and train AI agents in simulated workplaces (Anna Tong/Forbes)

IndustryDGX agent

Anna Tong / Forbes: AI labs are buying Slack, Jira, and email archives from defunct startups to build “reinforcement learning gyms” and train AI agents in simulated workplaces — Defunct startups are b

LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

Model ReleasesDGX agent

arXiv:2604.13072v1 Announce Type: new Abstract: LLM-based agents are increasingly expected to handle real-world assistant tasks, yet existing benchmarks typically evaluate them under isolated sources

⚡ Meet Qwen3.6-35B-A3B:Now Open-Source!🚀🚀 A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. 🔥 Agentic coding on par wi…

Model ReleasesDGX agent

⚡ Meet Qwen3.6-35B-A3B:Now Open-Source!🚀🚀 A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. 🔥 Agentic coding on par with models 10x its active size 📷 Strong multimodal perception an

Nous Research🤝fal Congratulations to @NousResearch on the Tool Gateway launch. Developer-first infrastructure is what will define the next …

AgentsDGX agent

Nous Research🤝fal Congratulations to @NousResearch on the Tool Gateway launch. Developer-first infrastructure is what will define the next phase of agentic applications. Tool Gateway is now live in No

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration

Model ReleasesDGX agent

arXiv:2604.14116v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have empowered AI research agents to perform isolated scientific tasks, automating complex, real-world workflows, s

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

Model ReleasesDGX agent

arXiv:2512.14234v2 Announce Type: replace Abstract: Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human

15 Apr 2026

AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition

Model ReleasesDGX agent

arXiv:2604.12735v1 Announce Type: new Abstract: LLM-based multimodal emotion recognition relies on static parametric memory and often hallucinates when interpreting nuanced affective states. In this p

Best part are retrieval and search details: >Agent queries differ from human queries & what good results means changes too >Parallel queries…

HardwareDGX agent

Best part are retrieval and search details: >Agent queries differ from human queries & what good results means changes too >Parallel queries and ranking are both tools to the same outcome >Top-K preci

`deepagents deploy` now supports user scoped memory! add a user/ directory in your project so each user gets their own writable AGENTS.md, s…

AgentsDGX agent

`deepagents deploy` now supports user scoped memory! add a user/ directory in your project so each user gets their own writable AGENTS.md, seeded on first deploy and persisted across conversations. yo

Drawing on Memory: Dual-Trace Encoding Improves Cross-Session Recall in LLM Agents

Model ReleasesDGX agent

arXiv:2604.12948v1 Announce Type: new Abstract: LLM agents with persistent memory store information as flat factual records, providing little context for temporal reasoning, change tracking, or cross-

Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

SafetyDGX agent

arXiv:2604.12616v1 Announce Type: new Abstract: The rapid evolution of Vision-Language Models (VLMs) has catalyzed unprecedented capabilities in artificial intelligence; however, this continuous modal

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization

Model ReleasesDGX agent

arXiv:2604.12290v1 Announce Type: new Abstract: Current LLM agent benchmarks, which predominantly focus on binary pass/fail tasks such as code generation or search-based question answering, often negl

Is Gemma 4 26B MoE or 31B good as an MCP agent for coding with Xcode?

Model ReleasesDGX agent

This r/ollama discussion explores the suitability of Google's Gemma 4 models — specifically the 26B Mixture of Experts (MoE) and 31B Dense variants — as MCP (Model Context Protocol) agents for coding

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning

SafetyDGX agent

arXiv:2601.06794v2 Announce Type: replace Abstract: Critique-guided reinforcement learning (RL) has emerged as a powerful paradigm for training LLM agents by augmenting sparse outcome rewards with nat

Policy-Invisible Violations in LLM-Based Agents

Model ReleasesDGX agent

arXiv:2604.12177v1 Announce Type: new Abstract: LLM-based agents can execute actions that are syntactically valid, user-sanctioned, and semantically appropriate, yet still violate organizational polic

Q&A with ElevenLabs co-founder Mati Staniszewski on how audio models work, the company's business model, the conversational Turing Test, voice agents, and more (John Collison/Cheeky Pint)

IndustryDGX agent

John Collison / Cheeky Pint: Q&A with ElevenLabs co-founder Mati Staniszewski on how audio models work, the company's business model, the conversational Turing Test, voice agents, and more — Mati Stan

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks

Model ReleasesDGX agent

arXiv:2604.12102v1 Announce Type: new Abstract: We introduce compute-grounded reasoning (CGR), a design paradigm for spatial-aware research agents in which every answerable sub-problem is resolved by

Thought-Retriever: Don't Just Retrieve Raw Data, Retrieve Thoughts for Memory-Augmented Agentic Systems

Model ReleasesDGX agent

arXiv:2604.12231v1 Announce Type: new Abstract: Large language models (LLMs) have transformed AI research thanks to their powerful internal capabilities and knowledge. However, existing LLMs still fai

Towards Long-horizon Agentic Multimodal Search

Model ReleasesDGX agent

arXiv:2604.12890v1 Announce Type: cross Abstract: Multimodal deep search agents have shown great potential in solving complex tasks by iteratively collecting textual and visual evidence. However, mana

14 Apr 2026

A Dual-Positive Monotone Parameterization for Multi-Segment Bids and a Validity Assessment Framework for Reinforcement Learning Agent-based Simulation of Electricity Markets

SafetyDGX agent

arXiv:2604.10252v1 Announce Type: new Abstract: Reinforcement learning agent-based simulation (RL-ABS) has become an important tool for electricity market mechanism analysis and evaluation. In the mod

From Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to Python

Model ReleasesDGX agent

arXiv:2604.11518v1 Announce Type: cross Abstract: Cross-language migration of large software systems is a persistent engineering challenge, particularly when the source codebase evolves rapidly. We pr

Hubble: An LLM-Driven Agentic Framework for Safe and Automated Alpha Factor Discovery

SafetyDGX agent

arXiv:2604.09601v1 Announce Type: new Abstract: Discovering predictive alpha factors in quantitative finance remains a formidable challenge due to the vast combinatorial search space and inherently lo

I have a Macbook AIR M5 Base and I want to run an Agentic Coding program, similar to Claude Code or Codex. Besides the model, how do I do it? I've already tried with Ollama, VS Code, Opencode, and haven't been able to. (I'm not a developer, sorry)

Model ReleasesDGX agent

This Reddit thread addresses a common challenge for non-developers trying to run a local agentic coding assistant on a MacBook Air M5: while tools like Ollama, VS Code, and OpenCode are the right piec

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization

AgentsDGX agent

arXiv:2604.11042v1 Announce Type: new Abstract: Fine-tuning object detection (OD) models on combined datasets assumes annotation compatibility, yet datasets often encode conflicting spatial definition

Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards

AgentsDGX agent

arXiv:2604.09855v1 Announce Type: new Abstract: The recent advancement of Large Language Models (LLMs) has established their potential as autonomous interactive agents. However, they often struggle in

Love seeing this open-sourced. Had a great chat with @nicoalbanese10 some weeks ago where he hinted to something like this. Great reference …

Local AiDGX agent

Love seeing this open-sourced. Had a great chat with @nicoalbanese10 some weeks ago where he hinted to something like this. Great reference architecture for cloud coding agents. Open Agents gives you

MADQRL: Distributed Quantum Reinforcement Learning Framework for Multi-Agent Environments

SafetyDGX agent

arXiv:2604.11131v1 Announce Type: new Abstract: Reinforcement learning (RL) is one of the most practical ways to learn from real-life use-cases. Motivated from the cognitive methods used by humans mak

← Previous
1…110111112113114…300
Next →