AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,958 results
Agents

Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests

DGX agent

arXiv:2606.07379v1 Announce Type: cross Abstract: A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead of solving t

agentsarxiv-cs-ai
8 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup​ 🔹Drag in videos as cod…

DGX agent

Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup​ 🔹Drag in videos as coding context: reference-to-LUT, long-video-to-short, screen-rec

agentskimi-moonshot--x
8 Jun 2026
Agents

LLM Agent-Assisted Reverse Engineering with Quantitative Readability Metrics

DGX agent

arXiv:2606.06838v1 Announce Type: cross Abstract: Automatic decompilers produce functionally correct but often unreadable C code. This paper addresses one stage of the reverse engineering workflow: im

agentsarxiv-cs-ai
8 Jun 2026
Local Ai

Queen-Bee Agents: A BeeSpec-Centered Architecture for Governed Enterprise MCP Orchestration

DGX agent

arXiv:2606.06545v1 Announce Type: cross Abstract: Enterprise agent systems increasingly need to connect large language models to private tools, internal knowledge, and Model Context Protocol (MCP) int

local-aiarxiv-cs-ai
8 Jun 2026
Agents

SCALE: Scalable Cross-Attention Learning with Extrapolation for Agentic Workflow Scheduling

DGX agent

arXiv:2606.06820v1 Announce Type: cross Abstract: Agentic Large Language Model (LLM) systems decompose complex tasks into workflow Directed Acyclic Graphs (DAGs) whose primitives must be scheduled on

agentsarxiv-cs-ai
8 Jun 2026
Agents

🔗Try it now: https://www.kimi.com/products/kimi-work We're just getting started. More data sources, more tools, more agent capabilities are…

DGX agent

Kimi is launching a new product called Kimi Work, an AI work platform with expandable capabilities including additional data sources, tools, and agent functionalities. The announcement indicates the p

agentskimi-moonshot--x
8 Jun 2026
Agents

We found that more autonomy with autonomous agents like Computer tracks with higher quality and satisfaction.

DGX agent

Research indicates that autonomous agents with greater operational autonomy, particularly in computer-based tasks, demonstrate improved performance quality and user satisfaction outcomes. The study su

agentsperplexity--x
8 Jun 2026
Agents

We got into Y Combinator! Agnost AI (YC S26) is the infra for self-improving AI agents. We plug into conversational AI companies, find what'…

DGX agent

We got into Y Combinator! Agnost AI (YC S26) is the infra for self-improving AI agents. We plug into conversational AI companies, find what's broken, and ship the fix as a PR. You just merge. DM if th

agentsyohei-nakajima--x
8 Jun 2026
Agents

We published new research with Harvard on the shift from chat interfaces to autonomous agents like Computer. Over 3 months, findings show wo…

DGX agent

We published new research with Harvard on the shift from chat interfaces to autonomous agents like Computer. Over 3 months, findings show workers using Computer finish tasks in 87% less time at 94% lo

agentsperplexity--x
8 Jun 2026
Agents

Agree with everything except the Markdown part There's got to be a better agent-native format for representing unstructured docs Not convinc…

DGX agent

Agree with everything except the Markdown part There's got to be a better agent-native format for representing unstructured docs Not convinced it's markdown or html I need Google Docs but just for mar

agentsjerry-liu--x
7 Jun 2026
Agents

Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning

DGX agent

arXiv:2503.01734v3 Announce Type: replace-cross Abstract: Attacks on machine learning models have been extensively studied through stateless optimization. In this paper, we demonstrate how a reinforce

agentsarxiv-cs-ai
6 Jun 2026
Model Releases

Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents

DGX agent

arXiv:2606.06036v1 Announce Type: new Abstract: Despite recent progress, LLM agents still struggle with reasoning over long interaction histories. While current memory-augmented agents rely on a stati

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation

DGX agent

arXiv:2606.05241v1 Announce Type: cross Abstract: Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

The End of Software Engineering: How AI Agents Are Fundamentally Restructuring the Software Paradigm

DGX agent

arXiv:2606.05608v1 Announce Type: cross Abstract: For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into s

model-releasesarxiv-cs-ai
6 Jun 2026
Agents

This chart from Anthropic is useful, since Agent Teams and Workflows are both very new and very powerful (and token hungry). On the other ha…

DGX agent

This chart from Anthropic is useful, since Agent Teams and Workflows are both very new and very powerful (and token hungry). On the other hand, maybe it doesn't matter as a lot of the decisions about

agentsethan-mollick--x
6 Jun 2026
Safety

Towards Healthy Evolution: Exploring the Role and Mechanisms of Human-Agent Interaction in Self-Evolving Systems

DGX agent

arXiv:2606.06114v1 Announce Type: new Abstract: Self-evolving agents improve through continual self-play and self-generated learning signals, but autonomous evolution can also cause capability degrada

safetyarxiv-cs-ai
6 Jun 2026
Model Releases

Arena AI Agentic User Benchmark Ranking

DGX agent

Arena AI's agentic benchmark ranks AI models on how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability. The leaderboard

model-releasesr-chatgpt
5 Jun 2026
Agents

EGTR-Review: Efficient Evidence-Grounded Scientific Peer Review Generation via Multi-Agent Teacher Distillation

DGX agent

arXiv:2606.06025v1 Announce Type: new Abstract: Scientific peer review generation has attracted increasing attention for reducing reviewing burdens and providing timely feedback. However, existing Lar

agentsarxiv-cs-cl
5 Jun 2026
Agents

Harnessing Generalist Agents for Contextualized Time Series

DGX agent

arXiv:2606.05404v1 Announce Type: cross Abstract: Time series are often embedded in rich contexts that are essential for holistic modeling. Moreover, real-world practitioners often require end-to-end

agentsarxiv-cs-cl
5 Jun 2026
Model Releases

Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

DGX agent

arXiv:2606.06388v1 Announce Type: cross Abstract: Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly pos

model-releasesarxiv-cs-cl
5 Jun 2026
Agents

In an internal message, Satya Nadella rebuked an internal memo that said Microsoft needs to 'make people addicted' to its new AI agent product called Scout (Aaron Holmes/The Information)

DGX agent

Aaron Holmes / The Information: In an internal message, Satya Nadella rebuked an internal memo that said Microsoft needs to “make people addicted” to its new AI agent product called Scout — Microsoft

agentstechmeme
5 Jun 2026
Agents

Need advice on building/training an AI Agent for fully automated blog generation

DGX agent

A discussion on building an AI-powered article generator using CrewAI and Ollama with specialized AI agents for research and writing to generate comprehensive articles on any topic. The solution runs

agentsr-ollama
5 Jun 2026
Agents

Shopify on Replit + the new SEO Agent https://x.com/i/broadcasts/1kJzDDopENZKv

DGX agent

This post likely covers a live broadcast or announcement discussing the integration of Shopify with Replit, along with information about a newly released SEO Agent tool. The content probably demonstra

agentsreplit--x
5 Jun 2026
Model Releases

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework

DGX agent

arXiv:2606.05570v1 Announce Type: new Abstract: Repository-level coding benchmarks face a trade-off between task difficulty and evaluation reliability: tasks that challenge frontier models often invol

model-releasesarxiv-cs-cl
5 Jun 2026
Agents

At @WitanLabs we're building the headless Office stack for AI agents. Here's what we worked on this week — and what we're building next ↓

DGX agent

Witan Labs is developing a headless Office stack designed specifically for AI agents, combining productivity tools and capabilities without a traditional user interface. This post appears to be a week

agentsharrison-chase--x
4 Jun 2026
Safety

Be Fair! Can Machine Learning Engineering Agents Adhere to Fairness Constraints?

DGX agent

arXiv:2606.04971v1 Announce Type: new Abstract: Machine learning engineering (MLE) agents promise to automate end-to-end ML pipeline development from raw data and natural language instructions, potent

safetyarxiv-cs-lg
4 Jun 2026
Agents

Cloudflare CEO Matthew Prince says agentic traffic is 'growing so fast that bots have now passed human traffic online for the first time' (Mark Tyson/Tom's Hardware)

DGX agent

Mark Tyson / Tom's Hardware: Cloudflare CEO Matthew Prince says agentic traffic is “growing so fast that bots have now passed human traffic online for the first time” — Bot (automated) vs. human HTTP

agentstechmeme
4 Jun 2026
Agents

Cursor can now show your agent's context usage as an interactive report in a canvas. The context explorer breaks down where tokens go across…

DGX agent

Cursor can now show your agent's context usage as an interactive report in a canvas. The context explorer breaks down where tokens go across the system prompt, tool definitions, rules, skills, and mor

agentscursor--x
4 Jun 2026
Agents

DAR: Deontic Reasoning with Agentic Harnesses

DGX agent

arXiv:2606.05009v1 Announce Type: cross Abstract: Deontic reasoning is the task of answering questions by applying explicit rules and policies to case-specific facts, for example computing tax liabili

agentsarxiv-cs-ai
4 Jun 2026
Safety

Fog of Love: Engineering Virtuous Agent Behavior with Affinity-based Reinforcement Learning in a Game Environment

DGX agent

arXiv:2606.04750v1 Announce Type: new Abstract: Instilling virtuous behavior in artificial intelligence has seen increasing interest. One of the techniques proposed is known as affinity-based reinforc

safetyarxiv-cs-ai
4 Jun 2026
Agents

Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning

DGX agent

arXiv:2507.21892v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) mitigates hallucination in LLMs by incorporating external knowledge, but relies on chunk-based retrieval that l

agentsarxiv-cs-cl
4 Jun 2026
Safety

Learning While Acting: A Skill-Enhanced Test-Time Co-Evolution Framework for Online Lifelong Learning Agents

DGX agent

arXiv:2606.04815v1 Announce Type: cross Abstract: Lifelong learning is essential for Large Language Model (LLM) agents operating in dynamic, interactive environments. However, existing lifelong learni

safetyarxiv-cs-ai
4 Jun 2026
Agents

On the latest episode of Max Agency, @hwchase17 sat down with @nlarusstone, Head of AI at @benchling for a conversation on building agents f…

DGX agent

On the latest episode of Max Agency, @hwchase17 sat down with @nlarusstone, Head of AI at @benchling for a conversation on building agents for scientific work. ⏯️ YouTube: https://www.youtube.com/watc

agentsharrison-chase--x
4 Jun 2026
Agents

Optimizing the Cost-Quality Tradeoff of Agentic Theorem Provers in Lean

DGX agent

arXiv:2606.04883v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in workflows for generating formal proofs in Lean. These workflows often decompose problems into smal

agentsarxiv-cs-cl
4 Jun 2026
Model Releases

Today I'm launching a new project called SynthTraces 🔥 It is a minimal codebase to generate synthetic coding agent session traces using Pi …

DGX agent

Today I'm launching a new project called SynthTraces 🔥 It is a minimal codebase to generate synthetic coding agent session traces using Pi (from @badlogicgames) I wanted a large number of coding-agent

model-releasesclem-delangue--x
4 Jun 2026
Agents

We partnered with Shopify so you can go from idea to live store in minutes Just tell Replit Agent what you want to sell. It will: - Build a …

DGX agent

We partnered with Shopify so you can go from idea to live store in minutes Just tell Replit Agent what you want to sell. It will: - Build a custom storefront - Create your Shopify store - Help you add

agentsreplit--x
4 Jun 2026
Agents

Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning

DGX agent

arXiv:2606.03965v1 Announce Type: cross Abstract: Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little

agentsarxiv-cs-ai
3 Jun 2026
Agents

DELTAMEM: Incremental Experience Memory for LLM Agents via Residual Trees

DGX agent

arXiv:2606.03083v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents increasingly rely on memory to learn from experiences over continual interactions. However, storing experiences

agentsarxiv-cs-ai
3 Jun 2026
Local Ai

InfoMem: Training Long-Context Memory Agents with Answer-Conditioned Information Gain

DGX agent

arXiv:2606.03329v1 Announce Type: new Abstract: Long-context tasks require LLMs to identify and preserve answer-relevant information from large contexts. Chunk-wise memory agents address this issue by

local-aiarxiv-cs-ai
3 Jun 2026
Agents

@jgreze will speak on this at https://ai.engineer/wf gathering all the top agent labs. lfg

DGX agent

Swyx announces that @jgreze will speak at the ai.engineer/wf gathering, which is positioned as an event bringing together leading agent research labs. The post expresses enthusiasm for the upcoming sp

agentsswyx--x
3 Jun 2026
Agents

Just pushed an update to help remote connecting with the Hermes Agent GUI over tailscale to function! Please update if you had any issues!

DGX agent

Nous Research released an update improving remote connectivity functionality for the Hermes Agent GUI when used over Tailscale, addressing previous compatibility issues. Users experiencing connection

agentsnous-research--x
3 Jun 2026
Model Releases

MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents

DGX agent

arXiv:2606.03203v1 Announce Type: new Abstract: Computer-use agents could automate repetitive screen-based clinical work, but their reliability in medical graphical user interfaces remains largely unv

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

PhotoCraft: Agentic Reasoning with Hierarchical Self-Evolving Memory for Deep Image Search

DGX agent

arXiv:2606.03099v1 Announce Type: cross Abstract: Deep Image Search requires multi-step reasoning over rich contextual cues, such as time, location, and event relations. However, most existing LLM-bas

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

The Impact of Configuring Agentic AI Coding Tools on Build-vs-Buy Decisions: A Study Protocol

DGX agent

arXiv:2606.03907v1 Announce Type: cross Abstract: Agentic AI coding tools write code with increasing autonomy and in doing so decide when to import a library and when to implement functionality from s

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size

DGX agent

arXiv:2606.02646v1 Announce Type: cross Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence. We derive a two-paramete

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

DGX agent

arXiv:2606.03137v1 Announce Type: new Abstract: LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existi

safetyarxiv-cs-ai
3 Jun 2026
Local Ai

Toward a Modular Architecture for Embedded AI Agent Systems at the Edge

DGX agent

arXiv:2606.02862v1 Announce Type: new Abstract: The rise of Large Language Models (LLMs) has enabled agentic AI capable of complex reasoning and tool use; however, deploying such autonomy in pervasive

local-aiarxiv-cs-ai
3 Jun 2026
Agents

Welp, that happened faster than I predicted. Thought it would be end of 2027, then early 2027, but agentic traffic growing so fast that bots…

DGX agent

Welp, that happened faster than I predicted. Thought it would be end of 2027, then early 2027, but agentic traffic growing so fast that bots have now passed human traffic online for the first time in

agentselon-musk--x
3 Jun 2026
← Previous
1…102103104105106…375
Next →