AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,959 results
2 Jun 2026

MemoNoveltyAgent: A Historical Research Memory-Aware Agent Workflow for Paper Novelty Assessment

Model ReleasesDGX agent

arXiv:2603.20884v2 Announce Type: replace Abstract: To alleviate the heavy burden of paper screening, researchers increasingly rely on existing AI agents, such as AI reviewers or DeepResearch, for pap

Microsoft launches Rayfin to let developers and agents build app backends on Fabric

Model ReleasesDGX agent

Microsoft Corp. today introduced Rayfin, an open-source software development kit and command-line interface that lets developers and coding agents define an entire application backend in code and depl

Microsoft unveils Microsoft Execution Containers for Windows, an OS-level sandbox for AI agents, with OpenAI, Nvidia, Manus, and Nous Research as partners (Michael Nuñez/VentureBeat)

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

Michael Nuñez / VentureBeat: Microsoft unveils Microsoft Execution Containers for Windows, an OS-level sandbox for AI agents, with OpenAI, Nvidia, Manus, and Nous Research as partners — For the past t

Microsoft's Project Solara is an Android OS designed for agents instead of apps

IndustryDGX agent

Microsoft has developed Project Solara, a platform for devices that run AI agents instead of apps, based on Android instead of Windows. The platform is Microsoft's bet that AI will open up entirely ne

MOSAIC: Modular Orchestration for Structured Agentic Intelligence and Composition

SafetyDGX agent

arXiv:2606.00708v1 Announce Type: new Abstract: Automated data science is a structured model-selection problem. A solution must choose data transformations, feature representations, architecture, trai

NestRL: A Nested Training Regime for Mutual Adaptation in Human-AI Teaming

AgentsDGX agent

arXiv:2602.17737v2 Announce Type: replace-cross Abstract: Mutual adaptation is a central challenge in human-AI teaming, as humans naturally adjust their strategies in response to an AI agent's behavio

RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates

SafetyDGX agent

arXiv:2506.11083v3 Announce Type: replace Abstract: We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate

Rehumanizing global health care with agentic AI

AgentsDGX agent

The global health care sector is under increasing strain. Decades of chronic underinvestment and constraints in recruitment have coincided with a surge in demand for services for aging populations. Ga

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents

Model ReleasesDGX agent

arXiv:2606.01552v1 Announce Type: new Abstract: Role-playing agents(RPAs) are widely used to steer large language models(LLMs) toward role-consistent behavior, yet existing benchmarks mainly evaluate

RubricMiddleware helps your agent verify task completion with a grader subagent This is similar to /goal in claude code or codex, but for de…

Model ReleasesDGX agent

RubricMiddleware is a LangChain feature that enables agents to verify task completion by delegating grading to a specialized subagent using predefined rubrics. This approach parallels the goal-verific

Self-Healing Agentic Orchestrators for Reliable Tool-Augmented Large Language Model Systems

Model ReleasesDGX agent

arXiv:2606.01416v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents rely on orchestration layers that coordinate planning, retrieval, tool invocation, validation, memory,

SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories

Model ReleasesDGX agent

arXiv:2606.01311v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on reusable external skills to solve long-horizon interactive tasks. Existing training-free skill

Strategizing at Speed: A Learned Model Predictive Game for Multi-Agent Drone Racing

Model ReleasesDGX agent

arXiv:2602.06925v2 Announce Type: replace Abstract: Autonomous drone racing pushes the boundaries of high-speed motion planning and multi-agent strategic decision-making. Success in this domain requir

Verification is the hidden bottleneck for knowledge work agents, especially in legal AI — complex, long-horizon work is graded by rubrics wi…

TutorialsDGX agent

Verification is the hidden bottleneck for knowledge work agents, especially in legal AI — complex, long-horizon work is graded by rubrics with dozens of strict criteria. In new research with @LangChai

1 Jun 2026

COMPASS: Cognitive MCTS-Guided Process Alignment for Safe Search Agents

SafetyDGX agent

arXiv:2605.30838v1 Announce Type: new Abstract: LLM-powered search agents enable multi-step reasoning and tool use. However, these capabilities introduce retrieval-induced safety degradation, as harmf

Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution

Model ReleasesDGX agent

arXiv:2605.30802v1 Announce Type: cross Abstract: Prediction markets aggregate collective intelligence to forecast uncertain events, but their utility depends on reliable outcome resolution. Existing

ElasticMem: Latent Memory as a Learnable Resource for LLM Agents

Model ReleasesDGX agent

arXiv:2605.30690v1 Announce Type: new Abstract: Long-term memory is essential for LLM agents to reason coherently across extended interactions, personalize responses, and reuse past experience. Howeve

Exploring Autonomous Agentic Data Engineering for Model Specialization

Model ReleasesDGX agent

arXiv:2605.30407v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without hig

From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents

Model ReleasesDGX agent

arXiv:2603.18382v2 Announce Type: replace Abstract: Anonymization is often assumed to protect privacy once explicit identifiers are removed, because re-identification has historically required special

Gemini’s new AI agent is about as good as Google’s demo

Model ReleasesDGX agent

Google's new '24/7' AI agent, Gemini Spark, can be shockingly good at doing things on your behalf. But I'm not sure it's worth the financial cost and potential privacy tradeoffs. The company gave me a

HADT: A Heterogeneous Multi-Agent Differential Transformer for Autonomous Earth Observation Satellite Cluster

AgentsDGX agent

arXiv:2605.31023v1 Announce Type: new Abstract: This work addresses the problem of autonomous resource management in heterogeneous satellite cluster conducting Earth Observation (EO) missions includin

just a small zoom out on the vibe shift: in Feb 2025 @soumithchintala was talking about his dream of personal, local, private agents, most p…

Model ReleasesDGX agent

just a small zoom out on the vibe shift: in Feb 2025 @soumithchintala was talking about his dream of personal, local, private agents, most people didn't believe him. it's June 2026 and @pewdiepie has

KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware

Model ReleasesDGX agent

arXiv:2603.08721v2 Announce Type: replace-cross Abstract: New AI accelerators with novel instruction set architectures (ISAs) often require developers to manually craft low-level kernels, a time-consu

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks

AgentsDGX agent

arXiv:2603.22744v2 Announce Type: replace Abstract: Large language models excel on objectively verifiable tasks such as math and programming, where evaluation reduces to unit tests or a single correct

NVIDIA Levels Up Local AI Agents Across RTX PCs and DGX Spark

Local AiDGX agent

Personal agents are exploding in popularity, with open source projects like OpenClaw and Hermes seeing rapid adoption by AI developer communities on GitHub. Built to adapt to individual preferences an

Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents

SafetyDGX agent

arXiv:2605.30723v1 Announce Type: new Abstract: LLM agents increasingly retrieve externally curated skills-procedural instructions retrieved at decision time-to improve performance on long-horizon int

31 May 2026

// The Efficiency Frontier // Cool paper on context management. As agents reuse the same documents and histories across many turns, the chea…

Model ReleasesDGX agent

// The Efficiency Frontier // Cool paper on context management. As agents reuse the same documents and histories across many turns, the cheapest context strategy is not fixed. This work describes a pr

29 May 2026

Adobe’s conversational AI agent is a mediocre design intern

AgentsDGX agent

AI image tools rarely make me feel like I'm part of the creative process. They are, after all, mostly designed so that people with no design experience can type in a few words and get back a usable re

AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.29643v1 Announce Type: new Abstract: Cross-Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence d

Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents

SafetyDGX agent

arXiv:2605.29910v1 Announce Type: cross Abstract: Consensus protocols form the backbone of distributed systems and blockchains, where implementation bugs can cause data corruption and financial losses

Auto-review mode is now available in Cursor. It allows agents to run tool calls with fewer approval prompts and safer execution.

ToolsDGX agent

Cursor has introduced an auto-review mode feature that enables AI agents to execute tool calls with reduced approval requirements while maintaining safer execution practices. This feature streamlines

Do Proactive Agents Really Need an LLM to Decide When to Wake and What to Anchor?

Local AiDGX agent

arXiv:2605.30152v1 Announce Type: cross Abstract: Proactive agents read user activity as text and call an LLM on every event to decide whether to act. But user activity is not natively text: it is a s

DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents

Model ReleasesDGX agent

arXiv:2605.29256v1 Announce Type: cross Abstract: Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality

Entity-Collision: A Stratified Protocol for Attributing Retrieval Lift in Agent Memory

Model ReleasesDGX agent

arXiv:2605.29630v1 Announce Type: cross Abstract: End-to-end agent-memory benchmarks report a single hit@k per retriever, confounding lexical leakage (uncontrolled query/gold/distractor entity overlap

Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension

Model ReleasesDGX agent

arXiv:2605.29874v1 Announce Type: cross Abstract: Do next-generation LLM agents inherit the cooperative biases documented in their predecessors, or does scale and provider diversity reshape equilibriu

Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes

Model ReleasesDGX agent

arXiv:2605.28965v1 Announce Type: new Abstract: Linking free-text phenotype descriptions to ontology terms, typically referred to as phenotype annotation, is essential for the cross-study integration

Having Grok Build sub-agents to iterate several ideas for me on data loading, batching, inference, and writing results to files for dense da…

TutorialsDGX agent

Having Grok Build sub-agents to iterate several ideas for me on data loading, batching, inference, and writing results to files for dense datasets before I went to sleep. It gave me a nice summary of

Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents

Local AiDGX agent

arXiv:2605.30159v1 Announce Type: new Abstract: Memory-augmented LLM agents tackle complex long-horizon tasks by recursively summarizing interaction trajectories into compact memory. However, existing

PhoneWorld: Scaling Phone-Use Agent Environments

Model ReleasesDGX agent

arXiv:2605.29486v1 Announce Type: cross Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Ex

PRO-CUA: Process-Reward Optimization for Computer Use Agents

SafetyDGX agent

arXiv:2605.29119v1 Announce Type: new Abstract: Computer use agents (CUAs) have shown strong potential for automating complex digital workflows, yet their training remains constrained by costly live e

Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents

SafetyDGX agent

arXiv:2605.28850v1 Announce Type: new Abstract: We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments. Using TradeArena, an

SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search

Model ReleasesDGX agent

arXiv:2605.29796v1 Announce Type: new Abstract: Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness, these syste

SkillsInjector: Dynamic Skill Context Construction for LLM Agents

Model ReleasesDGX agent

arXiv:2605.29794v1 Announce Type: new Abstract: LLM agents now draw on growing skill libraries to handle complex tasks. However, injecting more skills does not always improve task completion and can e

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2602.15382v2 Announce Type: replace Abstract: Multi-Agent Systems (MAS) powered by Large Language Models have unlocked advanced collaborative reasoning, yet they remain bottlenecked by discrete

Unifying Temporal and Structural Credit Assignment in LLM-Based Multi-Agent Prompt Optimization

Local AiDGX agent

arXiv:2605.30227v1 Announce Type: cross Abstract: While Multi-Agent Systems (MAS) empower Large Language Models to tackle complex reasoning tasks through collaborative interaction, optimizing their dy

28 May 2026

Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning

Model ReleasesDGX agent

arXiv:2605.28192v1 Announce Type: new Abstract: Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across b

AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications

Model ReleasesDGX agent

arXiv:2605.27761v1 Announce Type: new Abstract: The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments

Banger paper from Harvard. AutoScientists drops the central planner entirely. Agents interpret shared experimental data, self-organize aroun…

TutorialsDGX agent

Banger paper from Harvard. AutoScientists drops the central planner entirely. Agents interpret shared experimental data, self-organize around promising directions, evaluate proposals before resource a

Do Agents Think Deeper? A Mechanistic Investigation of Layer-Wise Dynamics in Sequential Planning

Model ReleasesDGX agent

arXiv:2605.27935v1 Announce Type: new Abstract: Recent mechanistic studies suggest that large language models (LLMs) may utilize their depth inefficiently in standard single-turn tasks. Whether this s

Dr-CiK: A Testbed for Foresight-Driven Agents

Model ReleasesDGX agent

arXiv:2605.27904v1 Announce Type: new Abstract: Time series forecasting in real-world settings often depends not only on historical observations, but also on external context that must be actively dis

DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM-based Scheduling Agents

Model ReleasesDGX agent

arXiv:2605.27566v1 Announce Type: new Abstract: Progress in neural combinatorial optimization for Dynamic Flexible Job Shop Scheduling Problem (DFJSP) is currently hindered by a methodological tension

Fine-Tuning Vision-Language Models for Understanding Current Damage and Scoring Priority with Quality Guard Agent

AgentsDGX agent

arXiv:2605.27452v1 Announce Type: new Abstract: Bridge inspection in Japan requires mandatory visual assessments every five years, yet qualitative damage ratings (levels a-e) assigned by different eng

London-based Geordie AI, which builds a security and governance platform for AI agents, raised a 30M Series A led by Balderton at an estimated 180M valuation (Jeremy Kahn/Fortune)

IndustryDGX agent

Jeremy Kahn / Fortune: London-based Geordie AI, which builds a security and governance platform for AI agents, raised a 30M Series A led by Balderton at an estimated 180M valuation — Geordie AI, a Lon

MaskClaw: Edge-Side Personalized Privacy Arbitration for GUI Agents with Behavior-Driven Skill Evolution

Model ReleasesDGX agent

arXiv:2605.28646v1 Announce Type: cross Abstract: GUI agents rely on screenshots to infer intent and operate across applications, but these screenshots often contain private messages, medical records,

🚨 This is EXACTLY WHY ICE agents are FORCED to wear masks ICE Newark rioter: “I HAVE YOUR FACE, MOTHERF***ER” “Your WHOLE F***ING FAMILY is…

IndustryDGX agent

🚨 This is EXACTLY WHY ICE agents are FORCED to wear masks ICE Newark rioter: “I HAVE YOUR FACE, MOTHERF***ER” “Your WHOLE F***ING FAMILY is DEAD!” “Your KIDS. Your WIFE. ALL DEAD!” This is the type of

Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution

Model ReleasesDGX agent

arXiv:2605.28000v1 Announce Type: cross Abstract: Large language model agents are increasingly expected to perform operational work: calling APIs, manipulating files, assembling workflows, and acting

27 May 2026

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation

Model ReleasesDGX agent

arXiv:2605.26321v1 Announce Type: new Abstract: AI agents are beginning to complete valuable, long-horizon business operations tasks, but training and evaluation environments for enterprise work still

APEX-Searcher: Refining Credit Assignment with Subgoaling for Agentic Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2603.13853v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) connects large language models (LLMs) to external knowledge, but single-round retrieval is often insuffic

AWS launches Agentic Shopping Assistant to help retailers build AI tools

Model ReleasesDGX agent

Amazon Web Services Inc. today introduced a new offering designed to help retailers integrate artificial intelligence features into their online stores. AWS Agentic Shopping Assistant, or ASA, combine

Doppel launches agentic email security to disrupt phishing campaigns at the source

Model ReleasesDGX agent

Social engineering defense startup Doppel Inc. today launched Doppel Email Security, an agentic artificial intelligence layer that traces phishing messages back to attacker infrastructure and orchestr

← Previous
1…117118119120121…300
Next →