AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,973 results
Safety

JACoP: Joint Alignment for Compliant Multi-Agent Prediction

DGX agent

arXiv:2605.11385v1 Announce Type: new Abstract: Stochastic Human Trajectory Prediction (HTP) using generative modeling has emerged as a significant area of research. Although state-of-the-art models e

safetyarxiv-cs-cv
13 May 2026
Local Ai
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Poolside is hosting a 2-day model research hackathon in London. Join us to push an open-weight agent model as far as you can. RL and fine-tu…

DGX agent

Poolside is hosting a 2-day model research hackathon in London. Join us to push an open-weight agent model as far as you can. RL and fine-tune Laguna XS.2, our latest-generation model, on Prime Intell

local-aiclem-delangue--x
13 May 2026
Research

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction

DGX agent

arXiv:2605.11212v1 Announce Type: new Abstract: Computer-use agents~(CUAs) rely on visual observations of graphical user interfaces, where each screenshot is encoded into a large number of visual toke

researcharxiv-cs-cl
13 May 2026
Model Releases

Starting June 15, paid Claude plans can claim a dedicated monthly credit for programmatic usage. The credit covers usage of: - Claude Agent …

DGX agent

Starting June 15, paid Claude plans can claim a dedicated monthly credit for programmatic usage. The credit covers usage of: - Claude Agent SDK - claude -p - Claude Code GitHub Actions - Third-party a

model-releasesboris-cherny--x
13 May 2026
Model Releases

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a…

DGX agent

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a human would never try on another human. In one run, GPT-5.1

model-releasesemad-mostaque--x
13 May 2026
Model Releases

Agentic MIP Research: Accelerated Constraint Handler Generation

DGX agent

arXiv:2605.09186v1 Announce Type: new Abstract: Mixed-integer programming (MIP) research is both mathematically sophisticated and engineering-intensive: testing an algorithmic hypothesis within a bran

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling

DGX agent

arXiv:2604.08178v2 Announce Type: replace Abstract: In classical Reinforcement Learning from Human Feedback (RLHF), Reward Models (RMs) serve as the fundamental signal provider for model alignment. As

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning

DGX agent

arXiv:2605.08271v1 Announce Type: cross Abstract: Understanding ultra-long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For

model-releasesarxiv-cs-ai
12 May 2026
Safety

CARL: Criticality-Aware Agentic Reinforcement Learning

DGX agent

arXiv:2512.04949v3 Announce Type: replace-cross Abstract: Agents capable of accomplishing complex tasks through multiple interactions with the environment have emerged as a popular research direction.

safetyarxiv-cs-ai
12 May 2026
Agents

CoCoDA: Co-evolving Compositional DAG for Tool-Augmented Agents

DGX agent

arXiv:2605.08399v1 Announce Type: new Abstract: Tool-augmented language models can extend small language models with external executable skills, but scaling the tool library creates a coupled challeng

agentsarxiv-cs-ai
12 May 2026
Model Releases

CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents

DGX agent

arXiv:2605.09675v1 Announce Type: new Abstract: Clinical reasoning agents based on large language models (LLMs) aim to automate tasks such as intensive care unit (ICU) monitoring and patient state tra

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Context-Augmented Code Generation: How Product Context Improves AI Coding Agent Decision Compliance by 49%

DGX agent

arXiv:2605.08112v1 Announce Type: cross Abstract: AI coding agents powered by large language models can read codebases and produce functional code, but they routinely violate team-specific product dec

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents

DGX agent

arXiv:2605.08747v1 Announce Type: new Abstract: Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call te

model-releasesarxiv-cs-ai
12 May 2026
Safety

EGL-SCA: Structural Credit Assignment for Co-Evolving Instructions and Tools in Graph Reasoning Agents

DGX agent

arXiv:2605.10366v1 Announce Type: new Abstract: Graph reasoning agents operating from natural-language inputs must solve a coupled problem: they must reconstruct a structured graph instance from text,

safetyarxiv-cs-ai
12 May 2026
Model Releases

Idira launches as Palo Alto Networks extends CyberArk tech to machine and agentic identities

DGX agent

Palo Alto Networks Inc. today launched Idira, a new identity security platform designed to manage human, machine and artificial intelligence agent identities across the enterprise under a single privi

model-releasessiliconangle
12 May 2026
Model Releases

Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables

DGX agent

arXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid

model-releasesarxiv-cs-cl
12 May 2026
Safety

MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL

DGX agent

arXiv:2511.01008v2 Announce Type: replace Abstract: Large Language Models (LLMs) often struggle with the precise logic and schema alignment required for complex Text-to-SQL tasks. While current method

safetyarxiv-cs-cl
12 May 2026
Model Releases

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

DGX agent

arXiv:2605.09131v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Meet physics-intern🧑‍🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA o…

DGX agent

Meet physics-intern🧑‍🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA on one of the hardest benchmarks for LLMs. Theoretical physics

model-releasesclem-delangue--x
12 May 2026
Model Releases

MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents

DGX agent

arXiv:2605.09530v1 Announce Type: cross Abstract: As LLM-powered agents are increasingly deployed in edge-cloud environments, personalized memory has become a key enabler of long-term adaptation and u

model-releasesarxiv-cs-cl
12 May 2026
Research

MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents

DGX agent

arXiv:2603.13131v3 Announce Type: replace Abstract: Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A centra

researcharxiv-cs-ai
12 May 2026
Local Ai

​[PoC] Building a Local Multi-Agent AI Dev Studio alpha version (Architect/Senior/Junior) on a 10-year-old Haswell & GTX 1050 Ti (No APIs, Full AirLLM + Ollama)

DGX agent

This post describes a proof-of-concept implementation of a multi-agent AI development studio with architect, senior, and junior role personas, built entirely locally using AirLLM and Ollama without re

local-air-ollama
12 May 2026
Model Releases

Priority-Driven Control and Communication in Decentralized Multi-Agent Systems via Reinforcement Learning

DGX agent

arXiv:2605.10482v1 Announce Type: cross Abstract: Event-triggered control provides a mechanism for avoiding excessive use of constrained communication bandwidth in networked multi-agent systems. Howev

model-releasesarxiv-cs-lg
12 May 2026
Safety

SceneFactory: GPU-Accelerated Multi-Agent Driving Simulation with Physics-Based Vehicle Dynamics

DGX agent

arXiv:2605.08528v1 Announce Type: cross Abstract: Autonomous-driving simulators typically trade physical fidelity for scalable parallelism. Physics-based platforms such as CARLA and MetaDrive provide

safetyarxiv-cs-ro
12 May 2026
Research

The Invisible Handshake: Persistent Overpricing by Adaptive Market Agents

DGX agent

arXiv:2510.15995v3 Announce Type: replace-cross Abstract: We study overpricing in a repeated game between two representative agents: a market maker, who controls market liquidity, and a market taker,

researcharxiv-cs-lg
12 May 2026
Agents

TimeClaw: A Time-Series AI Agent with Exploratory Execution Learning

DGX agent

arXiv:2605.10038v1 Announce Type: new Abstract: Time series analysis underpins forecasting, monitoring, and decision making in domains such as finance and weather, where solving a task often requires

agentsarxiv-cs-ai
12 May 2026
Safety

TodyComm: Task-Oriented Dynamic Communication for Multi-Round LLM-based Multi-Agent System

DGX agent

arXiv:2602.03688v2 Announce Type: replace Abstract: Multi-round LLM-based multi-agent systems rely on effective communication structures to support collaboration across rounds. However, most existing

safetyarxiv-cs-ai
12 May 2026
Model Releases

Unpredictability dissociates from structured control in language agents

DGX agent

arXiv:2605.09692v1 Announce Type: new Abstract: Unpredictable behavior is often taken as evidence of control, yet stochastic dispersion and structured action control need not coincide. This paper test

model-releasesarxiv-cs-ai
12 May 2026
Local Ai

Verifiable Process Rewards for Agentic Reasoning

DGX agent

arXiv:2605.10325v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has improved the reasoning abilities of large language models (LLMs), but most existing approaches

local-aiarxiv-cs-ai
12 May 2026
Model Releases

+1 to this. I was recently on a cross-continental flight without wifi, so I brought up Qwen3.6 & Gemma 4 (via @ollama) in Deep Agents on my …

DGX agent

+1 to this. I was recently on a cross-continental flight without wifi, so I brought up Qwen3.6 & Gemma 4 (via @ollama) in Deep Agents on my laptop. admittedly, they fell over on some more involved/com

model-releasesharrison-chase--x
11 May 2026
Model Releases

3 weeks since ml-intern launched and we just hit 1M messages exchanged. that's 3.3 agent-years of ML research in 21 days. 2 months worth of …

DGX agent

3 weeks since ml-intern launched and we just hit 1M messages exchanged. that's 3.3 agent-years of ML research in 21 days. 2 months worth of research every day. 17,383 training jobs total. talk about A

model-releasesclem-delangue--x
11 May 2026
Model Releases

Agent view is the best Claude Code native way to manage multiple sessions, kind of like tmux built for CC. We spent a lot of time getting th…

DGX agent

Agent view is the best Claude Code native way to manage multiple sessions, kind of like tmux built for CC. We spent a lot of time getting the details right, I hope you enjoy it. New in Claude Code: ag

model-releasesthariq--x
11 May 2026
Model Releases

AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents

DGX agent

arXiv:2605.06607v2 Announce Type: replace-cross Abstract: Recent LLM-based agents have closed substantial portions of the scientific discovery loop in software-only machine-learning research, in chemi

model-releasesarxiv-cs-ai
11 May 2026
Safety

Excellent explainer video by @FryRsquared on the risks of AI agents. She also raises a crucial point: we shouldn’t make the mistake of think…

DGX agent

Excellent explainer video by @FryRsquared on the risks of AI agents. She also raises a crucial point: we shouldn’t make the mistake of thinking current limitations will necessarily persist. As we’ve s

safetyyoshua-bengio--x
11 May 2026
Agents

FlightSense: An End-to-End MLOps Platform for Real-Time Flight Delay Prediction via Rotation-Chain Propagation Features and Agentic Conversational AI

DGX agent

arXiv:2605.07364v1 Announce Type: new Abstract: Flight delays impose cascading operational and financial burdens across the aviation network, costing the U.S. economy billions of dollars annually by d

agentsarxiv-cs-lg
11 May 2026
Model Releases

GTIG AI Threat Tracker: Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access

DGX agent

Executive Summary Since our February 2026 report on AI-related threat activity, Google Threat Intelligence Group (GTIG) has continued to track a maturing transition from nascent AI-enabled operations

model-releasesgoogle-cloud-ai
11 May 2026
Local Ai

HMACE: Heterogeneous Multi-Agent Collaborative Evolution for Combinatorial Optimization

DGX agent

arXiv:2605.07214v1 Announce Type: new Abstract: Large Language Models have recently emerged as a promising paradigm for automated heuristic design for NP-hard combinatorial optimization problems. Desp

local-aiarxiv-cs-ai
11 May 2026
Model Releases

Learning and Reusing Policy Decompositions for Hierarchical Generalized Planning with LLM Agents

DGX agent

arXiv:2605.06957v1 Announce Type: new Abstract: We present a dynamic policy-learning approach that combines generalized planning and hierarchical task decomposition for LLM-based agents. Our method, H

model-releasesarxiv-cs-ai
11 May 2026
Tutorials

// LLMs Improving LLMs // Interesting progress the past of couple of weeks around self-improving AI agents. If autoresearch was interesting,…

DGX agent

// LLMs Improving LLMs // Interesting progress the past of couple of weeks around self-improving AI agents. If autoresearch was interesting, you will like this read. (bookmark it) We've been hand-tuni

tutorialsdair-ai--x
11 May 2026
Local Ai

Reachy Mini ready to go! Audio on 🔉 Cleary I'll connect it to Local AI services and to my Hermes Agent really soon 💪

DGX agent

Clem Delangue announced the Reachy Mini robot is operational with audio capabilities enabled, with plans to integrate it with local AI services and a Hermes Agent in the near future. This indicates pr

local-aiclem-delangue--x
11 May 2026
Local Ai

Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair

DGX agent

arXiv:2605.07276v1 Announce Type: new Abstract: Code-agent RL often receives weak feedback: rollout-time signals are reliable and executable, but capture only necessary or surface conditions for task

local-aiarxiv-cs-ai
11 May 2026
Safety

Skill1: Unified Evolution of Skill-Augmented Agents via Reinforcement Learning

DGX agent

arXiv:2605.06130v2 Announce Type: replace Abstract: A persistent skill library allows language model agents to reuse successful strategies across tasks. Maintaining such a library requires three coupl

safetyarxiv-cs-ai
11 May 2026
Research

SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests

DGX agent

Using SocialReasoning Bench, we observed a stable pattern across models—agents execute competently, but fail to consistently improve the user’s position, even with explicit instructions to optimize fo

researchmicrosoft-research
11 May 2026
Model Releases

The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

DGX agent

arXiv:2605.08060v1 Announce Type: cross Abstract: Context window expansion is often treated as a straightforward capability upgrade for LLMs, but we find it systematically fails in multi-agent social

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning

DGX agent

arXiv:2605.06772v1 Announce Type: new Abstract: As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a practical questi

model-releasesarxiv-cs-ai
11 May 2026
Safety

> shipped the agent > opened the dashboard > latency: fine > error rate: fine > users: unhappy > checked the responses > technically correct…

DGX agent

> shipped the agent > opened the dashboard > latency: fine > error rate: fine > users: unhappy > checked the responses > technically correct > wrong tool called 3 steps earlier > no trace to follow >

safetyharrison-chase--x
10 May 2026
Model Releases

Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers (Anthropic)

DGX agent

Anthropic: Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers — Last year, we released a case study on

model-releasestechmeme
9 May 2026
Local Ai

Agentic AI Is Going to the Edge – But Not the Way You Think

DGX agent

This article examines how agentic AI systems are being deployed to edge computing environments, likely challenging common assumptions about what this deployment actually entails. It discusses practica

local-aivultr
8 May 2026
← Previous
1…150151152153154…375
Next →