AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,986 results
12 Aug 2026

MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games

AgentsDGX agent

arXiv:2602.24188v2 Announce Type: replace Abstract: We present a scalable and verifiable methodology for evaluating language models in multi-turn interactions, using a suite of collaborative games tha

What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model

AgentsDGX agent

arXiv:2608.10986v1 Announce Type: new Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops. We ask what such

11 Aug 2026

Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety
DGX agent

arXiv:2608.09857v1 Announce Type: cross Abstract: Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on

Agentic Visual Reasoning in Whole-Slide Pathology Images via Active Perception

SafetyDGX agent

arXiv:2608.08648v1 Announce Type: new Abstract: Whole-slide visual reasoning requires identifying sparse diagnostic evidence in gigapixel pathology slides and integrating observations across spatial s

Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents

ResearchDGX agent

arXiv:2608.09485v1 Announce Type: new Abstract: Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission,

PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs

Model ReleasesDGX agent

arXiv:2608.09028v1 Announce Type: new Abstract: Institutional policies stay in natural language while the systems that check compliance demand machine-readable constraints. Bridging that gap is still

really excited to see this release AND integrate it with deepagents!! https://www.langchain.com/blog/switchyard-agent-routing-benchmark

Model ReleasesDGX agent

really excited to see this release AND integrate it with deepagents!! https://www.langchain.com/blog/switchyard-agent-routing-benchmark Lightning strikes for continuous and long-run agents! Nemotron 3

v0.32.9

Model ReleasesDGX agent

NVIDIA Nemotron 3.5 Lightning NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with 3B active parameters built for that execution layer of always-on agents. It is designed f

10 Aug 2026

Autonomous discovery of accelerator commissioning algorithms

AgentsDGX agent

arXiv:2608.07138v1 Announce Type: cross Abstract: Simulated commissioning has become essential for de-risking modern light-source design and commissioning, but the procedures being simulated are still

KNOWPLAN: Knowledge-Driven AI Agents for Smart Degree Pathway Planning

ApplicationsDGX agent

arXiv:2608.06530v1 Announce Type: new Abstract: Planning a degree from official university sources requires solving two problems in order. The institution's curriculum must first be reconstructed from

Tested Muse Glimmer locally on coding with OpenCode & agentic work

Model ReleasesDGX agent

Ran the model with quants (Q4) by Unsloth with latest (build from master) llama.cpp server. It takes ~20GB ram running on M5 Pro with 48GB at about 17t/s. Didn't do any reasoning loops/overthinking. O

Towards Assurance Closure in AI-Native Large-Scale Agile Software Development

AgentsDGX agent

arXiv:2608.07317v1 Announce Type: cross Abstract: The AI-Native Manifesto envisions large-scale agile software development in which humans increasingly govern intent, risk, and exceptions while agents

TRIBE: Predicting Team Performance via Communication Behavior Ensembles

Model ReleasesDGX agent

arXiv:2608.06926v1 Announce Type: new Abstract: Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present

7 Aug 2026

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation

ResearchDGX agent

arXiv:2608.05999v1 Announce Type: new Abstract: Vision-language-action (VLA) models have demonstrated remarkable capabilities in robotic manipulation by leveraging pretrained vision-language models. H

Managing AI Coding Costs at Scale

AgentsDGX agent

**Managing AI Coding Costs at Scale** This article addresses the economic challenges of deploying AI coding tools across large teams or organizations. It explores strategies for tracking resource usag

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing

Model ReleasesDGX agent

arXiv:2608.05485v1 Announce Type: new Abstract: Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and edit

6 Aug 2026

Congrats to @mattrubens and the Roomote team on the launch. Builders can use Together AI as an inference provider in Roomote and assign diff…

AgentsDGX agent

Congrats to @mattrubens and the Roomote team on the launch. Builders can use Together AI as an inference provider in Roomote and assign different open models to coding, planning, vision, and review ac

Formal Analysis and Supply Chain Security for Agentic AI Skills

Model ReleasesDGX agent

arXiv:2603.00195v2 Announce Type: replace-cross Abstract: 32 pages, 5 theorems with full proofs, 68 references, open-source tool: https://github.com/qualixar/skillfortify. v2: corrects the bibliograph

Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First

Model ReleasesDGX agent

arXiv:2608.04804v1 Announce Type: cross Abstract: Frontier language models can resolve repository-level software issues, but each attempt is expensive, and existing routers select a model from the iss

5 Aug 2026

ANCHOR-RE: An Agentic Neuro-Symbolic Framework for Grounded Biomedical Relation Extraction

Model ReleasesDGX agent

arXiv:2608.03154v1 Announce Type: new Abstract: Biomedical relation extraction (BioRE) extracts structured knowledge from biomedical literature for applications such as knowledge base construction and

How Mobileye transformed support operations using Amazon Bedrock AgentCore

AgentsDGX agent

In this post, we'll explore how Mobileye deployed an AI support agentic solution on Amazon Bedrock AgentCore - from the support bottleneck that sparked the idea, through the proof of concept that vali

Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations

ResearchDGX agent

arXiv:2608.03970v1 Announce Type: new Abstract: Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, di

4 Aug 2026

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning

AgentsDGX agent

arXiv:2608.00301v1 Announce Type: cross Abstract: Error-penalized scoring rules (+1 for a correct answer, -lambda for a wrong one, 0 for abstaining) are increasingly prescribed against hallucination:

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

Model ReleasesDGX agent

First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that: https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_2

Deepseek V4 flash 0731 ranks #21 on Agent Arena

Model ReleasesDGX agent

https://preview.redd.it/522fsdwvtdhh1.png?width=1200&format=png&auto=webp&s=6a6cf7a467514167a8193029dbd20fb3a9ba4f6c It ranks lower than both Sonnet 4.6 and Luna. I'd wager Luna costs in the same ball

LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference

Model ReleasesDGX agent

arXiv:2608.02515v1 Announce Type: new Abstract: Long-running assistants and agents consume interaction streams that eventually outgrow the context. Existing context retention, summarization, and retri

PackingGPT: 3D Packing Agent for Real Furniture in Last-Mile Delivery

Model ReleasesDGX agent

arXiv:2608.01427v1 Announce Type: new Abstract: 3D bin packing rectangular items into standardised containers to maximise space utilisation under geometric shipping automation. Loading a furniture pur

SIPTraj: Map-Free End-to-End Trajectory Prediction via Physics-Guided Scene Interaction

Model ReleasesDGX agent

arXiv:2608.00779v1 Announce Type: new Abstract: Trajectory prediction of surrounding agents is a prerequisite for safe planning and decision making in autonomous driving. Without high-definition (HD)

3 Aug 2026

Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving

AgentsDGX agent

arXiv:2607.29031v1 Announce Type: cross Abstract: Existing autonomous-driving world models typically perform dense prediction of future videos, occupancy states, BEV representations, or agent motion.

Here’s why AI agents lie and cheat to reach their goals

ResearchDGX agent

MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here. When two OpenAI model

MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft

Model ReleasesDGX agent

arXiv:2607.29218v1 Announce Type: new Abstract: With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unfortunately, m

2 Aug 2026

Fascinating to see @ClementDelangue, CEO of @huggingface speaking on @FaceTheNation. Excellent points and solid advocacy around the benefits…

AgentsDGX agent

Fascinating to see @ClementDelangue, CEO of @huggingface speaking on @FaceTheNation. Excellent points and solid advocacy around the benefits of open AI models, which helped him defend against a rogue

1 Aug 2026

If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, …

Model ReleasesDGX agent

If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, 17 real tasks from 3 repositories, with context-injection st

ThreatLocker raised a $190M Series F led by Elephant as it looks to extend its zero-trust enterprise security platform to protect against AI-related risks (Kyle Alspach/CRN)

AgentsDGX agent

Kyle Alspach / CRN: ThreatLocker raised a $190M Series F led by Elephant as it looks to extend its zero-trust enterprise security platform to protect against AI-related risks — The cybersecurity vendo

31 Jul 2026

AI as Friction for Reflection Support in Ideation

AgentsDGX agent

arXiv:2607.26827v1 Announce Type: cross Abstract: Generative AI tools for creative work tend to be designed around the goal of removing friction, on the assumption that smoother iteration and faster o

Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer

Local AiDGX agent

arXiv:2607.17100v2 Announce Type: replace-cross Abstract: Auto Research uses language-model agents to propose, implement, and evaluate machine-learning changes in a closed loop, but is usually judged

One Run Is Not an Idea: The Implementation Lottery in Automated Research

AgentsDGX agent

arXiv:2607.26587v1 Announce Type: cross Abstract: Automated research systems use experimental scores both to deliver artifacts and to decide which ideas to retain, transfer, and pursue. Yet one run sc

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Model ReleasesDGX agent

arXiv:2607.28609v1 Announce Type: cross Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Veri

Pushing the Frontier on Approximate EFX Allocations

ResearchDGX agent

arXiv:2406.12413v3 Announce Type: replace-cross Abstract: We study the problem of allocating a set of indivisible goods to a set of agents with additive valuation functions, aiming to achieve approxim

30 Jul 2026

HeteroPROPMT: A Real-time and Privacy-Preserving Heterogeneous Collaborative Perception Framework

AgentsDGX agent

arXiv:2607.26283v1 Announce Type: new Abstract: Collaborative Perception (CP) improves autonomous systems' awareness of their surroundings by sharing sensor data, intermediate features, and detection

29 Jul 2026

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

SafetyDGX agent

arXiv:2607.25308v1 Announce Type: cross Abstract: Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning w

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

Model ReleasesDGX agent

arXiv:2607.25091v1 Announce Type: new Abstract: The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the unde

28 Jul 2026

CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases

AgentsDGX agent

arXiv:2408.03910v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories. Thi

MARS: Multi-hop Adaptive Retrieval and SPARQL Generation for KGQA

AgentsDGX agent

arXiv:2607.14561v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated strong reasoning performance, but their tendency to hallucinate limits their reliability in knowledge

Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations

Model ReleasesDGX agent

arXiv:2607.22610v1 Announce Type: new Abstract: When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those dependencies

27 Jul 2026

One Hand Watches The Other: Dynamic Multi-Agent Cooperation for Sample-Efficient Bimanual Manipulation in Dynamic Environments

Model ReleasesDGX agent

arXiv:2607.22119v1 Announce Type: cross Abstract: Multi-stream robot manipulation policies achieve unparalleled sample efficiency and generalization by modeling actions relative to environmental refer

24 Jul 2026

Our team just shipped Fugu-Ultra v1.1! 🐡 By dynamically orchestrating the latest frontier models, we pushed performance up by 7.9 points. W…

AgentsDGX agent

Our team just shipped Fugu-Ultra v1.1! 🐡 By dynamically orchestrating the latest frontier models, we pushed performance up by 7.9 points. We are now beating Fable 5 in complex coding and reasoning tas

23 Jul 2026

Ling-3.0-flash, the new MoE model from @AntLingAGI, is now free in Nous Portal for the next week! At 124B parameters and 5.1B active, it's q…

AgentsDGX agent

Ling-3.0-flash, the new MoE model from @AntLingAGI, is now free in Nous Portal for the next week! At 124B parameters and 5.1B active, it's quick to run and built for agent workloads: coding, search, r

22 Jul 2026

v0.32.2

Model ReleasesDGX agent

What's Changed launch: keep Claude Code channels available by @hoyyeva in #17210 cmd: remove dead agent prompt wrappers by @ParthSareen in #17227 agent: reorder working directory instruction by @Parth

21 Jul 2026

text/image-to-sim

AgentsDGX agent

Gizmo is a simulation-authoring agent that converts textual descriptions and reference images into structured, editable 3D scenes tailored for robotics workflows. It was publicly released as a beta on

16 Jul 2026

Active Trust Management for Successful Human-Robot Teaming: Moving from a Trust Repair to a Trust Satisficing Perspective

AgentsDGX agent

arXiv:2607.13595v1 Announce Type: new Abstract: Integrating mobile robots into human teams promises significant capability improvements for tasks such as searching hazardous environments. Unlike exist

SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing

SafetyDGX agent

arXiv:2607.13594v1 Announce Type: new Abstract: LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a gua

15 Jul 2026

CityBehavEx: A Scalable and Empirically Validated LLM-Assisted Urban Simulation Platform

SafetyDGX agent

arXiv:2607.12086v1 Announce Type: new Abstract: Recent LLM-based multi-agent urban simulators can generate semantically rich city routines, but they remain costly to scale and are often weakly validat

RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls

Model ReleasesDGX agent

arXiv:2607.12216v1 Announce Type: cross Abstract: Multi-agent and memory-augmented LLM systems often place coordination content, shared state, prior discussion, tool outputs, summaries, and role instr

14 Jul 2026

Google named a Leader in the 2026 IDC MarketScape for Worldwide Foundation Model Software

Model ReleasesDGX agent

For years, we’ve built with a clear priority: putting the practical needs of the enterprise first. Long before generative AI dominated the headlines, we were focused on building the global infrastruct

some good self improvement research here

AgentsDGX agent

some good self improvement research here The first experimental evidence of recursive self-improvement (RSI). Autoresearching the autoresearch agent for eight days. The result beats the harness we han

Text match filters for agents

ToolsDGX agent

Text match filters are a feature in Pinecone that allow users to filter vector search results based on exact text matching criteria, enabling more precise control over which documents or records are r

13 Jul 2026

How do physical systems achieve collective intelligence and self-repair without a central brain? A new paper published today in Nature Commu…

AgentsDGX agent

How do physical systems achieve collective intelligence and self-repair without a central brain? A new paper published today in Nature Communications by my Sakana AI colleague Sebastian Risi (@risi197

I guess image input is the big capability of the models, and tool use can be a substitute for non-omni model output. Still, multimodal voice…

AgentsDGX agent

Ethan Mollick notes that image input represents the primary advanced capability of current AI models, and that tool‑use can effectively replace outputs from non‑omni models. He observes that multimoda

10 Jul 2026

LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis

SafetyDGX agent

arXiv:2606.16149v2 Announce Type: replace Abstract: Rare disease diagnosis involves interpreting clinical and genetic findings through complex diagnostic reasoning. We investigated whether this reason

← Previous
1…150151152153154…300
Next →