AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,724 results
28 Jun 2026

Harness: ✅ (DeepAgents) Sandboxes: ✅ (LangSmith Sandboxes) Eval: ✅ (LangSmith Sandboxes Model: integrate with all the popular models and pro…

ApplicationsDGX agent

Harness: ✅ (DeepAgents) Sandboxes: ✅ (LangSmith Sandboxes) Eval: ✅ (LangSmith Sandboxes Model: integrate with all the popular models and providers Plus we have the engine that helps you turn this flyw

My conversation with @ScottWu46, founder and CEO of @Cognition, the company behind Devin, the first AI software engineer. 0:00 Scott Wu's Ob…

ApplicationsDGX agent

My conversation with @ScottWu46, founder and CEO of @Cognition, the company behind Devin, the first AI software engineer. 0:00 Scott Wu's Obsession With Winning 2:06 Competitive Programming, Games And

NEW paper worth reading. Reasoning-data curation is expensive because scoring a trace usually means reading it to the end. This new work fro…

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
TutorialsDGX agent

NEW paper worth reading. Reasoning-data curation is expensive because scoring a trace usually means reading it to the end. This new work from UCLA shows you may not have to. The quality of a reasoning

Why do RL runs on LLMs blow up even when the recipe looks right? GEOALIGN, from the Alibaba team behind Qwen, points at the rollouts. A hand…

Model ReleasesDGX agent

Why do RL runs on LLMs blow up even when the recipe looks right? GEOALIGN, from the Alibaba team behind Qwen, points at the rollouts. A handful of bad batches push the policy in incoherent directions,

27 Jun 2026

If you use LLM-as-judge, this one is worth reading. (bookmark it) It's actually one of the most effective ways to use LLM-as-a-Judge for eva…

TutorialsDGX agent

If you use LLM-as-judge, this one is worth reading. (bookmark it) It's actually one of the most effective ways to use LLM-as-a-Judge for evals. Holistic judge scores hide both their reasoning and thei

LiteParse is unreasonably good for document parsing ✅ It is the fastest document parsing tool out there - average parse time per page is 3ms…

Model ReleasesDGX agent

LiteParse is unreasonably good for document parsing ✅ It is the fastest document parsing tool out there - average parse time per page is 3ms ⚡️⚡️ ✅ Now that we support markdown, it tops opendataloader

NEW paper from NVIDIA. (bookmark it) Speed-of-light performance analysis tells you the theoretical floor of a workload, but teams still deri…

HardwareDGX agent

NEW paper from NVIDIA. (bookmark it) Speed-of-light performance analysis tells you the theoretical floor of a workload, but teams still derive it by hand and freeze it. SOLAR automates the whole thing

Run Ornith with Ollama: ollama run ornith For coding, use it with Claude or Pi: ollama launch claude --model ornith ollama launch pi --model…

Model ReleasesDGX agent

Run Ornith with Ollama: ollama run ornith For coding, use it with Claude or Pi: ollama launch claude --model ornith ollama launch pi --model ornith For the more capable 35B model, use: ollama launch c

26 Jun 2026

Adaptive Utility driven Resource Orchestration for Resilient AI (AURORA-AI)

SafetyDGX agent

arXiv:2606.27005v1 Announce Type: new Abstract: Modern AI systems are increasingly deployed under non-stationary computational, demographic, and operational conditions in which static resource allocat

Beyond Feedforward Networks: Reentry Neural Systems as the Fundamental Basis of Subjecthood and Intrinsic Safety of Next-Generation AGI

SafetyDGX agent

arXiv:2606.26406v1 Announce Type: cross Abstract: We propose a complete architectural blueprint for safe artificial general intelligence based on a closed reentry loop (D I cycle). In contrast to feed

Confidence-Aware Tool Orchestration for Robust Video Understanding

Model ReleasesDGX agent

arXiv:2606.26904v1 Announce Type: cross Abstract: Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we term the Blind Trust Problem: unde

Context Recycling for Long-Horizon LLM Inference

Model ReleasesDGX agent

arXiv:2606.26105v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong capabilities in short-context reasoning but degrade in performance over long conversational horizons due t

Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning

SafetyDGX agent

arXiv:2606.26397v1 Announce Type: cross Abstract: Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL) frequently

Empirical Software Engineering TerraProbe: A Layered-Oracle Framework for Detecting Deceptive Fixes in LLM-Assisted Terraform

Model ReleasesDGX agent

arXiv:2606.26590v1 Announce Type: new Abstract: Security misconfigurations in Terraform Infrastructure-as-Code are a growing risk in cloud deployments, and large language models are increasingly used

Forward Deployed Engineering is the most critical ingredient for Enterprise AI adoption. It’s why the most important companies on the planet…

TutorialsDGX agent

Forward Deployed Engineering is the most critical ingredient for Enterprise AI adoption. It’s why the most important companies on the planet are investing so heavily into it. That’s why we put togethe

I was at the first AI Engineer Summit @aiDotEngineer ~500 people, limited admission, felt like a secret. Next week it takes over Moscone Wes…

ToolsDGX agent

I was at the first AI Engineer Summit @aiDotEngineer ~500 people, limited admission, felt like a secret. Next week it takes over Moscone West: thousands of engineers, 400+ sessions. Huge props to @swy

Learning Adversarial Augmentation Policies for Robust Garlic Seedling Detection

Local AiDGX agent

arXiv:2606.26828v1 Announce Type: new Abstract: Accurate seedling detection during early growth stages is essential for timely replanting and effective crop management in precision agriculture. Howeve

Reducing Conversational Escalation in Large Language Model Dialogue with Nonviolent Communication Constraints

SafetyDGX agent

arXiv:2606.26106v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in emotionally charged situations involving interpersonal conflict, frustration, and distress. Whil

The Capability Frontier: Benchmarks Miss 82% of Model Performance

Model ReleasesDGX agent

arXiv:2606.26836v1 Announce Type: new Abstract: Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabilities, particula

The Open Source Economic Index of AI Adoption and Capability

Model ReleasesDGX agent

arXiv:2606.26118v1 Announce Type: cross Abstract: We work towards measuring both AI adoption and the capability of AI to perform discrete labor tasks across various occupations. To measure adoption, w

Theory-Scale Auto-Formalization of Logics for Computer Science

Model ReleasesDGX agent

arXiv:2606.26525v1 Announce Type: new Abstract: Auto-formalization is critical for scalable formal verification, but existing progress largely focuses on isolated statements, while theory-scale auto-f

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

Model ReleasesDGX agent

arXiv:2606.26443v1 Announce Type: cross Abstract: A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object

25 Jun 2026

Bias-Controlled Primal-Dual Natural Actor-Critic: Optimal Rates for Constrained Multi-Objective Average-Reward RL

SafetyDGX agent

arXiv:2606.25012v1 Announce Type: new Abstract: Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisf

CustomX: Unified Character, Action, and Scene Customization in Video World Models

ResearchDGX agent

arXiv:2512.17796v2 Announce Type: replace Abstract: Recent advances in world models have greatly enhanced interactive environment simulation. Existing methods mainly fall into two categories: (1) stat

Event-Adaptive Motion Planning with Distilled Vision-Language Model in Safety-Critical Situations

SafetyDGX agent

arXiv:2606.25629v1 Announce Type: new Abstract: Robot navigation in safety-critical scenarios faces significant challenges from unforeseen semantic events, where collisions arise primarily from the un

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation

SafetyDGX agent

arXiv:2606.26006v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are often constrained by the imitation ceiling imposed by sub-optimal data. While Reinforcement Learning (RL) fine-t

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

Model ReleasesDGX agent

arXiv:2606.26071v1 Announce Type: new Abstract: A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior. But beh

Ok, I am interviewing the legend @trq212 on Friday. I'm planning to ask him to show me: → His Claude Code setup and how he uses /goal and /l…

Model ReleasesDGX agent

Ok, I am interviewing the legend @trq212 on Friday. I'm planning to ask him to show me: → His Claude Code setup and how he uses /goal and /loop and dynamic workflows → How he does planning with HTML +

Paris-based Tsuga, whose observability software runs in customers' own clouds to help them skip per-byte ingestion costs, raised a $35M Series A led by Singular (Ana Maria Constantin/The Next Web)

IndustryDGX agent

Ana Maria Constantin / The Next Web: Paris-based Tsuga, whose observability software runs in customers' own clouds to help them skip per-byte ingestion costs, raised a 35M Series A led by Singular — T

Project Auto-World: Towards Automated Benchmarking of Neural Relational Reasoners

Model ReleasesDGX agent

arXiv:2606.24965v1 Announce Type: cross Abstract: Reasoning about relational structures remains a significant challenge for neural models, particularly when they must systematically apply learned know

RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments

Model ReleasesDGX agent

arXiv:2606.26094v1 Announce Type: new Abstract: For most of scientific history, researchers studying behavior could only infer hidden mechanisms from outward actions: an inverse problem that becomes m

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

SafetyDGX agent

arXiv:2601.16529v3 Announce Type: replace Abstract: Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-based gui

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning

Model ReleasesDGX agent

arXiv:2606.25319v1 Announce Type: new Abstract: Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evidence and ground their reasoning in

What happens when Claude Code gets an experiment tracker

Model ReleasesDGX agent

At CVPR 2026, Lambda ran a live demo for two and a half days: Claude Code teaching Google's Gemma 4 to play a Tetris-like game. Claude Code started with a Gemma 4 model that couldn't play at all. It p

24 Jun 2026

A global log for medical AI

SafetyDGX agent

arXiv:2510.04033v2 Announce Type: replace Abstract: Modern computer systems rely on syslog, a universal protocol that records critical events across heterogeneous infrastructure. Medicine's rapidly gr

Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment

Local AiDGX agent

arXiv:2606.24834v1 Announce Type: new Abstract: LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on functional co

Anthropic debuts Claude Tag, a more capable AI teammate that lives within Slack

Model ReleasesDGX agent

Anthropic PBC today unveiled a new version of its chatbot Claude that lives inside Slack, where it operates like a virtual employee. It’s called Claude Tag, and it’s designed to work across entire org

Beyond the Autoregressive Horizon: A Comprehensive Survey of Diffusion Models, World Modelling, and State Space Models for Code

ResearchDGX agent

arXiv:2606.23690v1 Announce Type: cross Abstract: Autoregressive (AR) language models have driven significant progress in automated software engineering, enabling powerful code generation and assistan

Breaking the Filter Bubble: A Semantic Pareto-DQN Framework for Multi-Objective Recommendation

SafetyDGX agent

arXiv:2606.24042v1 Announce Type: new Abstract: Recommender systems often induce filter bubbles and semantic homogenization by monolithically optimizing for immediate user engagement. Standard single-

Decentralized Pose Graph Riemannian Optimization for Object-based Multi-Robot SLAM

ApplicationsDGX agent

arXiv:2606.24489v1 Announce Type: new Abstract: Pose graph optimization (PGO) is a key back-end component for state estimation in networked multi-robot simultaneous localization and mapping (SLAM). In

DramaDirector: Geometry-Guided Short Drama Generation

Model ReleasesDGX agent

arXiv:2606.24107v1 Announce Type: cross Abstract: Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation

Model ReleasesDGX agent

arXiv:2512.24731v2 Announce Type: replace Abstract: Sound effects build an essential layer of multimodal storytelling, shaping the emotional atmosphere and the narrative semantics of videos. Despite r

EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games

Model ReleasesDGX agent

arXiv:2606.23995v1 Announce Type: cross Abstract: Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized game-theoret

FALCON: Transforming Cyber Threat Intelligence into Deployable IDS Rules with Self-Reflection

SafetyDGX agent

arXiv:2508.18684v2 Announce Type: replace-cross Abstract: Signature-based Intrusion Detection Systems (IDS) detect malicious activity by matching network or host events against predefined rules. Secur

Figma now has AI motion graphics and shader tools

IndustryDGX agent

Figma has revealed some new design and coding product updates at its annual Config conference that aim to help creatives 'push their ideas further' and automate tedious tasks with AI. Part of this is

Knowledge-Graph Grounding Helps LLMs Only for Out-of-Training Knowledge: A Controlled Study on Clinical Question Answering

Model ReleasesDGX agent

arXiv:2606.22419v2 Announce Type: replace Abstract: A recent Nature Medicine study reports that general-purpose frontier LLMs outperform specialized retrieval-augmented clinical tools on medical bench

NoContactNoWorries: Estimating Contact through Vision and Proprioception for In-Hand Dexterous Manipulation

TutorialsDGX agent

arXiv:2606.24450v1 Announce Type: cross Abstract: Perceiving physical contact is fundamental to dexterous manipulation. While robots often rely on dedicated hardware tactile sensors, humans exhibit a

PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation

SafetyDGX agent

arXiv:2606.24081v1 Announce Type: cross Abstract: As Text-to-Image (T2I) jailbreak techniques evolve rapidly, existing benchmarks and reproduction workflows often struggle to keep pace. More important

Q&A with AWS CEO Matt Garman on the parallels between early AWS and AI, Quick, AI coding, Amazon's $200B capex in 2026, hiring entry-level staff, and more (Casey Newton/Platformer)

IndustryDGX agent

Casey Newton / Platformer: Q&A with AWS CEO Matt Garman on the parallels between early AWS and AI, Quick, AI coding, Amazon's $200B capex in 2026, hiring entry-level staff, and more — Matt Garman argu

Qdrant Lands in SF: Vector Space Day 2026 Recap

ToolsDGX agent

On June 11th, 2026, over 350 developers, researchers, and engineers came together at The Midway in San Francisco for Vector Space Day, our first event of its kind in the United States and our first ma

REALM: A Unified Red-Teaming Benchmark for Physical-World VLMs

Model ReleasesDGX agent

arXiv:2606.23892v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used as perception-reasoning backbones for embodied intelligence in safety-critical physical systems, whe

RoBoSR: Structured Scene Representations for Embodied Robotic Reasoning

SafetyDGX agent

arXiv:2606.24338v1 Announce Type: new Abstract: Despite rapid progress, embodied reasoning under real-world variability remains challenging. Existing approaches rely on demonstration-driven sequential

Topological Online Learning for Displacement-based Formation Control

ResearchDGX agent

arXiv:2606.23901v1 Announce Type: new Abstract: This paper addresses the problem of robust formation control by introducing Topological Online Learning for Displacement-based (TOLD) formation control,

We open-source Qwen-AgentWorld-35B-A3B (MoE, 35B/3B active, 256K context) and AgentWorldBench. Two routes, one roadmap: 🔬 Build the simulat…

Model ReleasesDGX agent

We open-source Qwen-AgentWorld-35B-A3B (MoE, 35B/3B active, 256K context) and AgentWorldBench. Two routes, one roadmap: 🔬 Build the simulator — scalable, controllable, surpassing real environments 🧠 I

When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs

Model ReleasesDGX agent

arXiv:2606.24370v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into decision-support roles in business and policy contexts. While prior benchmark studies have

23 Jun 2026

A Differentiable Atari VCS:A Complex, Fully Known Ground Truth for Explainable AI

HardwareDGX agent

arXiv:2606.22447v1 Announce Type: cross Abstract: Explanation requires ground truth: to verify an account of a system we must know its inner functioning-just what is missing where explainable AI (XAI)

A DVDrive Approach for doScenes Instructed Driving Challenge

Local AiDGX agent

arXiv:2606.21623v1 Announce Type: new Abstract: Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only fr

Beehiiv adds Cloudflare AI Crawl Control so writers can block or allow bots

IndustryDGX agent

Cloudflare Inc. and newsletter platform beehiiv Inc. today launched an integration that hands independent publishers a single toggle to decide whether artificial intelligence crawlers can reach their

Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies

Model ReleasesDGX agent

arXiv:2606.20599v1 Announce Type: cross Abstract: Tree of Thought (ToT) search has become a promising direction for improving the reasoning capabilities of large language models, but deploying these m

Conflict-Aware Switching for CBF-CLF-Based Multi-Goal Navigation

SafetyDGX agent

arXiv:2606.21577v1 Announce Type: new Abstract: Quadratic programs (QPs) using Control Barrier Functions (CBFs) and Control Lyapunov Functions (CLFs) are widely used for safe control in reach-and-avoi

← Previous
1…271272273274275…296
Next →