AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,612 results
4 Jun 2026

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

Model ReleasesDGX agent

arXiv:2603.10044v2 Announce Type: replace-cross Abstract: A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark never

SAM 3D: 3Dfy Anything in Images

Model ReleasesDGX agent

arXiv:2511.16624v2 Announce Type: replace-cross Abstract: We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single i

Scaling AI Agents: A Step-by-Step Guide to Deploying ADK on GKE Autopilot

Model ReleasesDGX agent

While building AI agents locally using Google’s Agent Development Kit (ADK) is an excellent way to prototype, production-ready agents require a robust, scalable infrastructure. For developers looking


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Scene-Centric Unsupervised Video Panoptic Segmentation

Model ReleasesDGX agent

arXiv:2606.04925v1 Announce Type: new Abstract: Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regio

Self-Evolving Deep Research via Joint Generation and Evaluation

Model ReleasesDGX agent

arXiv:2606.04507v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capab

Shifting the Breaking Point of Flow Matching for Multi-Instance Editing

Model ReleasesDGX agent

arXiv:2602.08749v3 Announce Type: replace Abstract: Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offeri

Signed Dual Attention: Capturing Signed Dependencies in Time Series Forecasting

Model ReleasesDGX agent

arXiv:2606.04833v1 Announce Type: cross Abstract: Initially developed for natural language processing, Transformer architectures and attention mechanisms are now central to a wide range of deep learni

SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language Models

Model ReleasesDGX agent

arXiv:2606.04202v1 Announce Type: new Abstract: As LLMs become more widely deployed, they are increasingly expected to work alongside other AI agents rather than operating in isolation. Effective coor

SMADE-IE: Sparse Multi-Agent Framework with Evidence-Driven Debate for Zero-Shot Information Extraction

Model ReleasesDGX agent

arXiv:2606.04691v1 Announce Type: new Abstract: Zero-shot information extraction (IE) with large language models (LLMs) has attracted increasing attention due to its flexibility in adapting to new sch

Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2407.03956v3 Announce Type: replace-cross Abstract: Prior research has enhanced the ability of Large Language Models (LLMs) to solve logic puzzles using techniques such as chain-of-thought promp

Sources: Anthropic has embedded around half a dozen forward-deployed engineers within the NSA to help the agency deploy Mythos for offensive cyber operations (Financial Times)

Model ReleasesDGX agent

Financial Times: Sources: Anthropic has embedded around half a dozen forward-deployed engineers within the NSA to help the agency deploy Mythos for offensive cyber operations — Arrangement comes as AI

Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time

Model ReleasesDGX agent

arXiv:2504.12329v2 Announce Type: replace-cross Abstract: Recent advances leverage post-training to enhance model reasoning performance, which typically requires costly training pipelines and still su

SPLIT-PINN: Separable Probability Learning Technique via Physics-Informed Neural Networks for High-Dimensional Probabilistic Modeling

Model ReleasesDGX agent

arXiv:2606.04000v1 Announce Type: cross Abstract: We present a probabilistic modeling framework for incorporating small-scale spatial heterogeneity into macroscopic descriptions of material behavior f

StandardE2E: A Unified Framework for End-to-End Autonomous Driving Datasets

Model ReleasesDGX agent

arXiv:2606.04271v1 Announce Type: cross Abstract: Autonomous driving has shifted from modular perception-prediction-planning stacks toward end-to-end (E2E) models that map sensor inputs directly to ve

StepPRM-RTL: Stepwise Process-Reward Guided LLM Fine-Tuning for Enhanced RTL Synthesis

Model ReleasesDGX agent

arXiv:2606.04246v1 Announce Type: new Abstract: Automatic generation of RTL code for digital hardware designs remains challenging due to long-horizon reasoning, multi-step dependencies, and strict cor

Stepwise Reasoning Enhancement for LLMs via External Subgraph Generation

Model ReleasesDGX agent

arXiv:2606.04454v1 Announce Type: new Abstract: Large language models have shown strong performance in natural language generation and downstream reasoning tasks, but they still struggle with logical

Streaming Communication in Multi-Agent Reasoning

Model ReleasesDGX agent

arXiv:2606.05158v1 Announce Type: cross Abstract: Multi-agent reasoning systems adopt a 'generate-then-transfer' paradigm that forces end-to-end latency to scale linearly with pipeline depth. We intro

STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations

Model ReleasesDGX agent

arXiv:2606.05165v1 Announce Type: cross Abstract: Training Data Attribution (TDA) seeks to trace a model's predictions back to its training data. The gold standard for TDA relies on causal interventio

Structure-Aware Prediction of PROTAC-Mediated Protein Degradability via Graph Neural Networks

Model ReleasesDGX agent

arXiv:2606.04021v1 Announce Type: cross Abstract: Proteolysis-targeting chimeras (PROTACs) can selectively degrade disease-causing proteins, yet predicting which targets are amenable to degradation re

Symbolic Regression for Shared Expressions: Introducing Partial Parameter Sharing

Model ReleasesDGX agent

arXiv:2601.04051v3 Announce Type: replace Abstract: Symbolic regression aims to find symbolic expressions that describe datasets. Due to its inherent interpretability, symbolic regression (SR) is a po

SymTRELLIS: Symmetry-Enforced Voxel Latents for 3D Generation

Model ReleasesDGX agent

arXiv:2606.04108v1 Announce Type: cross Abstract: Single-view 3D generative models have achieved impressive visual quality, yet they are not designed to satisfy structural or functional requirements,

TaDA: Calibrated Probe Gating for Task-Domain LoRA Merging

Model ReleasesDGX agent

arXiv:2606.05016v1 Announce Type: new Abstract: Combining a task LoRA adapter with a domain LoRA adapter into a single unified model is a practical yet largely unexplored challenge. Existing methods t

Test-Time Compute Scaling for ASR with Depth-Conditioned Looped Transformers

Model ReleasesDGX agent

arXiv:2606.04678v1 Announce Type: new Abstract: End-to-end ASR systems typically use fixed-depth acoustic encoders at inference, making it difficult to trade additional test-time computation for impro

That's a badass title, and it's true! Every day, it gets harder and harder to create tests that AI models can't beat. Reality is humanity's …

Model ReleasesDGX agent

That's a badass title, and it's true! Every day, it gets harder and harder to create tests that AI models can't beat. Reality is humanity's real last exam. Andon Labs' Real-World AI Evals: Claude call

The Canadian AI strategy unveiled today advocates for the development of technology that is safe, ethical, trustworthy, and that benefits so…

Model ReleasesDGX agent

The Canadian AI strategy unveiled today advocates for the development of technology that is safe, ethical, trustworthy, and that benefits society as a whole—these are exactly the principles that need

The capabilities of Claude Code and Codex have expanded a lot in recent months, they added many ways to approach work (subagents, skills, go…

Model ReleasesDGX agent

The capabilities of Claude Code and Codex have expanded a lot in recent months, they added many ways to approach work (subagents, skills, goal, workflows, plugins, etc). Given the AI labs can use thei

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?

Model ReleasesDGX agent

arXiv:2606.04455v1 Announce Type: new Abstract: Current AI benchmarks evaluate agents on task execution within human-designed workflows. These evaluations fundamentally fail to measure a critical next

The new memory system will keep track of important details automatically. If you prefer the legacy saved memories experience, you can switch…

Model ReleasesDGX agent

The new memory system will keep track of important details automatically. If you prefer the legacy saved memories experience, you can switch back in settings. The new memory system is rolling out to P

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents

Model ReleasesDGX agent

arXiv:2606.04296v1 Announce Type: new Abstract: As autonomous AI agents move from conversational systems to long-horizon software execution, runtime safety layers that decide when to interrupt an agen

The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail

Model ReleasesDGX agent

arXiv:2606.04010v1 Announce Type: cross Abstract: Brain foundation models (BFMs) are self-supervised Transformers pretrained on fMRI data. We posit that these models should capture each subject's cogn

Thinking Through Signs: PEEL as a Semiotic Scaffolding for Epistemically Accountable AI-Enabled Research

Model ReleasesDGX agent

arXiv:2606.04152v1 Announce Type: new Abstract: Large language models are reshaping research practice while quietly eroding researchers epistemic accountability. This commentary introduces PEEL - Prot

Today I'm launching a new project called SynthTraces 🔥 It is a minimal codebase to generate synthetic coding agent session traces using Pi …

Model ReleasesDGX agent

Today I'm launching a new project called SynthTraces 🔥 It is a minimal codebase to generate synthetic coding agent session traces using Pi (from @badlogicgames) I wanted a large number of coding-agent

Toward a Generalized Defense Across Sparse, Continuous, and Structured Parameter Attacks

Model ReleasesDGX agent

arXiv:2606.04317v1 Announce Type: cross Abstract: Deep neural networks are increasingly deployed across heterogeneous and partially untrusted environments, where models are distributed through cloud s

Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification

Model ReleasesDGX agent

arXiv:2606.04037v1 Announce Type: new Abstract: Pre-deployment verification of enterprise artificial intelligence (AI) agents remains a critical gap between large language model (LLM) capability bench

Toward Trustworthy Portrait Editing: Evaluation of Demographic Misrepresentation in I2I Models

Model ReleasesDGX agent

arXiv:2602.16149v2 Announce Type: replace Abstract: Instruction-guided image-to-image (I2I) editors are increasingly used in consumer and professional visual workflows, where trustworthiness depends n

Towards Efficient and Evidence-grounded Mobility Prediction with LLM-Driven Agent

Model ReleasesDGX agent

arXiv:2606.05130v1 Announce Type: cross Abstract: Individual-level mobility prediction is central to urban simulation, transportation planning, and policy analysis. Supervised sequence models achieve

transitions like this are why we think it's helpful to have a provider-agnostic harness we used to talk more about swapping models when the …

Model ReleasesDGX agent

transitions like this are why we think it's helpful to have a provider-agnostic harness we used to talk more about swapping models when the latest and greatest came out -- but the latest and greatest

Treat Traffic Like Trees: A Semantic-Preserving Hierarchical Graph-Based Expert Framework for Encrypted Traffic Analysis

Model ReleasesDGX agent

arXiv:2606.04517v1 Announce Type: cross Abstract: Graph-based deep learning methods have been widely employed in encrypted traffic analysis to exploit latent correlations across different granularitie

Tree-Based Formalization of Multi-Agent Complementarity in Human-AI Interactions

Model ReleasesDGX agent

arXiv:2606.04779v1 Announce Type: new Abstract: Complementarity is the case in which a human--AI interaction (HAI) outperforms the best prediction benchmark available among its members. Although this

Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from k-Parity

Model ReleasesDGX agent

arXiv:2601.22450v2 Announce Type: replace-cross Abstract: Masked Diffusion Language Models have recently emerged as a powerful generative paradigm, yet their generalization properties remain understud

UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

Model ReleasesDGX agent

arXiv:2606.05058v1 Announce Type: cross Abstract: Computer-Aided Design (CAD) underpins modern engineering and manufacturing by enabling the creation of precise, editable 3D models. However, CAD resea

Unifying Model-Free Efficiency and Model-Based Representations via Latent Dynamics

Model ReleasesDGX agent

arXiv:2602.12643v2 Announce Type: replace-cross Abstract: We present Unified Latent Dynamics (ULD), a novel reinforcement learning algorithm that unifies the efficiency of model-free methods with the

Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs

Model ReleasesDGX agent

arXiv:2606.04035v1 Announce Type: cross Abstract: We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5

VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark

Model ReleasesDGX agent

arXiv:2606.04244v1 Announce Type: new Abstract: Multimodal large language models are increasingly capable of complex reasoning, yet their performance often degrades when they must externalize a proble

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

Model ReleasesDGX agent

arXiv:2606.04588v1 Announce Type: new Abstract: Multimodal large language models have made rapid progress in video understanding, yet existing benchmarks largely rely on simple prompts and provide lim

VGGSounder: Audio-Visual Evaluations for Foundation Models

Model ReleasesDGX agent

arXiv:2508.08237v4 Announce Type: replace-cross Abstract: The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound

Video2LoRA: Parametric Video Internalization for Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.04351v1 Announce Type: cross Abstract: Processing video in vision-language models is expensive: each frame occupies hundreds of tokens, and inference cost scales with every frame and every

We are excited to join Nvidia's Nemotron Coalition of leading AI labs working together to advance open frontier foundation models. To celebr…

Model ReleasesDGX agent

We are excited to join Nvidia's Nemotron Coalition of leading AI labs working together to advance open frontier foundation models. To celebrate we have partnered with @nvidia and @nebiustf to provide

We're building in Canada. 🇨🇦

Model ReleasesDGX agent

We're building in Canada. 🇨🇦 For decades, Canada invested to build the research foundations that made modern AI possible. Now we have to build, train, and scale what comes next here at home. Canada's

We're presenting ParseBench at CVPR 2026! ParseBench is the most comprehensive document understanding benchmark for VLMs. ✅ It contains 2k p…

Model ReleasesDGX agent

We're presenting ParseBench at CVPR 2026! ParseBench is the most comprehensive document understanding benchmark for VLMs. ✅ It contains 2k pages of real-world enterprise documents ✅ It has comprehensi

We're presenting ParseBench at CVPR 2026 today. 🦙 Come learn why document understanding is an AGI-complete problem (an agent can't act on a…

Model ReleasesDGX agent

We're presenting ParseBench at CVPR 2026 today. 🦙 Come learn why document understanding is an AGI-complete problem (an agent can't act on a doc it can't correctly read, and reading a real enterprise t

WETBench: A Benchmark for Detecting Task-Specific Machine-Generated Text on Wikipedia

Model ReleasesDGX agent

arXiv:2507.03373v2 Announce Type: replace Abstract: Given Wikipedia's role as a trusted source of high-quality, reliable content, concerns are growing about the proliferation of low-quality machine-ge

We’ve been researching new ways for ChatGPT memory to carry context across conversations and keep it useful over time. Today, that work is r…

Model ReleasesDGX agent

We’ve been researching new ways for ChatGPT memory to carry context across conversations and keep it useful over time. Today, that work is rolling out as a more capable memory system in ChatGPT. https

What Are We Actually Benchmarking in Robot Manipulation?

Model ReleasesDGX agent

arXiv:2606.04233v1 Announce Type: new Abstract: A robotics benchmark score measures success under one fixed evaluation setup, yet is routinely treated as evidence of general manipulation capability. W

What happened when one of our models found a counterexample to an 80-year-old Erdős conjecture? Researchers @alexwei_, @HongxunWu, and @wjmz…

Model ReleasesDGX agent

What happened when one of our models found a counterexample to an 80-year-old Erdős conjecture? Researchers @alexwei_, @HongxunWu, and @wjmzbmr1 shared the story on the OpenAI Podcast with @AndrewMayn

What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in Agentic Systems

Model ReleasesDGX agent

arXiv:2606.04425v1 Announce Type: cross Abstract: Modern agentic systems transform LLMs from session-bounded assistants into stateful systems that persist and evolve shared world state across sessions

What's new for Managed Service for Apache Spark clusters

Model ReleasesDGX agent

At Google Cloud, our goal is to let you run large-scale analytical and data science workloads with maximum efficiency so you can process big data pipelines, machine learning, and ETL tasks. We recentl

When Do Fewer Coordinates Suffice in DP-SGD?

Model ReleasesDGX agent

arXiv:2606.04375v1 Announce Type: new Abstract: Differentially private stochastic gradient descent (DP-SGD) injects noise into every updated coordinate, making the injected noise energy scale with the

When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation Detection

Model ReleasesDGX agent

arXiv:2606.04098v1 Announce Type: new Abstract: Video misinformation increasingly operates at the semantic and evidential level: authentic footage may be selectively edited, temporally reordered, spli

When you burn so much money you run out of options…

Model ReleasesDGX agent

When you burn so much money you run out of options… Anthropic co-founder and President Daniela Amodei said the high cost of developing AI models is driving firms like hers to look to the public market

← Previous
1…171172173174175…377
Next →