AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,569 results
12 May 2026

What Will Happen Next: Large Models-Driven Deduction for Emergency Instances

Model ReleasesDGX agent

arXiv:2605.08599v1 Announce Type: new Abstract: Traditional simulation methods reproduce occurred emergency instances through presetting to assist people in risk assessment and emergency decision-maki

What’s new in Microsoft Foundry | April 2026

Model ReleasesDGX agent

April brings Foundry Local GA for local AI development, GPT-5.5 model support with Tier 5 and Tier 6 default quota in Microsoft Foundry, new tracing paths for Microsoft Agent Framework and hosted agen

What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2601.20164v2 Announce Type: replace-cross Abstract: Prior work suggests that language models, while trained on next token prediction, show implicit planning behavior: they may select the next to

When Adaptation Fails: A Gradient-Based Diagnosis of Collapsed Gating in Vision-Language Prompt Learning

Model ReleasesDGX agent

arXiv:2605.09549v1 Announce Type: new Abstract: Adaptive prompting mechanisms have been proposed to enhance vision-language models by dynamically tailoring prompts to inputs. However, in frozen few-sh

When (and How) to Trust the Expert: Diagnosing Query-Time Expert-Guided Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.09109v1 Announce Type: new Abstract: Many continuous-control problems ship with a competent but suboptimal controller (a tuned PID, a hand-designed gait). A growing family of methods uses s

When Attention Beats Fourier: Multi-Scale Transformers for PDE Solving on Irregular Domains

Model ReleasesDGX agent

arXiv:2605.08318v1 Announce Type: cross Abstract: We study the problem of architecture selection for deep learning models trained to solve partial differential equations (PDEs), asking when transforme

When Does Non-Uniform Replay Matter in Reinforcement Learning?

Model ReleasesDGX agent

arXiv:2605.10236v1 Announce Type: cross Abstract: Modern off-policy reinforcement learning algorithms often rely on simple uniform replay sampling and it remains unclear when and why non-uniform repla

When is the last time a general purpose LLM (putting aside hybrid systems like Claude Code with special purpose symbolic harnesses) last com…

Model ReleasesDGX agent

When is the last time a general purpose LLM (putting aside hybrid systems like Claude Code with special purpose symbolic harnesses) last completely blew away all competing prior models? GPT-4 relative

When Prompts Become Payloads: A Framework for Mitigating SQL Injection Attacks in Large Language Model-Driven Applications

Model ReleasesDGX agent

arXiv:2605.10176v1 Announce Type: cross Abstract: Natural language interfaces to structured databases are becoming increasingly common, largely due to advances in large language models (LLMs) that ena

When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews

Model ReleasesDGX agent

arXiv:2605.10171v1 Announce Type: cross Abstract: Scientific peer reviews frequently contain conflicting expert judgments, and the increasing scale of conference submissions makes it challenging for A

When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning

Model ReleasesDGX agent

arXiv:2605.09860v1 Announce Type: new Abstract: Long-horizon reasoning requires deciding not only what actions to take, but how deeply to commit before the next observation. We formalize this as commi

When to Trust Imagination: Adaptive Action Execution for World Action Models

Model ReleasesDGX agent

arXiv:2605.06222v2 Announce Type: replace-cross Abstract: World Action Models (WAMs) have recently emerged as a promising paradigm for robotic manipulation by jointly predicting future visual observat

Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing

Model ReleasesDGX agent

arXiv:2605.10544v1 Announce Type: new Abstract: Long-context adaptation is often viewed as window scaling, but this misses a token-level supervision mismatch: in packed training with document masking,

Why Retrying Fails: Context Contamination in LLM Agent Pipelines

Model ReleasesDGX agent

arXiv:2605.08563v1 Announce Type: new Abstract: When an LLM agent fails a multi-step tool-augmented task and retries, the failed attempt typically remains in its context window -- contaminating the ne

Why Zeroth-Order Adaptation May Forget Less: A Randomized Shaping Theory

Model ReleasesDGX agent

arXiv:2605.10658v1 Announce Type: new Abstract: Continual learning requires new-task adaptation without damaging previously acquired capabilities. Recent forward-pass and zeroth-order (ZO) results sho

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation

Model ReleasesDGX agent

arXiv:2605.10912v1 Announce Type: new Abstract: Large language and vision-language models increasingly power agents that act on a user's behalf through command-line interface (CLI) harnesses. However,

WindINR: Latent-State INR for Fast Local Wind Query and Correction in Complex Terrain

Model ReleasesDGX agent

arXiv:2605.09511v1 Announce Type: new Abstract: Many downstream decisions in complex terrain require fast wind estimates at a small number of user-specified locations and heights for a given forecast

Wordle 1,787 3/6 ⬛⬛⬛⬛🟨 🟨⬛🟨⬛⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result where the player solved puzzle #1,787 in three attempts, using the color-coded emoji system to show letter placement feedback (black for wrong letters, yellow

Wordle 1,788 5/6 ⬛⬛⬛⬛⬛ 🟨⬛⬛⬛🟨 ⬛🟨🟨⬛⬛ 🟩🟩🟩⬛⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a completed game of Wordle (puzzle #1,788) played by Anthropic, showing the progression of guesses across six attempts with color-coded feedback indicating correct letter positions

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

Model ReleasesDGX agent

arXiv:2605.10434v1 Announce Type: new Abstract: Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving i

You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models

Model ReleasesDGX agent

arXiv:2510.03761v2 Announce Type: replace-cross Abstract: The widespread use of preprint repositories such as arXiv has accelerated the communication of scientific results but also introduced overlook

Your Simulation Runs but Solves the Wrong Physics: PDE-Grounded Intent Verification for LLM-Generated Multiphysics Simulation Code

Model ReleasesDGX agent

arXiv:2605.09360v1 Announce Type: cross Abstract: Execution-based evaluation of LLM-generated code implicitly treats successful execution as a proxy for correctness. In scientific simulation, this pro

Zero-Shot Chinese Character Recognition via Global-Local Dual-Branch Alignment and Hierarchical Inference

Model ReleasesDGX agent

arXiv:2605.08814v1 Announce Type: new Abstract: Chinese character categories are extremely large, and unseen characters frequently arise in open-world scenarios, making zero-shot Chinese character rec

11 May 2026

+1 to this. I was recently on a cross-continental flight without wifi, so I brought up Qwen3.6 & Gemma 4 (via @ollama) in Deep Agents on my …

Model ReleasesDGX agent

+1 to this. I was recently on a cross-continental flight without wifi, so I brought up Qwen3.6 & Gemma 4 (via @ollama) in Deep Agents on my laptop. admittedly, they fell over on some more involved/com

2.5-D Decomposition for LLM-Based Spatial Construction

Model ReleasesDGX agent

arXiv:2605.07066v1 Announce Type: new Abstract: Autonomous systems that build structures from natural-language instructions need reliable spatial reasoning, yet large language models (LLMs) make syste

3 weeks since ml-intern launched and we just hit 1M messages exchanged. that's 3.3 agent-years of ML research in 21 days. 2 months worth of …

Model ReleasesDGX agent

3 weeks since ml-intern launched and we just hit 1M messages exchanged. that's 3.3 agent-years of ML research in 21 days. 2 months worth of research every day. 17,383 training jobs total. talk about A

A Causal Diffusion Model for Video Reconstruction from Ultra-Low-Bitrate Representations

Model ReleasesDGX agent

arXiv:2602.13837v2 Announce Type: replace Abstract: We study video reconstruction from ultra-low-bitrate representations, where the primary challenge shifts from encoding to decoding. In this regime,

A Hierarchical Ensemble Pipeline for Anomaly Detection in ESA Satellite Telemetry

Model ReleasesDGX agent

arXiv:2605.06681v1 Announce Type: cross Abstract: A hierarchical ensemble pipeline is introduced to address anomaly detection in multivariate telemetry data provided by European Space Agency (ESA). Th

A Marine Debris Detection Framework for Ocean Robots via Self-Attention Enhancement and Feature Interaction Optimization

Model ReleasesDGX agent

arXiv:2605.07388v1 Announce Type: new Abstract: Marine debris detection for ocean robot is crucial for ecological protection, yet performance is often degraded by low-quality images with blur, complex

A Reproducible Multi-Architecture Baseline for Token-Level Chinese Metaphor Identification under the MIPVU Framework

Model ReleasesDGX agent

arXiv:2605.07170v1 Announce Type: new Abstract: Metaphor is pervasive in everyday language, yet token-level computational identification of metaphor-related words in Chinese under the MIPVU framework

A Reproducible Optimisation Protocol for Calibrating Prompt-Based Large Language Model Workflows in Evidence Synthesis

Model ReleasesDGX agent

arXiv:2605.06937v1 Announce Type: new Abstract: This methods article presents a reproducible calibration workflow for prompt-based large language models (LLMs) in structured evidence-synthesis tasks.

A Unified and Controllable Framework for Layered Image Generation with Visual Effects

Model ReleasesDGX agent

arXiv:2601.15507v2 Announce Type: replace Abstract: Recent image generation models produce impressive composites, but often fail to preserve the identity of user-provided content when editing specific

A^2RD: Agentic Autoregressive Diffusion for Long Video Consistency

Model ReleasesDGX agent

arXiv:2605.06924v1 Announce Type: cross Abstract: Synthesizing consistent and coherent long video remains a fundamental challenge. Existing methods suffer from semantic drift and narrative collapse ov

Adapting Vision-Language Models for Neutrino Event Classification in High-Energy Physics

Model ReleasesDGX agent

arXiv:2509.08461v3 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLMs) have demonstrated their remarkable capacity to process and reason over structured and unstruct

Adaptive Memory Decay for Log-Linear Attention

Model ReleasesDGX agent

arXiv:2605.06946v1 Announce Type: cross Abstract: Sequence models face a fundamental tradeoff between memory capacity and computational efficiency. Transformers achieve expressive context modeling at

Adaptive Regularization for Sparsity Control in Bregman-Based Optimizers

Model ReleasesDGX agent

arXiv:2605.07892v1 Announce Type: new Abstract: Sparse training reduces the memory and computational costs of deep neural networks. However, sparse optimization methods, e.g., those adding an ell_1 pe

Agent view is the best Claude Code native way to manage multiple sessions, kind of like tmux built for CC. We spent a lot of time getting th…

Model ReleasesDGX agent

Agent view is the best Claude Code native way to manage multiple sessions, kind of like tmux built for CC. We spent a lot of time getting the details right, I hope you enjoy it. New in Claude Code: ag

AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents

Model ReleasesDGX agent

arXiv:2605.07926v1 Announce Type: new Abstract: As LLM-based agents increasingly rely on external tools, it is important to evaluate their ability to sustain tool-grounded reasoning beyond familiar wo

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

Model ReleasesDGX agent

arXiv:2605.06869v1 Announce Type: new Abstract: AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no u

Agents + file sandboxes are all in the range in 2026 🤖🗃️ This is a nifty reference implementation by @itsclelia showing you how to run you…

Model ReleasesDGX agent

Agents + file sandboxes are all in the range in 2026 🤖🗃️ This is a nifty reference implementation by @itsclelia showing you how to run your agent over a collection of docs (PDFs, images, Office) with

AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents

Model ReleasesDGX agent

arXiv:2605.06607v2 Announce Type: replace-cross Abstract: Recent LLM-based agents have closed substantial portions of the scientific discovery loop in software-only machine-learning research, in chemi

@Alibaba_Qwen Sign up to get free access to Qwen 3.6 Plus and much more at http://portal.nousresearch.com/manage-subscription!

Model ReleasesDGX agent

Nous Research announced free access to Qwen 3.6 Plus and additional features available through their subscription portal at portal.nousresearch.com. This promotion likely provides users with complimen

Amortized Molecular Optimization via Group Relative Policy Optimization

Model ReleasesDGX agent

arXiv:2602.12162v3 Announce Type: replace Abstract: In structurally constrained molecular optimization, state-of-the-art methods restart an expensive oracle-driven search from scratch for every new in

Amortized Multi-Objective Optimization Across Tasks with Generative Solution Modeling

Model ReleasesDGX agent

arXiv:2511.09598v5 Announce Type: replace Abstract: Many real-world applications require solving families of expensive multi-objective optimization problems~(EMOPs) under varying operational condition

An Anthropic engineer argues HTML is a better output format for AI agents than Markdown, citing information density, ease of sharing, and two-way interaction (@trq212)

Model ReleasesDGX agent

@trq212: An Anthropic engineer argues HTML is a better output format for AI agents than Markdown, citing information density, ease of sharing, and two-way interaction — Using Claude Code: The Unreason

An Embarrassingly Simple Graph Heuristic Reveals Shortcut-Solvable Benchmarks for Sequential Recommendation

Model ReleasesDGX agent

arXiv:2605.07125v1 Announce Type: cross Abstract: Sequential recommendation has increasingly shifted toward generative recommenders that combine sequential patterns with semantic item information. Yet

An Interpretable and Scalable Framework for Evaluating Large Language Models

Model ReleasesDGX agent

arXiv:2605.07046v1 Announce Type: cross Abstract: Evaluation of large language models (LLMs) is increasingly critical, yet standard benchmarking methods rely on average accuracy, overlooking both the

Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning

Model ReleasesDGX agent

arXiv:2602.19612v3 Announce Type: replace Abstract: Machine Unlearning (MU) enables Large Language Models (LLMs) to remove unsafe or outdated information. However, existing work assumes that all facts

Architecting a resilient, scalable and secure foundation for the agentic era

Model ReleasesDGX agent

Across the public sector, the conversation has shifted; we are no longer just talking about the potential of AI, we are already seeing the impact. While visionary leadership and cultural buy-in are cr

Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?

Model ReleasesDGX agent

arXiv:2605.07937v1 Announce Type: new Abstract: Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreve

Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning

Model ReleasesDGX agent

arXiv:2502.07143v3 Announce Type: replace Abstract: The severe shortage of medical doctors limits access to timely and reliable healthcare, leaving millions underserved. Large language models (LLMs) o

Attention Transfer Is Not Universally Effective for Vision Transformers

Model ReleasesDGX agent

arXiv:2605.07191v1 Announce Type: new Abstract: A recent work shows that Attention Transfer, which transfers only the attention patterns from a pre-trained teacher Vision Transformer (ViT) to a random

Attribution-Based Neuron Utility for Plasticity Restoration in Deep Networks

Model ReleasesDGX agent

arXiv:2605.06834v1 Announce Type: new Abstract: Continual learning research attempts to conserve two fundamental capabilities: new knowledge acquisition and the preservation of previously acquired kno

Automate security detection, validation, and response with Daybreak

Model ReleasesDGX agent

Daybreak is an automated security system that detects potential security threats, validates their legitimacy, and executes response actions with minimal human intervention. Developed by OpenAI, it rep

Bayesian Fine-tuning in Projected Subspaces

Model ReleasesDGX agent

arXiv:2605.07706v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning of large models by decomposing weight updates into low-rank matrices, significantly r

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight

Model ReleasesDGX agent

arXiv:2605.07021v1 Announce Type: new Abstract: Reasoning in Large Language Models (LLMs) poses a challenge for oversight as many misaligned behaviors do not surface until reasoning concludes. To addr

Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility

Model ReleasesDGX agent

arXiv:2605.06856v1 Announce Type: cross Abstract: Generative AI systems achieve impressive performance on standard benchmarks yet fail to deliver real-world utility, a disconnect we identify across 28

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs

Model ReleasesDGX agent

arXiv:2605.07731v1 Announce Type: cross Abstract: This report benchmarks the performance of ENGINEERING Ingegneria Informatica S.p.A.'s EngGPT2MoE-16B-A3B LLM, a 16B parameter Mixture of Experts (MoE)

Benchmarking Foundation Models for Renal Lesion Stratification in CT

Model ReleasesDGX agent

arXiv:2605.07749v1 Announce Type: new Abstract: The rapid proliferation of open-source medical foundation models (FMs) raises a practical question: how well do their pre-trained representations transf

Benchmarking World-Model Learning with Environment-Level Queries

Model ReleasesDGX agent

arXiv:2510.19788v4 Announce Type: replace Abstract: World models are central to building AI agents capable of flexible reasoning and planning. Yet current evaluations (i) test only properties measurab

← Previous
1…270271272273274…377
Next →