AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,569 results
12 May 2026

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare

Model ReleasesDGX agent

arXiv:2605.08445v1 Announce Type: new Abstract: AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard t

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

Model ReleasesDGX agent

arXiv:2507.23511v3 Announce Type: replace-cross Abstract: While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. Th

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.10002v1 Announce Type: new Abstract: Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet inco

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

Model ReleasesDGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies

Model ReleasesDGX agent

arXiv:2605.09661v1 Announce Type: cross Abstract: Large language models (LLMs) have saturated standard medical benchmarks that test factual recall, yet their ability to perform higher-order reasoning,

Meet physics-intern🧑‍🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA o…

Model ReleasesDGX agent

Meet physics-intern🧑‍🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA on one of the hardest benchmarks for LLMs. Theoretical physics

MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents

Model ReleasesDGX agent

arXiv:2605.09530v1 Announce Type: cross Abstract: As LLM-powered agents are increasingly deployed in edge-cloud environments, personalized memory has become a key enabler of long-term adaptation and u

MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs

Model ReleasesDGX agent

arXiv:2605.08374v1 Announce Type: new Abstract: Episodic memory allows LLM agents to accumulate and retrieve experience, but current methods treat each memory independently, i.e., evaluating retrieval

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology

Model ReleasesDGX agent

arXiv:2605.09152v1 Announce Type: new Abstract: Deciphering animal intent is a fundamental challenge in computational ethology, largely because of semantic aliasing, the phenomenon where identical ext

MESD: A Risk-Sensitive Metric for Explanation Fairness Across Intersectional Subgroups

Model ReleasesDGX agent

arXiv:2603.13452v2 Announce Type: replace Abstract: Fairness in machine learning is predominantly evaluated through outcome-oriented metrics, such as Demographic parity, which measure whether predicti

Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon

Model ReleasesDGX agent

arXiv:2605.09708v1 Announce Type: cross Abstract: We present Metal-Sci, a 10-task benchmark of scientific Apple Silicon Metal compute kernels spanning six optimization regimes (stencils, all-pairs in

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

Model ReleasesDGX agent

arXiv:2605.10067v1 Announce Type: cross Abstract: Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing ap

MicroFuse: Protein-to-Genome Expert Fusion for Microbial Operon Reasoning

Model ReleasesDGX agent

arXiv:2605.08815v1 Announce Type: new Abstract: Predicting microbial operon co-membership requires integrating two complementary biological signals: protein-scale molecular identity and genome-context

MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph

Model ReleasesDGX agent

arXiv:2605.10120v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) show remarkable potential for scientific reasoning, yet their performance in specialized domains such as micr

MIDUS: Memory-Infused Depth Up-Scaling

Model ReleasesDGX agent

arXiv:2512.13751v2 Announce Type: replace-cross Abstract: Expanding pre-trained language models offers a practical way to increase capacity without training larger models from scratch. Depth Up-Scalin

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

Model ReleasesDGX agent

arXiv:2510.09592v2 Announce Type: replace Abstract: Real-time Spoken Language Models (SLMs) struggle to leverage Chain-of-Thought (CoT) reasoning due to the prohibitive latency of generating the entir

Minimizing Worst-Case Weighted Latency for Multi-Robot Persistent Monitoring: Theory and RL-Based Solutions

Model ReleasesDGX agent

arXiv:2605.09633v1 Announce Type: new Abstract: We study multi-robot persistent monitoring on weighted graphs, where node weights encode monitoring priorities and edge weights encode travel distances.

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

Model ReleasesDGX agent

arXiv:2605.08816v1 Announce Type: new Abstract: In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogou

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

Model ReleasesDGX agent

arXiv:2605.08678v1 Announce Type: new Abstract: Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonst

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection

Model ReleasesDGX agent

arXiv:2605.10833v1 Announce Type: cross Abstract: Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which

Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds

Model ReleasesDGX agent

arXiv:2605.09724v1 Announce Type: new Abstract: Existing accounts of grokking explain the phenomena in terms of mechanistic frameworks such as circuit efficiency or lazy-to-rich transitions. However,

Model-Free Neural Filtering: A Comparison with Classical Filters in Nonlinear Systems

Model ReleasesDGX agent

arXiv:2601.21266v3 Announce Type: replace Abstract: Neural network models are increasingly used for state estimation in control and decision-making, yet it remains unclear to what extent they behave a

Models got an order of magnitude better at following instructions in one year

Model ReleasesDGX agent

A year ago, frontier models started losing track of instructions somewhere around 200–300 simultaneous constraints. With 2026 models, that ceiling is closer to 2,000 — an order-of-magnitude jump. We r

MolRGen: A Training and Evaluation Setting for De Novo Molecular Generation with Reasonning Models

Model ReleasesDGX agent

arXiv:2603.18256v2 Announce Type: replace-cross Abstract: Recent reasoning-based large language models have shown strong performance on tasks with verifiable outcomes, but their use in de novo molecul

MolSight: Molecular Property Prediction with Images

Model ReleasesDGX agent

arXiv:2605.10157v1 Announce Type: cross Abstract: Every molecule ever synthesised can be drawn as a 2D skeletal diagram, yet in modern property prediction this universally available representation has

MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring

Model ReleasesDGX agent

arXiv:2605.09684v1 Announce Type: cross Abstract: We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-eli

More cardputer code incoming. Thank you Claude.

Model ReleasesDGX agent

Boris Cherny announced upcoming releases of code related to 'cardputer,' likely a computing or card-based project, and expressed gratitude to Claude (presumably the AI assistant). The post suggests ne

MOTOR-Bench: A Real-world Dataset and Multi-agent Framework for Zero-shot Human Mental State Understanding

Model ReleasesDGX agent

arXiv:2605.09703v1 Announce Type: new Abstract: Understanding human mental states from natural behavior is crucial for intelligent systems in the real world. However, most current research focuses on

MPerS: Dynamic MLLM MixExperts Perception-Guided Remote Sensing Scene Segmentation

Model ReleasesDGX agent

arXiv:2605.10769v1 Announce Type: cross Abstract: The multimodal fusion of images and scene captions has been extensively explored and applied in various fields. However, when dealing with complex rem

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

Model ReleasesDGX agent

arXiv:2605.10616v1 Announce Type: cross Abstract: Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generaliza

Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy

Model ReleasesDGX agent

arXiv:2605.10550v1 Announce Type: new Abstract: Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -

Multi-Tier Labeling and Physics-Informed Learning for Orbital Anomaly Detection at Scale

Model ReleasesDGX agent

arXiv:2605.09790v1 Announce Type: cross Abstract: Detecting orbital anomalies, such as maneuvers, atmospheric decay, and attitude upsets, across the rapidly growing population of low-Earth-orbit (LEO)

MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing

Model ReleasesDGX agent

arXiv:2605.08163v1 Announce Type: cross Abstract: Text-in-image editing has become a key capability for visual content creation, yet existing benchmarks remain overwhelmingly English-centric and often

Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning

Model ReleasesDGX agent

arXiv:2605.08949v1 Announce Type: new Abstract: A central challenge in continual learning for large language models (LLMs) is catastrophic forgetting, where adapting to new tasks can substantially deg

NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation

Model ReleasesDGX agent

arXiv:2605.10813v1 Announce Type: new Abstract: LLM-powered multi-agent systems can now automate the full research pipeline from ideation to paper writing, but a fundamental question remains: automati

NARRA-Gym for Evaluating Interactive Narrative Agents

Model ReleasesDGX agent

arXiv:2605.08503v1 Announce Type: new Abstract: Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmark

Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents

Model ReleasesDGX agent

arXiv:2605.09863v1 Announce Type: cross Abstract: Production LLM coding agents drift over long sessions: they forget user-specified constraints, slip into mistakes the user already flagged, and confab

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks

Model ReleasesDGX agent

arXiv:2605.10639v1 Announce Type: new Abstract: The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluati

$NBIS announced a partnership with LangChain to integrate Nebius Token Factory with LangChain’s Deep Agents. The goal is to make it easier f…

Model ReleasesDGX agent

$NBIS announced a partnership with LangChain to integrate Nebius Token Factory with LangChain’s Deep Agents. The goal is to make it easier for teams building AI agents on LangChain to run those worklo

Need document parsing that stays fully local and private? 👀 Meet liteparse-server, a self-hostable, open-source HTTP server for parsing doc…

Model ReleasesDGX agent

Need document parsing that stays fully local and private? 👀 Meet liteparse-server, a self-hostable, open-source HTTP server for parsing documents and generating screenshots from PDFs, Office files, an

Nested Slice Sampling: Vectorized Nested Sampling for GPU-Accelerated Inference

Model ReleasesDGX agent

arXiv:2601.23252v2 Announce Type: replace-cross Abstract: Model comparison and calibrated uncertainty quantification often require integrating over parameters, but scalable inference can be challengin

Neural Cluster First, Route Second: One-Shot Capacitated Vehicle Routing via Differentiable Optimal Transport

Model ReleasesDGX agent

arXiv:2605.09301v1 Announce Type: cross Abstract: The Capacitated Vehicle Routing Problem (CVRP) underpins modern last-mile logistics. Current Neural Combinatorial Optimization (NCO) methods construct

Neural Information Causality

Model ReleasesDGX agent

arXiv:2605.09316v1 Announce Type: cross Abstract: Query-separated computation forces a representation to play an operational role: data are encoded before a query is known, and a later decoder can ans

Neural Posterior Estimation of Terrain Parameters from Radar Sounder Data

Model ReleasesDGX agent

arXiv:2605.08179v1 Announce Type: cross Abstract: Radar sounders are electromagnetic instruments that can probe deep into the subsurface of Earth and other planetary bodies by processing the echo of t

Neural Weight Norm = Kolmogorov Complexity

Model ReleasesDGX agent

arXiv:2605.10878v1 Announce Type: new Abstract: Why does weight decay work? We prove that, in any fixed-precision regime, the smallest weight norm of a looped neural network outputting a binary string

NeuralBench: A Unifying Framework to Benchmark NeuroAI Models

Model ReleasesDGX agent

arXiv:2605.08495v1 Announce Type: new Abstract: Deep learning and large public datasets have recently catalyzed the proliferation of AI models for processing brain recordings. However, systematically

NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims

Model ReleasesDGX agent

arXiv:2605.08192v1 Announce Type: cross Abstract: Frontier AI safety claims - published assertions that a highly capable general-purpose model is below a threshold of concern, adequately mitigated, or

New Signadot skill lets Claude Code, Codex and Cursor validate changes in live Kubernetes environments

Model ReleasesDGX agent

Microservices testing company Signadot Inc. today launched /signadot-validate, a new skill that lets coding agents such as Anthropic PBC’s Claude Code, OpenAI Group PBC’s Codex and Cursor validate the

Normalization Equivariance for Arbitrary Backbones, with Application to Image Denoising

Model ReleasesDGX agent

arXiv:2605.08193v1 Announce Type: cross Abstract: Normalization Equivariance (NE), equivariance to global contrast and brightness transforms, improves robustness to distribution shift in image-to-imag

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness

Model ReleasesDGX agent

arXiv:2605.10379v1 Announce Type: new Abstract: Large language models (LLMs) have become capable mathematical problem-solvers, often producing correct proofs for challenging problems. However, correct

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

Model ReleasesDGX agent

arXiv:2605.08762v1 Announce Type: cross Abstract: Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

Model ReleasesDGX agent

arXiv:2605.09996v1 Announce Type: new Abstract: While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision-language, wit

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.09727v1 Announce Type: cross Abstract: A central challenge in reinforcement learning (RL) is to learn models that generalize beyond the tasks on which they are trained, a goal traditionally

OpenAI contacted me to say “Study Mode is still live and accessible via /study and /learn shortcuts” so that’s good, although the official s…

Model ReleasesDGX agent

OpenAI contacted me to say “Study Mode is still live and accessible via /study and /learn shortcuts” so that’s good, although the official study mode page doesn’t mention that. (I don’t think slash co

OpenSGA: Efficient 3D Scene Graph Alignment in the Open World

Model ReleasesDGX agent

arXiv:2605.10484v1 Announce Type: new Abstract: Scene graph alignment establishes object correspondences between two 3D scene graphs constructed from partially overlapping observations. This enables e

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces

Model ReleasesDGX agent

arXiv:2605.08904v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and tool use. However, the fundamental cognitive faculties essential

Optimal FALQON for Quantum Approximate Optimization via Layer-wise Parameter Tuning

Model ReleasesDGX agent

arXiv:2605.08332v1 Announce Type: cross Abstract: Feedback-based adaptive quantum optimization (FALQON) is a promising approach for solving combinatorial problems on noisy intermediate-scale quantum (

Optimality of Sub-network Laplace Approximations: New Results and Methods

Model ReleasesDGX agent

arXiv:2605.09075v1 Announce Type: cross Abstract: Although the Laplace approximation offers a simple route to uncertainty quantification in deep neural networks, its reliance on inverting large Hessia

Optimised Support Vector Regression for California Housing Price Prediction: The Critical Role of Feature Engineering and Hyperparameter Tuning

Model ReleasesDGX agent

arXiv:2605.08660v1 Announce Type: new Abstract: In the recent literature, Support Vector Regression (SVR) has been cited as one of the weakest performers on the California Housing benchmark dataset, w

Optimized Culprit Identification Using Mobilenet and Attention Mechanisms

Model ReleasesDGX agent

arXiv:2605.08169v1 Announce Type: cross Abstract: Automated culprit identification in surveillance systems is a critical task that requires high accuracy along with computational efficiency for real-t

← Previous
1…265266267268269…377
Next →