AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
All
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,628 results
Model Releases

MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science

DGX agent

arXiv:2510.12171v2 Announce Type: replace Abstract: Large Language Models have shown strong scientific reasoning ability, but their performance on materials science problems remains less studied. To f

model-releasesarxiv-cs-ai
9 Jun 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance

DGX agent

arXiv:2605.22664v2 Announce Type: replace Abstract: LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet ente

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

MedVision: Benchmarking Quantitative Medical Image Analysis

DGX agent

arXiv:2511.18676v2 Announce Type: replace-cross Abstract: Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., 'Is this normal or abnormal

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Memory Beyond Recall: A Dual-Process Cognitive Memory System for Self-Evolving LLM Agents

DGX agent

arXiv:2606.09483v1 Announce Type: cross Abstract: Long-term memory for an LLM agent is more than retrieving the right passage at the right time. Current memory systems collapse belief revision, causal

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering

DGX agent

arXiv:2601.22859v3 Announce Type: replace-cross Abstract: The evolution of Large Language Model (LLM) agents for software engineering (SWE) is constrained by the scarcity of verifiable datasets, a bot

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Microsoft AI head calls out Anthropic for acting like Claude is conscious

DGX agent

Microsoft AI CEO Mustafa Suleyman says it's 'really, really dangerous' for Anthropic to speculate about Claude's consciousness inside its 'constitution,' or the instructions that tell the model how to

model-releasesthe-verge-ai
9 Jun 2026
Model Releases

Minibatch Selection via Partition Matroid Constrained Gradient Matching

DGX agent

arXiv:2606.07954v1 Announce Type: cross Abstract: Training large language models (LLMs) on heterogeneous data requires selecting minibatches that balance convergence speed with coverage across domains

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

MLingualFC: Evaluating Jailbreak Vulnerabilities in Multilingual Vision-Language Models

DGX agent

arXiv:2606.07706v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated strong performance across multimodal tasks, yet their safety robustness remains an open challenge. Whi

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting

DGX agent

arXiv:2601.09085v2 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) has become a standard approach for training mathematical reasoning models; however, its reliance on

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Model-Based Learning of Whittle indices

DGX agent

arXiv:2511.20397v2 Announce Type: replace Abstract: We present BLINQ, a new model-based algorithm that learns the Whittle indices of an indexable, communicating and unichain Markov Decision Process (M

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Devices

DGX agent

arXiv:2606.07857v1 Announce Type: cross Abstract: The rise of edge-based machine learning has enabled distributed adaptation of language models across mobile and IoT devices, offering privacy preserva

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

MOLOT System Card: Malicious Operational Logic Observation Transformer

DGX agent

arXiv:2606.07792v1 Announce Type: cross Abstract: MOLOT (Malicious Operational Logic Observation Transformer) is a static malicious-code detection system designed for SAST setup where package metadata

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Multi-Armed Bandits with Arriving Arms: Sequential Screening, Dynamic Regret, and Sublinear Guarantees

DGX agent

arXiv:2606.09002v1 Announce Type: cross Abstract: We study a stochastic multi-armed bandit problem in which the set of available arms expands over time. This setting arises in sequential experimentati

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Multimodal Large Language Models as Synthetic Participants in Video-Based Studies: An Evaluation

DGX agent

arXiv:2606.07541v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong performance on objective tasks such as video understanding and reasoning. However, it remai

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

mythos will be bad ON PURPOSE on ai 'frontier llm research' tasks, this is very very sad for the research community also the fact that this …

DGX agent

mythos will be bad ON PURPOSE on ai 'frontier llm research' tasks, this is very very sad for the research community also the fact that this is un purpose not visible to the user is crazy Introducing C

model-releasesjeremy-howard--x
9 Jun 2026
Model Releases

NEW: Anthropic introduces Claude Fable 5, a Mythos-class model for general use. Beginning of a new class of frontier models.

DGX agent

NEW: Anthropic introduces Claude Fable 5, a Mythos-class model for general use. Beginning of a new class of frontier models. Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for g

model-releasesdair-ai--x
9 Jun 2026
Model Releases

NGram-MoSE: Efficient Remote Sensing Super-Resolution via N-Gram Context and Mixture-of-Experts

DGX agent

arXiv:2606.08535v1 Announce Type: new Abstract: Remote sensing applications for environmental monitoring and disaster management are frequently constrained by a spatial--temporal trade-off: imagery wi

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs

DGX agent

arXiv:2606.09411v1 Announce Type: cross Abstract: Large language models can be fine-tuned to encode prompt-borne secrets into fluent, seemingly benign outputs. This creates a steganographic exfiltrati

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

NutriMLLM: Multimodal Large Language Models for Dietary Micronutrient Analysis

DGX agent

arXiv:2606.08948v1 Announce Type: cross Abstract: Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but training such models requires large mul

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Offline Reinforcement Learning for Plasma Control in Nuclear Fusion: Codebase and Benchmark

DGX agent

arXiv:2606.07550v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) offers a promising route for developing plasma controllers from historical tokamak data, since online trial-and-er

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

DGX agent

arXiv:2606.08572v1 Announce Type: new Abstract: While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability t

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

OmniFaceRig: Fully Automatic Inner-Mouth-Aware Face Rigging Across Diverse 3D Character Topologies

DGX agent

arXiv:2606.08043v1 Announce Type: cross Abstract: Facial rigging - creating FACS-based blendshapes together with inner-mouth geometry (teeth, gums, and tongue) - remains a major bottleneck in 3D chara

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics

DGX agent

arXiv:2606.09826v1 Announce Type: cross Abstract: Vision-language model (VLM) agents are increasingly deployed in interactive game environments. Yet game benchmarks for VLM agents typically report a s

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

OmniGen-AR: AutoRegressive Any-to-Image Generation

DGX agent

arXiv:2606.09156v1 Announce Type: new Abstract: Autoregressive (AR) models have demonstrated strong potential in visual generation, offering superior performance with simple architectures and optimiza

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs

DGX agent

arXiv:2606.07577v1 Announce Type: new Abstract: Audio-visual large language models (LLMs) hold strong promise for long-form video understanding, yet their long-video inference is fundamentally limited

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

OmniTryOn: Video Try-On Anything at Once!

DGX agent

arXiv:2606.08514v1 Announce Type: new Abstract: Although video virtual try-on (VVT) has achieved significant progress, existing methods still exhibit two fundamental limitations: first, they are restr

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

On Choosing the mu Parameter in Gaussian Differential Privacy

DGX agent

arXiv:2606.09582v1 Announce Type: new Abstract: Recent work argues for using Gaussian differential privacy (GDP) to report the privacy guarantees in privacy-preserving machine learning. We provide pri

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Online Learning with Recency: Algorithms for Sliding-window Streaming Multi-armed Bandits

DGX agent

arXiv:2606.08977v1 Announce Type: new Abstract: Motivated by the recency effect in online learning, we study algorithms for single-pass *sliding-window streaming multi-armed bandits (MABs)* in this pa

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

OpenAI is hosed. There is almost no reason to prefer them over Anthropic, they have lost their lead (despite every advantage in the world), …

DGX agent

OpenAI is hosed. There is almost no reason to prefer them over Anthropic, they have lost their lead (despite every advantage in the world), they have made commitments far far beyond their means, their

model-releasesgary-marcus--x
9 Jun 2026
Model Releases

Operator learning for the 2D incompressible Navier-Stokes equations: a conformal prediction approach in the data-scarce regime

DGX agent

arXiv:2606.08654v1 Announce Type: new Abstract: In this paper, we propose a perturbation-based conformal prediction framework for uncertainty quantification in operator learning, with a focus on the 2

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

OptMuon: Closed-Loop Orthogonalized Momentum Methods for Stochastic Optimization with Zero-Noise Optimality

DGX agent

arXiv:2606.08783v1 Announce Type: cross Abstract: Orthogonalized momentum updates, as used in Muon-style optimizers, have recently shown strong empirical stability in large-scale deep learning. Howeve

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

PACT: Learning Diverse Diagnostic Strategies via Privileged Synthesis and Branch Consensus

DGX agent

arXiv:2606.08938v1 Announce Type: cross Abstract: Clinical diagnosis requires flexible use of multiple reasoning paradigms under incomplete patient information. Existing LLM-based medical agents show

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Parameter Tuning with Generalization Guarantees for GPU-Accelerated Linear Programming

DGX agent

arXiv:2606.08638v1 Announce Type: cross Abstract: Recent research has developed practical, parallelizable first-order methods for large scale linear programming, but performance is highly dependent on

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation

DGX agent

arXiv:2510.20182v2 Announce Type: replace Abstract: Pedestrian simulation traditionally relies on expert-tuned, hand-crafted models that limit scalability and generalization. Meanwhile, large-scale vi

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

PereStruct: Multimodal Semantic Assembly for Robust Historical Document Parsing

DGX agent

arXiv:2606.07661v1 Announce Type: new Abstract: Parsing historical documents with complex, non-standard layouts remains a fundamental bottleneck in large-scale archival digitization. Unlike modern typ

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

DGX agent

arXiv:2606.09038v1 Announce Type: new Abstract: Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. H

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Phantom transitions in language model fine-tuning

DGX agent

arXiv:2606.07559v1 Announce Type: cross Abstract: Fine-tuning a language model on contexts whose correct completion has a near-synonym competitor often fails silently. The cross-entropy loss decreases

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Pharmacogenomic Knowledge Graph Augmentation for Graph Neural Network-Based Drug-Drug Interaction Prediction

DGX agent

arXiv:2606.07698v1 Announce Type: cross Abstract: Graph neural networks (GNNs) applied to drug-drug interaction (DDI) prediction rely exclusively on molecular structure encoded as SMILES-derived graph

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Phase transition in large language models and the criticality of natural languages

DGX agent

arXiv:2406.05335v3 Announce Type: replace-cross Abstract: Generation of text and speech in natural languages can be modeled as a stochastic process. This idea dates back to the seminal work of Markov

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

phepy: Visual benchmarks and improvements for out-of-distribution detectors

DGX agent

arXiv:2503.05169v2 Announce Type: replace Abstract: Applying machine learning to increasingly high-dimensional problems with sparse or biased training data increases the risk that a model is used on i

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher Systems

DGX agent

arXiv:2606.08481v1 Announce Type: cross Abstract: Enterprise property graphs vary widely in schema structure, internal terminology, domain assumptions, governance constraints, and user interaction pat

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits

DGX agent

arXiv:2510.17947v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are improving at an exceptional rate. With the advent of agentic workflows, multi-turn dialogue has become the de

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation

DGX agent

arXiv:2603.05500v2 Announce Type: replace-cross Abstract: Efficient and stable training of large language models (LLMs) remains a core challenge in modern machine learning systems. To address this cha

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

POISE: Position-Aware Undetectable Skill Injection on LLM Agents

DGX agent

arXiv:2606.07943v1 Announce Type: cross Abstract: Agent skills provide a lightweight mechanism for extending general-purpose agents, but their open format exposes them to skill-poisoning attacks. A pr

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction

DGX agent

arXiv:2606.09788v1 Announce Type: new Abstract: Large-scale document processing requires contextually aware table extraction (TE) that is both accurate and efficient. Yet current approaches require bi

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Powering the future of robotics in Europe

DGX agent

Google DeepMind is launching a three-month accelerator program for early-stage robotics startups across Europe, designed to support the next generation of physical AI. Selected startups receive hands-

model-releasesgoogle-deepmind
9 Jun 2026
Model Releases

Pre-Intervention Prediction of Sparse Autoencoder Steering Side Effects

DGX agent

arXiv:2606.08365v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are increasingly used to steer language models, but feature steering is rarely clean: the same intervention can beha

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Prescriptive Scaling Reveals the Evolution of Language Model Capabilities

DGX agent

arXiv:2602.15327v2 Announce Type: replace-cross Abstract: Machine learning model performance improvements tend to arise from competition and application. For deployment, we consider prescriptive scali

model-releasesarxiv-cs-ai
9 Jun 2026
← Previous
1…198199200201202…472
Next →