AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,520 results
24 Jul 2026

SESaMo: Symmetry-Enforcing Stochastic Modulation for Normalizing Flows

Model ReleasesDGX agent

arXiv:2505.19619v3 Announce Type: replace Abstract: Deep generative models have recently garnered significant attention across various fields, from physics to chemistry, where sampling from unnormaliz

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Model ReleasesDGX agent

arXiv:2607.21072v1 Announce Type: new Abstract: Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks a

Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2607.09999v2 Announce Type: replace Abstract: We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved. Using a six-catego

SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents

Model ReleasesDGX agent

arXiv:2607.15557v4 Announce Type: replace Abstract: Agent skills, SKILL files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Pub

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Model ReleasesDGX agent

arXiv:2607.20548v1 Announce Type: cross Abstract: Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges hav

SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification

Model ReleasesDGX agent

arXiv:2607.20475v1 Announce Type: new Abstract: Sampling in LLM inference comprises a combinatorial set of logit processing, token selection, and verification operations for speculative decoding. Howe

Source-Prior-Driven Selective Adaptation for Efficient Diffusion Model Finetuning

Model ReleasesDGX agent

arXiv:2607.20913v1 Announce Type: new Abstract: Fine-tuning large diffusion models for new domains or styles involves a trade-off: improving target-specific generation often degrades the pretrained mo

Spectral functions in Minkowski quantum electrodynamics from neural reconstruction

Model ReleasesDGX agent

arXiv:2510.24728v2 Announce Type: replace-cross Abstract: We study neural reconstructions of quenched rainbow quantum electrodynamics (QED) Dyson--Schwinger benchmarks in Minkowski-related kinematics.

Spectral-Spatial Synergistic Guided Network for Hyperspectral Salient Object Detection

Model ReleasesDGX agent

arXiv:2607.21032v1 Announce Type: new Abstract: Hyperspectral salient object detection aims to identify visually salient regions from hyperspectral images. Existing methods often fail because they fun

StabilityBench: Benchmarking Instability in LLMs

Model ReleasesDGX agent

arXiv:2607.20558v1 Announce Type: cross Abstract: AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poor

Statistical Inference for Generative Model Comparison

Model ReleasesDGX agent

arXiv:2501.18897v4 Announce Type: replace-cross Abstract: Generative models have achieved remarkable success across a range of applications, yet their evaluation still lacks principled uncertainty qua

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs

Model ReleasesDGX agent

arXiv:2509.14257v3 Announce Type: replace-cross Abstract: Large Language Model agents achieve strong performance on multi-step reasoning and tool-use tasks, but their impressive capabilities typically

swiss-ai/Apertus-v1.5 70B/8B

Model ReleasesDGX agent

https://huggingface.co/swiss-ai/Apertus-v1.5-70B https://huggingface.co/swiss-ai/Apertus-v1.5-8B Apertus 1.5 is a family of 8B and 70B parameter language models designed to advance the state of multil

T-STAR: A Large-Scale Benchmark for Spatio-Temporal Panoptic Scene Graph Generation in Satellite Video

Model ReleasesDGX agent

arXiv:2607.21228v1 Announce Type: new Abstract: Structured understanding of satellite video is essential for advancing dynamic geospatial scene analysis from low-level perception to high-level cogniti

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

Model ReleasesDGX agent

arXiv:2607.20510v1 Announce Type: new Abstract: We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Te

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Model ReleasesDGX agent

arXiv:2607.20911v1 Announce Type: new Abstract: We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring pro

The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path

Model ReleasesDGX agent

arXiv:2607.20484v1 Announce Type: new Abstract: Large Language Models (LLMs) are fundamentally limited by representation collapse, a bottleneck that severely degrades long-context performance. We iden

The Geometry of Personality: Activation Steering with Jungian Cognitive Functions

Model ReleasesDGX agent

arXiv:2607.20803v1 Announce Type: cross Abstract: Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as

The GPT-5.x Pro series has remained the best models for hard technical problems since they launched. There is some parallel model magic goin…

Model ReleasesDGX agent

The GPT-5.x Pro series has remained the best models for hard technical problems since they launched. There is some parallel model magic going on that is not well-explained. Anthropic has never had an

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.11149v3 Announce Type: replace Abstract: LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including lo

The RealDefocus Benchmark for Defocus Deblurring

Model ReleasesDGX agent

arXiv:2607.21078v1 Announce Type: new Abstract: Single-Image Defocus Deblurring (SIDD) aims to recover an all-in-focus image from a single defocused observation, but rigorous and reproducible evaluati

The Second LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results

Model ReleasesDGX agent

arXiv:2607.21118v1 Announce Type: new Abstract: This paper presents a review of the second LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aims to advance unified image resto

The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production wit…

Model ReleasesDGX agent

The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production with predictable costs and full control over the model stack. T

This is a big jump in ARC-AGI-3.

Model ReleasesDGX agent

This is a big jump in ARC-AGI-3. Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2% The previous high score (7.8%) was set by GPT-5.6 Sol (Max) Throughout our analysis, we observed no

Three-Pronged Spectral Control for Federated Parameter Efficient Fine Tuning

Model ReleasesDGX agent

arXiv:2607.20914v1 Announce Type: new Abstract: Federated parameter-efficient fine-tuning (PEFT) enables communication-efficient adaptation of large pretrained models on decentralized edge data, but i

Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models

Model ReleasesDGX agent

arXiv:2607.21433v1 Announce Type: cross Abstract: Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate within a tok

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

Model ReleasesDGX agent

arXiv:2602.19313v2 Announce Type: replace-cross Abstract: General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, fa

torchsom: The Reference PyTorch Library for Self-Organizing Maps

Model ReleasesDGX agent

arXiv:2510.11147v2 Announce Type: replace-cross Abstract: This paper introduces torchsom, an open-source Python library that provides a reference implementation of the Self-Organizing Map (SOM) in PyT

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.21111v1 Announce Type: cross Abstract: Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected

Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2607.21496v1 Announce Type: cross Abstract: Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for enabling timely intervention and improving

Toward Mechanistic Interpretability of an AI Foundation Model Fine-Tuned for Atmospheric Chemistry

Model ReleasesDGX agent

arXiv:2607.20778v1 Announce Type: new Abstract: Weather forecasting foundation models (FMs) are increasingly fine-tuned to predict air quality, offering fast global pollution forecasts at lower comput

Towards an Automated Test of LLM Security Knowledge

Model ReleasesDGX agent

arXiv:2607.18496v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM perf

Towards Faithful Graph Explanations with Synergistic Edge Effects via Granular Balls

Model ReleasesDGX agent

arXiv:2607.21381v1 Announce Type: new Abstract: Instance-level explanations aim to reveal the rationale behind a model's decisions for a specific graph. Previous methods explain graph neural networks

Training Large Language Models for Self-Explanation Faithfulness

Model ReleasesDGX agent

arXiv:2607.21090v1 Announce Type: cross Abstract: We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated r

TransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects

Model ReleasesDGX agent

arXiv:2607.21071v1 Announce Type: new Abstract: Autonomous biomedical laboratories increasingly rely on visual perception to recognize, localize, and manipulate transparent plasticware, yet high-quali

U-CFR: Uncertainty-Guided Cascade Forward Refinement for Interactive Segmentation

Model ReleasesDGX agent

arXiv:2607.20705v1 Announce Type: cross Abstract: Interactive image segmentation is critical for efficient image annotation; however, existing methods often require many corrective clicks or rely on p

Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning

Model ReleasesDGX agent

arXiv:2607.21300v1 Announce Type: cross Abstract: Machine unlearning has emerged as a tool for removing personal data from trained models to comply with recent AI regulations. To evaluate unlearning e

Verifier-First Evaluation of Agentic LLMs for Infrastructure-as-Code Generation

Model ReleasesDGX agent

arXiv:2607.20478v1 Announce Type: cross Abstract: Infrastructure-as-Code (IaC) generation from natural language requires satisfying provider schemas, dependency planning, and organizational policy con

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?

Model ReleasesDGX agent

arXiv:2607.20868v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental ab

Visual Contrastive Self-Distillation

Model ReleasesDGX agent

arXiv:2607.21556v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymme

VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method

Model ReleasesDGX agent

arXiv:2607.21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions commonly encod

VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory

Model ReleasesDGX agent

arXiv:2603.04910v2 Announce Type: replace-cross Abstract: Imitation learning from human demonstrations has achieved significant success in robotic control, yet most visuomotor policies still condition

WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms

Model ReleasesDGX agent

arXiv:2607.20638v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in code generation and reasoning, yet their ability to perform temporal reasoning ove

We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few point…

Model ReleasesDGX agent

We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few points worse on dense tables, but does slightly better on parsing

WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance

Model ReleasesDGX agent

arXiv:2511.12997v2 Announce Type: replace Abstract: Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex browsing tas

Webly Supervised Multi-Label Recognition: Evaluation Benchmark and Dual-Branch Multi-Label Contrastive Learning

Model ReleasesDGX agent

arXiv:2607.20874v1 Announce Type: new Abstract: Training deep learning models with freely available web images can reduce their dependence on costly manual annotations. Although webly supervised learn

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay

Model ReleasesDGX agent

arXiv:2607.21005v1 Announce Type: new Abstract: Most explanations of training instability focus on learning-rate criticality, typically characterized by the Edge of Stability, beyond which optimizatio

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces

Model ReleasesDGX agent

arXiv:2607.20425v1 Announce Type: new Abstract: What makes writing 'good' remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how r

What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations

Model ReleasesDGX agent

arXiv:2607.21491v1 Announce Type: new Abstract: Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-c

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation

Model ReleasesDGX agent

arXiv:2607.21401v1 Announce Type: cross Abstract: A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard that watches that answer has to keep up w

When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers

Model ReleasesDGX agent

arXiv:2607.20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm? We answer with four findings from controlled

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion

Model ReleasesDGX agent

arXiv:2607.20543v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. We study thi

When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs

Model ReleasesDGX agent

arXiv:2607.21445v1 Announce Type: new Abstract: Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this pap

WhereEdit: Mask-aware Local Latent Editing for One-Step Image Editing

Model ReleasesDGX agent

arXiv:2607.20883v1 Announce Type: new Abstract: Recent one-step text-to-image (T2I) models enable efficient image synthesis and provide new opportunities for real-time image editing. However, existing

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning

Model ReleasesDGX agent

arXiv:2607.09328v2 Announce Type: replace-cross Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across dis

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

Model ReleasesDGX agent

arXiv:2607.21535v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models

Workflow-Localized Mechanism Learning: Attribution-Guided Repair and Knowledge Reuse for Structured Agent Skills

Model ReleasesDGX agent

arXiv:2607.20999v1 Announce Type: new Abstract: Agent Skills package reusable procedural knowledge as external artifacts for frozen language-model agents, yet existing optimizers do not jointly resolv

Zagreus-0.4B-por a small open source language model for Portuguese

Model ReleasesDGX agent

mii-llm, an open source AI lab, released Zagreus-0.4B-por, a compact bilingual Portuguese–English language model pretrained entirely from scratch. The model has approximately 400 million parameters an

ZONDA: Zero-shot Object Navigation with Dynamic Avoidance in Multi-floor Environments

Model ReleasesDGX agent

arXiv:2607.21025v1 Announce Type: new Abstract: In Object Goal Navigation task, existing methods are typically restricted to static and single-floor environments, ignoring cross-floor topologies and d

23 Jul 2026

4DGS360: 360{eg} Gaussian Reconstruction of Dynamic Objects from a Single Video

Model ReleasesDGX agent

arXiv:2603.21618v2 Announce Type: replace Abstract: We introduce 4DGS360, a diffusion-free framework for 360^{irc} dynamic object reconstruction from casual monocular video. Existing methods often fai

← Previous
1…7273747576…376
Next →