AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,611 results
9 Jun 2026

REFINE: Super-efficient 3D Gaussian Splatting Pruning via Rendering-Free Primitive Importance

Model ReleasesDGX agent

arXiv:2606.09074v1 Announce Type: new Abstract: Existing pruning methods for 3D Gaussian splatting (3DGS) suffer from either severe quality degradation or prohibitive computational overhead. In this p

Reformulate LLM Reinforcement Learning for Efficient Training under Black-box Discrepancy

Model ReleasesDGX agent

arXiv:2606.08779v1 Announce Type: new Abstract: Reinforcement Learning (RL) has emerged as a pivotal post-training paradigm, yet it frequently suffers from unpredictable sub-optimum performance or eve

Remember with Confidence: Uncertainty Quantification for Spatio-temporal Memory with Probabilistic Guarantees

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2606.08277v1 Announce Type: new Abstract: Long-horizon robot operation requires spatio-temporal memory to record the environment state and recall it for downstream reasoning. Scene graphs and re

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them

Model ReleasesDGX agent

arXiv:2606.07597v1 Announce Type: cross Abstract: Pre-training data mixtures are commonly tuned by running small-scale experiments and extrapolating to the target training budget. When high-quality da

Report: GKE Inference Gateway delivers up to 92% faster AI responses

Model ReleasesDGX agent

As generative AI moves from experimental pilots to massive production environments, the efficiency of your infrastructure becomes the ultimate differentiator. One way to get the most out of it and min

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

Model ReleasesDGX agent

arXiv:2606.07591v1 Announce Type: cross Abstract: AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We presen

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency

Model ReleasesDGX agent

arXiv:2601.06649v2 Announce Type: replace-cross Abstract: Research in machine learning has questioned whether increases in training token counts reliably produce proportional performance gains in larg

RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations

Model ReleasesDGX agent

arXiv:2606.08376v1 Announce Type: cross Abstract: As artificial intelligence (AI) systems are increasingly deployed across socially consequential domains, reports of AI-related harms and failures have

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

Model ReleasesDGX agent

arXiv:2606.08063v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in visual understanding, yet their performance degrades significantly un

Rosetta Memory: Adaptive Memory for Cross-LLM Agents

Model ReleasesDGX agent

arXiv:2606.07711v1 Announce Type: cross Abstract: Memory is the key component for transforming a stateless LLM into a persistent, evolving agent through experience accumulation, long-horizon planning,

RTL-BenchLS: A Large-Scale Benchmark for RTL Reasoning and Generation with Large Language Models

Model ReleasesDGX agent

arXiv:2606.08976v1 Announce Type: new Abstract: LLM-based RTL generation and reasoning is a promising direction for hardware design automation. High-quality benchmarks are critical infrastructure for

Rubrik turns its platform into an AI agent and ships Agent Cloud for Claude

Model ReleasesDGX agent

Rubrik Inc. today turned its data security platform into an autonomous agent and made its control layer for Anthropic PBC’s Claude generally available, the headline items in a wave of announcements at

RunAgent SuperBrowser: A Theory of Autonomous Web Navigation Grounded in Human Browsing Behaviour

Model ReleasesDGX agent

arXiv:2606.09399v1 Announce Type: new Abstract: We present SUPERBROWSER, an autonomous web-navigation agent designed against a single guiding hypothesis: a web agent should browse the way a person bro

Safe-RULE: Safe Reinforcement UnLEarning

Model ReleasesDGX agent

arXiv:2606.09559v1 Announce Type: cross Abstract: Offline safe reinforcement learning (Safe RL) enables policy learning without online interactions, making it suitable for safety-critical systems such

SafeRun: Enabling Determinism in LLM Planning for Running

Model ReleasesDGX agent

arXiv:2606.09027v1 Announce Type: cross Abstract: Large Language Models enable flexible natural-language planning but remain unreliable in determinism-critical domains due to their probabilistic natur

SC3: The Multi-Solvent Solubility Challenge and Benchmark

Model ReleasesDGX agent

arXiv:2606.07656v1 Announce Type: cross Abstract: Solubility prediction is a standard benchmark in computational chemistry, yet multi-solvent models which reportedly approach the experimental-noise ce

Scaffold Effects on GAIA: A Controlled Comparison

Model ReleasesDGX agent

arXiv:2606.08529v1 Announce Type: new Abstract: Published agent capability scores conflate what a model can do with what its scaffold lets it do, and the magnitude of this elicitation gap is not well

ScaleSweep: Accurate NVFP4 Post-Training Quantization of LLMs via Block Scale Initialization

Model ReleasesDGX agent

arXiv:2606.07618v1 Announce Type: cross Abstract: NVFP4 is a recently introduced hardware-supported FP4 format that improves the fidelity of 4-bit quantization through fine-grained block scales. Howev

Scaling by Diversified Experience for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.09009v1 Announce Type: new Abstract: Vision-Language-Action models face significant challenges in real-world deployment due to the entanglement of high-level reasoning with low-level contro

Scaling Laws for Masked-Reconstruction Transformers on Single-Cell Transcriptomics

Model ReleasesDGX agent

arXiv:2602.15253v2 Announce Type: replace Abstract: Neural scaling laws -- power-law relationships between loss, model size, and data -- have been extensively documented for language and vision transf

scCBGM: Interpretable Single-Cell Counterfactual Editing

Model ReleasesDGX agent

arXiv:2606.07760v1 Announce Type: new Abstract: Understanding cellular phenotypes and how they respond to perturbations is critical for disease biology and therapeutic design. Single-cell RNA sequenci

SceneConductor: 3D Scene Generation from Single Image with Multi-Agent Orchestration

Model ReleasesDGX agent

arXiv:2606.08402v1 Announce Type: cross Abstract: Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental context fro

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems

Model ReleasesDGX agent

arXiv:2606.08034v1 Announce Type: cross Abstract: Symbolic benchmarks have emerged as a key approach to assess model robustness under minor modifications to STEM-related questions. However, existing s

SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing

Model ReleasesDGX agent

arXiv:2602.09809v2 Announce Type: replace Abstract: Scientific diagrams convey explicit structural information, yet modern text-to-image models often produce visually plausible but structurally incorr

See how Claude Fable 5 compares across every model: http://cursor.com/evals

Model ReleasesDGX agent

Claude Fable 5 is compared against other AI models on various evaluation metrics through Cursor's benchmarking tool. The evaluation likely covers performance across different tasks such as coding, rea

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding

Model ReleasesDGX agent

arXiv:2606.09064v1 Announce Type: cross Abstract: Recent advances in Video Large Language Models (Video-LLMs) have enabled performance on long-video understanding tasks. However, existing methods stil

SegmentAnyTreeV2: Scaling Transformer-Based Tree Instance Segmentation Across Sensors, Platforms, and Forests

Model ReleasesDGX agent

arXiv:2606.08206v1 Announce Type: new Abstract: We present SegmentAnyTreeV2, a sensor- and platform-agnostic framework for semantic and instance segmentation of forest point clouds. The model combines

Semi-supervised Source Detection in Astronomical Images: New Benchmark and Strong Baseline

Model ReleasesDGX agent

arXiv:2606.09219v1 Announce Type: new Abstract: Source detection in modern observational astronomy is a cornerstone for localizing and identifying stellar sources accurately. It is crucial for studies

SENTRY: Statistical Reliability Analysis of Vision Transformers Under Soft Errors

Model ReleasesDGX agent

arXiv:2606.07620v1 Announce Type: cross Abstract: With the growth of Vision Transformers in safety-critical domains like autonomous systems and medical imaging, ensuring their reliability against soft

Seq103: A Unified Neuroevolution Framework for Compact Sequence Architecture Discovery

Model ReleasesDGX agent

arXiv:2606.07664v1 Announce Type: cross Abstract: Neuroevolution is a representative neural architecture search paradigm that evolves both network topology and weights through evolutionary algorithms.

Setting a custom price for a model in AgentsView

Model ReleasesDGX agent

TIL: Setting a custom price for a model in AgentsView I've been really enjoying AgentsView by Wes McKinney as a tool for exploring my token usage across different coding agents running on my laptop. C

Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs

Model ReleasesDGX agent

arXiv:2606.07963v1 Announce Type: new Abstract: Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific trigg

Shift-Dependent Asymmetry: Orthogonal Inverse Low-Rank Adaptation for Federated Medical Segmentation

Model ReleasesDGX agent

arXiv:2606.08687v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) enables efficient federated fine-tuning of segmentation foundation models for medical imaging. However, most federated LoRA m

Signals Are Not States: Neuro-Symbolic Safeguards for Culturally Aware Classroom AI

Model ReleasesDGX agent

arXiv:2603.22793v2 Announce Type: replace Abstract: Classroom AI systems increasingly infer high-level educational states such as engagement, confusion, collaboration, participation, and instructional

SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation

Model ReleasesDGX agent

arXiv:2606.08278v1 Announce Type: new Abstract: Humanoid foundation models are advancing faster than we can evaluate them. While real-world testing is expensive and difficult to reproduce, existing si

SLMJury: Can Small Language Models Judge as Well as Large Ones?

Model ReleasesDGX agent

arXiv:2606.07810v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used as judges for evaluating model outputs, but their high cost, latency, and opacity limit scalability. We i

SNN-MLIR: An MLIR Dialect for Compiling Neuromorphic SNNs from NIR to Bare-Metal C

Model ReleasesDGX agent

arXiv:2606.09213v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) are increasingly trained in a wide range of frameworks (SnnTorch, Lava, Norse, and others) each with its own model form

SoK: Reconstruction Attacks on Synthetic Tabular Data (Insights from Winning the NIST CRC)

Model ReleasesDGX agent

arXiv:2606.08372v1 Announce Type: cross Abstract: Synthetic data is increasingly promoted as a privacy-preserving substitute for releasing sensitive tabular records, yet its central adversarial threat

Solving Inverse Problems with Flow-based Models via Model Predictive Control

Model ReleasesDGX agent

arXiv:2601.23231v2 Announce Type: replace-cross Abstract: Flow-based generative models provide strong unconditional priors for inverse problems, but guiding their dynamics for conditional generation r

Sovereign AI for all.

Model ReleasesDGX agent

Cohere advocates for democratizing access to sovereign AI systems, enabling organizations and nations to develop and deploy their own AI models independently rather than relying on centralized provide

Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models

Model ReleasesDGX agent

arXiv:2603.19183v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have emerged as a promising approach for general-purpose robot manipulation. However, little research has mechan

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

Model ReleasesDGX agent

arXiv:2606.09669v1 Announce Type: new Abstract: Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However,

SpectrumKV: Per-Token Mixed-Precision KV Cache Transfer for Prefill-Decode Disaggregated LLM Serving

Model ReleasesDGX agent

arXiv:2606.08635v1 Announce Type: new Abstract: Prefill-decode (PD) disaggregation decouples prompt processing from token generation, but it also turns the key-value (KV) cache into a network payload.

SSAFE: Simple and Strong AI-Generated Image Detection via Frozen Vision Encoders

Model ReleasesDGX agent

arXiv:2606.08634v1 Announce Type: new Abstract: The rapid advancement of generative models has blurred the boundary between synthetic and real imagery, creating an urgent need for reliable deepfake de

Stabilizing On-Policy Distillation for MLLM Reasoning with Global Normalization

Model ReleasesDGX agent

arXiv:2606.09091v1 Announce Type: cross Abstract: On-policy distillation (OPD) has recently emerged as an important post-training paradigm. By using a stronger teacher model to provide dense, fine-gra

Still: Amortized KV Cache Compaction in a Single Forward Pass

Model ReleasesDGX agent

arXiv:2606.07878v1 Announce Type: new Abstract: The KV cache is the memory bottleneck of long-horizon language model deployment. Practically, a deployable compactor must be lightweight enough to call

Storage Insights datasets: Enabling org-wide operational discovery with activity insights

Model ReleasesDGX agent

As enterprise storage footprints scale to billions of objects, AI applications and agentic workloads are fundamentally shifting the role of storage from a passive repository to the foundation of the d

Strained Coherence: A Pre-Failure Signal in Coding Agent Execution Trajectories

Model ReleasesDGX agent

arXiv:2606.07889v1 Announce Type: cross Abstract: LLM-based coding agents sometimes acknowledge a problem in their own reasoning and then proceed anyway. We call this pattern strained coherence: a saf

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?

Model ReleasesDGX agent

arXiv:2606.09547v1 Announce Type: new Abstract: Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video

Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy

Model ReleasesDGX agent

arXiv:2606.07929v1 Announce Type: new Abstract: Large language models (LLMs) are entering clinical practice based on benchmark accuracy that may fail to detect safety-relevant failure modes. Here we p

Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking

Model ReleasesDGX agent

arXiv:2606.07689v1 Announce Type: new Abstract: Deep research agents have attracted increasing attention for their ability to collect large-scale online information to acquire target knowledge, with r

Structured Neuron Pruning in Deep Neural Networks Using Multi-Armed Bandits

Model ReleasesDGX agent

arXiv:2606.07615v1 Announce Type: cross Abstract: Deep neural networks often contain redundant hidden units. Removing individual weights can reduce parameter count, but unstructured sparsity is not al

Students without access to LLMs are 2 to 8 times more creative than students with access. That is the finding of a new paper comparing 2,200…

Model ReleasesDGX agent

Students without access to LLMs are 2 to 8 times more creative than students with access. That is the finding of a new paper comparing 2,200 college admissions essays written by humans before ChatGPT

Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER)

Model ReleasesDGX agent

arXiv:2606.07608v1 Announce Type: cross Abstract: We present a systematic study of fine-tuning OpenAI's Whisper large-v3 for Swiss German ASR, using 1,367 hours of broadcast speech paired with Standar

Supracompetitive Pricing Under AI Monoculture

Model ReleasesDGX agent

arXiv:2601.01279v3 Announce Type: replace-cross Abstract: When competing sellers delegate pricing to a shared AI model, such as a large language model, correlated recommendations combined with perform

SurfDesign: Effective Protein Design on Molecular Surfaces

Model ReleasesDGX agent

arXiv:2606.07567v1 Announce Type: cross Abstract: Protein function is largely determined by molecular surface geometry and physicochemical complementarity, yet most protein design methods condition on

SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

Model ReleasesDGX agent

arXiv:2606.07682v1 Announce Type: cross Abstract: AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex env

Systematic LLM Translation of Legacy Scientific Code to Differentiable Frameworks: Application to a Land Surface Model

Model ReleasesDGX agent

arXiv:2606.07681v1 Announce Type: cross Abstract: Differentiable programming offers transformative capabilities for scientific modeling, enabling gradient-based parameter estimation, sensitivity analy

TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs

Model ReleasesDGX agent

arXiv:2606.09578v1 Announce Type: new Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly evaluated on table reasoning tasks, but the role of table representation

TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

Model ReleasesDGX agent

arXiv:2602.03224v2 Announce Type: replace Abstract: Test-time evolution of agent memory represents a pivotal paradigm for advancing AGI, as it strengthens complex reasoning through experience accumula

← Previous
1…159160161162163…377
Next →