AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,597 results
9 Jun 2026

Correcting Mean Bias in Text Embeddings: A Refined Renormalization with Training-Free Improvements on MMTEB

Model ReleasesDGX agent

arXiv:2511.11041v2 Announce Type: replace-cross Abstract: We find that current sentence-embedding models produce outputs with a consistent bias: every embedding e decomposes as ilde e + mu, where the

8 Jun 2026

UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

Model ReleasesDGX agent

arXiv:2606.06622v1 Announce Type: new Abstract: We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are

6 Jun 2026

TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2606.06133v1 Announce Type: cross Abstract: TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequently produ

3 Jun 2026

Predicting Inference-Time Scaling Gains from Labeled Validation-Set Output Statistics

ResearchDGX agent

arXiv:2606.02981v1 Announce Type: new Abstract: Best-of-N inference scaling (drawing N candidate answers from a language model and returning the one a reward model ranks highest) improves accuracy by

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

Model ReleasesDGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

SenseJudge: Human-Centric Preference-Driven Judgment Framework

Model ReleasesDGX agent

arXiv:2606.03189v1 Announce Type: new Abstract: Large Language Models (LLMs) as judges across various scenarios such as assessing model responses is becoming an increasingly accepted paradigm. However

2 Jun 2026

Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction

Model ReleasesDGX agent

arXiv:2606.01850v1 Announce Type: new Abstract: Model compression techniques such as quantization and pruning are widely used to reduce the deployment cost of large language models (LLMs), with existi

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

Model ReleasesDGX agent

arXiv:2606.00616v1 Announce Type: cross Abstract: Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos. We introduce pause-

Physics-Guided Recurrent State-Space Neural Networks for Multi-Step Prediction

ResearchDGX agent

arXiv:2606.02278v1 Announce Type: cross Abstract: State-space models are traditionally based on physical knowledge, but multi-step predictions from these physical models can be poor due to model inacc

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning

Model ReleasesDGX agent

arXiv:2606.00148v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) often know the rule but pick the wrong answer: on abstract visual reasoning (AVR) tasks, a model can describe

This is actually one of the main advantages startups have over frontier labs, as long as there's a healthy spectrum of open-weight to closed…

Model ReleasesDGX agent

This is actually one of the main advantages startups have over frontier labs, as long as there's a healthy spectrum of open-weight to closed-weight models on the cost-performance curve. Building a mod

Towards Multidisciplinary Summarization of Hospital Stays: Efficient Sentence-Level Clinical Provenance Categorization

Model ReleasesDGX agent

arXiv:2606.02487v1 Announce Type: new Abstract: Effective 'all-team' summarization in high-complexity settings like the Neonatal Intensive Care Unit (NICU) requires aggregating insights from diverse d

1 Jun 2026

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs

Model ReleasesDGX agent

arXiv:2512.20732v2 Announce Type: replace-cross Abstract: As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to gene

Mellum2 Technical Report

Model ReleasesDGX agent

arXiv:2605.31268v1 Announce Type: new Abstract: We present Mellum 2, an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters per token. Mellum 2 is a general-p

Very good advice on self-improving agents. (bookmark it) This is something I am seeing in my own experiments with coding agents and harnesse…

Model ReleasesDGX agent

Very good advice on self-improving agents. (bookmark it) This is something I am seeing in my own experiments with coding agents and harnesses for long-horizon tasks. What I have found is that stronger

29 May 2026

Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations

ResearchDGX agent

arXiv:2605.29976v1 Announce Type: cross Abstract: We evaluate the climate simulation capabilities of ArchesWeather and ArchesWeatherGen, two machine learning models originally trained for weather fore

FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

Model ReleasesDGX agent

arXiv:2605.29001v1 Announce Type: cross Abstract: A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o fro

Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives

Model ReleasesDGX agent

arXiv:2505.21627v4 Announce Type: replace-cross Abstract: State-of-the-art large language models require specialized hardware and substantial energy to operate. As a consequence, cloud-based services

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

Model ReleasesDGX agent

arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak

ReasonOps: Operator Segmentation for LLM Reasoning Traces

Model ReleasesDGX agent

arXiv:2605.29192v1 Announce Type: new Abstract: Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structu

28 May 2026

Adversarial Fine-tuning of Compressed Neural Networks for Joint Improvement of Robustness and Efficiency

Model ReleasesDGX agent

arXiv:2403.09441v2 Announce Type: replace Abstract: As deep learning (DL) models are increasingly being integrated into our everyday lives, ensuring their safety by making them robust against adversar

Do We Really Need Quantum Machine Learning?: A Multidimensional Empirical Study

Model ReleasesDGX agent

arXiv:2605.27923v1 Announce Type: cross Abstract: The rapid growth of computer vision and increasingly complex image recognition tasks has exposed fundamental computational limitations of classical ma

Laguna M.1/XS.2 Technical Report

Model ReleasesDGX agent

arXiv:2605.27605v1 Announce Type: new Abstract: We present Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has 225.8B total parameters

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

Model ReleasesDGX agent

arXiv:2605.27901v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. Howeve

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness

Model ReleasesDGX agent

arXiv:2605.28190v1 Announce Type: new Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding rob

27 May 2026

HiSpec: Hierarchical Speculative Decoding for LLMs

ResearchDGX agent

arXiv:2510.01336v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates LLM inference by using a smaller draft model to speculate tokens that a larger target model verifies. Verific

26 May 2026

AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2605.25166v1 Announce Type: cross Abstract: Time series forecasting models are increasingly scaled through large Transformer backbones, yet most existing approaches process all series through a

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

Model ReleasesDGX agent

arXiv:2605.25246v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimizat

On the Sample Complexity of Robust Binary Hypothesis Testing

Model ReleasesDGX agent

arXiv:2605.24741v1 Announce Type: cross Abstract: We study the sample complexity of robust binary hypothesis testing under three standard contamination models: arepsilon-additive (Huber), arepsilon-su

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing

ResearchDGX agent

arXiv:2605.25902v1 Announce Type: new Abstract: Narrowly finetuned language models memorize implanted content verbatim, but auditing what a deployed model has been taught, without access to its weight

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

Model ReleasesDGX agent

arXiv:2605.24703v1 Announce Type: cross Abstract: Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA). Unlike text-on

25 May 2026

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

Model ReleasesDGX agent

arXiv:2605.23281v1 Announce Type: new Abstract: Monocular metric depth estimation has achieved strong progress with large-scale training and universal-camera modeling, yet robust deployment across div

23 May 2026

Physics Priors Offer Useful Accuracy-Carbon Trade-Offs in Spatio-Temporal Forecasting

ResearchDGX agent

arXiv:2509.24517v2 Announce Type: replace Abstract: Development of modern deep learning methods has been driven primarily by the push for improving model efficacy (accuracy metrics). This sole focus o

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation

Model ReleasesDGX agent

arXiv:2605.21856v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive reasoning abilities across a wide range of tasks, but data contamination undermines the object

21 May 2026

DrugRAG: Enhancing Pharmacy LLM Performance Through A Novel Retrieval-Augmented Generation Pipeline

Model ReleasesDGX agent

arXiv:2512.14896v2 Announce Type: replace Abstract: In our study, we evaluated large language model (LLM) performance on pharmacy licensure-style question-answering tasks and developed an external kno

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

Model ReleasesDGX agent

arXiv:2502.12120v3 Announce Type: replace-cross Abstract: Scaling laws guide the development of large language models (LLMs) by offering estimates for the optimal balance of model size, tokens, and co

TextSculptor: Training and Benchmarking Scene Text Editing

Model ReleasesDGX agent

arXiv:2605.21090v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editin

20 May 2026

K-Quantization and its Impact on Output Performance

Model ReleasesDGX agent

arXiv:2605.19645v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) have shown their remarkable capacities in many NLP tasks. However, their substantial size often pres

m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder

Model ReleasesDGX agent

arXiv:2605.19568v1 Announce Type: new Abstract: Embedding models are pivotal in industrial information retrieval systems like search and advertising. However, existing pretrained models often exhibit

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

Model ReleasesDGX agent

arXiv:2605.18956v1 Announce Type: new Abstract: Recent motion-language models unify tasks like comprehension and generation but operate at a coarse granularity, lacking fine-grained understanding and

19 May 2026

3DPhysVideo: Consistency-Guided Flow SDE for Video Generation via 3D Scene Reconstruction and Physical Simulation

Model ReleasesDGX agent

arXiv:2605.16795v1 Announce Type: cross Abstract: Video generative models have made remarkable progress, yet they often yield visual artifacts that violate grounding in physical dynamics. Recent works

E-PMQ: Expert-Guided Post-Merge Quantization with Merged-Weight Anchoring

ResearchDGX agent

arXiv:2605.16882v1 Announce Type: new Abstract: Low-resource deployment constraints have made model quantization essential for deploying neural networks while preserving performance. Meanwhile, model

Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-worl…

Model ReleasesDGX agent

Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-world agentic tasks), but is 5x the cost of Gemini 3 Flash @Goog

HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools

Model ReleasesDGX agent

arXiv:2605.17106v1 Announce Type: new Abstract: Production LLM deployments increasingly maintain heterogeneous model pools spanning order-of-magnitude cost differences. Existing routers make binary st

SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning

Model ReleasesDGX agent

arXiv:2605.17949v1 Announce Type: new Abstract: Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model reasoni

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

Model ReleasesDGX agent

arXiv:2601.12696v2 Announce Type: replace Abstract: Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerab

18 May 2026

3D Segmentation Using Viewpoint-Dependent Spatial Relationships

Model ReleasesDGX agent

arXiv:2605.15708v1 Announce Type: new Abstract: Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentat

Algorithmic Simplification of Neural Networks with Mosaic-of-Motifs

Model ReleasesDGX agent

arXiv:2602.14896v2 Announce Type: replace Abstract: Large-scale deep learning models are well-suited for compression. Across a variety of tasks, methods like pruning, quantization, and knowledge disti

15 May 2026

Masked Next-Scale Prediction for Self-supervised Scene Text Recognition

Model ReleasesDGX agent

arXiv:2605.14885v1 Announce Type: new Abstract: Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine-grained character strokes. Training such models relie

MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval

Model ReleasesDGX agent

arXiv:2605.06132v2 Announce Type: replace Abstract: In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the 're

PEML: Parameter-efficient Multi-Task Learning with Optimized Continuous Prompts

Model ReleasesDGX agent

arXiv:2605.14055v1 Announce Type: cross Abstract: Parameter-Efficient Fine-Tuning (PEFT) is widely used for adapting Large Language Models (LLMs) for various tasks. Recently, there has been an increas

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

Model ReleasesDGX agent

arXiv:2605.14152v1 Announce Type: cross Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual

Stateful Reasoning via Insight Replay

Model ReleasesDGX agent

arXiv:2605.14457v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has become a foundation for eliciting multi-step reasoning in large language models, but recent studies show that its b

TabPFN-3: Technical Report

Model ReleasesDGX agent

arXiv:2605.13986v1 Announce Type: new Abstract: Tabular data underpins most high-value prediction problems in science and industry, and TabPFN has driven the foundation model revolution for this modal

13 May 2026

A lot of people have been wondering about Mythos, Glasswing, and the vulns we / our partners are fixing. Today, I’m excited for us to start …

Model ReleasesDGX agent

A lot of people have been wondering about Mythos, Glasswing, and the vulns we / our partners are fixing. Today, I’m excited for us to start sharing more. (For context, I lead Glasswing @AnthropicAI.)

Beyond Parameter Aggregation: Semantic Consensus for Federated Fine-Tuning of LLMs

Model ReleasesDGX agent

arXiv:2605.11857v1 Announce Type: new Abstract: Federated fine-tuning of large language models is commonly formulated as a parameter aggregation problem. However, even parameter-efficient methods requ

12 May 2026

Collective Alignment in LLM Multi-Agent Systems: Disentangling Bias from Cooperation via Statistical Physics

Model ReleasesDGX agent

arXiv:2605.10528v1 Announce Type: cross Abstract: We investigate the emergent collective dynamics of LLM-based multi-agent systems on a 2D square lattice and present a model-agnostic statistical-physi

Hint Tuning: Less Data Makes Better Reasoners

Model ReleasesDGX agent

arXiv:2605.08665v1 Announce Type: new Abstract: Large reasoning models achieve high accuracy through extended chain-of-thought but generate 5--8 more tokens than necessary, applying verbose reasoning

LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering

Model ReleasesDGX agent

arXiv:2605.09384v1 Announce Type: cross Abstract: The reasoning gap between large and compact vision-language models (VLMs) limits the deployment of medical AI on portable clinical devices. Compact VL

Locking Pretrained Weights via Deep Low-Rank Residual Distillation

Model ReleasesDGX agent

arXiv:2605.10777v1 Announce Type: new Abstract: The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling the

← Previous
1…196197198199200…1010
Next →