AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
49,435 results
Model Releases

Towards Multidisciplinary Summarization of Hospital Stays: Efficient Sentence-Level Clinical Provenance Categorization

DGX agent

arXiv:2606.02487v1 Announce Type: new Abstract: Effective 'all-team' summarization in high-complexity settings like the Neonatal Intensive Care Unit (NICU) requires aggregating insights from diverse d

model-releasesarxiv-cs-cl
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs

DGX agent

arXiv:2512.20732v2 Announce Type: replace-cross Abstract: As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to gene

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Mellum2 Technical Report

DGX agent

arXiv:2605.31268v1 Announce Type: new Abstract: We present Mellum 2, an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters per token. Mellum 2 is a general-p

model-releasesarxiv-cs-cl
1 Jun 2026
Research

Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations

DGX agent

arXiv:2605.29976v1 Announce Type: cross Abstract: We evaluate the climate simulation capabilities of ArchesWeather and ArchesWeatherGen, two machine learning models originally trained for weather fore

researcharxiv-cs-ai
29 May 2026
Model Releases

FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

DGX agent

arXiv:2605.29001v1 Announce Type: cross Abstract: A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o fro

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives

DGX agent

arXiv:2505.21627v4 Announce Type: replace-cross Abstract: State-of-the-art large language models require specialized hardware and substantial energy to operate. As a consequence, cloud-based services

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

DGX agent

arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

ReasonOps: Operator Segmentation for LLM Reasoning Traces

DGX agent

arXiv:2605.29192v1 Announce Type: new Abstract: Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structu

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Adversarial Fine-tuning of Compressed Neural Networks for Joint Improvement of Robustness and Efficiency

DGX agent

arXiv:2403.09441v2 Announce Type: replace Abstract: As deep learning (DL) models are increasingly being integrated into our everyday lives, ensuring their safety by making them robust against adversar

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Do We Really Need Quantum Machine Learning?: A Multidimensional Empirical Study

DGX agent

arXiv:2605.27923v1 Announce Type: cross Abstract: The rapid growth of computer vision and increasingly complex image recognition tasks has exposed fundamental computational limitations of classical ma

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Laguna M.1/XS.2 Technical Report

DGX agent

arXiv:2605.27605v1 Announce Type: new Abstract: We present Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has 225.8B total parameters

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

DGX agent

arXiv:2605.27901v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. Howeve

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness

DGX agent

arXiv:2605.28190v1 Announce Type: new Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding rob

model-releasesarxiv-cs-cl
28 May 2026
Research

HiSpec: Hierarchical Speculative Decoding for LLMs

DGX agent

arXiv:2510.01336v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates LLM inference by using a smaller draft model to speculate tokens that a larger target model verifies. Verific

researcharxiv-cs-ai
27 May 2026
Model Releases

AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting

DGX agent

arXiv:2605.25166v1 Announce Type: cross Abstract: Time series forecasting models are increasingly scaled through large Transformer backbones, yet most existing approaches process all series through a

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

DGX agent

arXiv:2605.25246v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimizat

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

On the Sample Complexity of Robust Binary Hypothesis Testing

DGX agent

arXiv:2605.24741v1 Announce Type: cross Abstract: We study the sample complexity of robust binary hypothesis testing under three standard contamination models: arepsilon-additive (Huber), arepsilon-su

model-releasesarxiv-cs-lg
26 May 2026
Research

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing

DGX agent

arXiv:2605.25902v1 Announce Type: new Abstract: Narrowly finetuned language models memorize implanted content verbatim, but auditing what a deployed model has been taught, without access to its weight

researcharxiv-cs-lg
26 May 2026
Model Releases

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

DGX agent

arXiv:2605.24703v1 Announce Type: cross Abstract: Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA). Unlike text-on

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

DGX agent

arXiv:2605.23281v1 Announce Type: new Abstract: Monocular metric depth estimation has achieved strong progress with large-scale training and universal-camera modeling, yet robust deployment across div

model-releasesarxiv-cs-cv
25 May 2026
Research

Physics Priors Offer Useful Accuracy-Carbon Trade-Offs in Spatio-Temporal Forecasting

DGX agent

arXiv:2509.24517v2 Announce Type: replace Abstract: Development of modern deep learning methods has been driven primarily by the push for improving model efficacy (accuracy metrics). This sole focus o

researcharxiv-cs-lg
23 May 2026
Model Releases

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation

DGX agent

arXiv:2605.21856v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive reasoning abilities across a wide range of tasks, but data contamination undermines the object

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

DrugRAG: Enhancing Pharmacy LLM Performance Through A Novel Retrieval-Augmented Generation Pipeline

DGX agent

arXiv:2512.14896v2 Announce Type: replace Abstract: In our study, we evaluated large language model (LLM) performance on pharmacy licensure-style question-answering tasks and developed an external kno

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

DGX agent

arXiv:2502.12120v3 Announce Type: replace-cross Abstract: Scaling laws guide the development of large language models (LLMs) by offering estimates for the optimal balance of model size, tokens, and co

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

TextSculptor: Training and Benchmarking Scene Text Editing

DGX agent

arXiv:2605.21090v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editin

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

K-Quantization and its Impact on Output Performance

DGX agent

arXiv:2605.19645v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) have shown their remarkable capacities in many NLP tasks. However, their substantial size often pres

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder

DGX agent

arXiv:2605.19568v1 Announce Type: new Abstract: Embedding models are pivotal in industrial information retrieval systems like search and advertising. However, existing pretrained models often exhibit

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

DGX agent

arXiv:2605.18956v1 Announce Type: new Abstract: Recent motion-language models unify tasks like comprehension and generation but operate at a coarse granularity, lacking fine-grained understanding and

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

3DPhysVideo: Consistency-Guided Flow SDE for Video Generation via 3D Scene Reconstruction and Physical Simulation

DGX agent

arXiv:2605.16795v1 Announce Type: cross Abstract: Video generative models have made remarkable progress, yet they often yield visual artifacts that violate grounding in physical dynamics. Recent works

model-releasesarxiv-cs-ai
19 May 2026
Research

E-PMQ: Expert-Guided Post-Merge Quantization with Merged-Weight Anchoring

DGX agent

arXiv:2605.16882v1 Announce Type: new Abstract: Low-resource deployment constraints have made model quantization essential for deploying neural networks while preserving performance. Meanwhile, model

researcharxiv-cs-cl
19 May 2026
Model Releases

HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools

DGX agent

arXiv:2605.17106v1 Announce Type: new Abstract: Production LLM deployments increasingly maintain heterogeneous model pools spanning order-of-magnitude cost differences. Existing routers make binary st

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning

DGX agent

arXiv:2605.17949v1 Announce Type: new Abstract: Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model reasoni

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

DGX agent

arXiv:2601.12696v2 Announce Type: replace Abstract: Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerab

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

3D Segmentation Using Viewpoint-Dependent Spatial Relationships

DGX agent

arXiv:2605.15708v1 Announce Type: new Abstract: Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentat

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

Algorithmic Simplification of Neural Networks with Mosaic-of-Motifs

DGX agent

arXiv:2602.14896v2 Announce Type: replace Abstract: Large-scale deep learning models are well-suited for compression. Across a variety of tasks, methods like pruning, quantization, and knowledge disti

model-releasesarxiv-cs-lg
18 May 2026
Model Releases

Masked Next-Scale Prediction for Self-supervised Scene Text Recognition

DGX agent

arXiv:2605.14885v1 Announce Type: new Abstract: Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine-grained character strokes. Training such models relie

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval

DGX agent

arXiv:2605.06132v2 Announce Type: replace Abstract: In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the 're

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

PEML: Parameter-efficient Multi-Task Learning with Optimized Continuous Prompts

DGX agent

arXiv:2605.14055v1 Announce Type: cross Abstract: Parameter-Efficient Fine-Tuning (PEFT) is widely used for adapting Large Language Models (LLMs) for various tasks. Recently, there has been an increas

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

DGX agent

arXiv:2605.14152v1 Announce Type: cross Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Stateful Reasoning via Insight Replay

DGX agent

arXiv:2605.14457v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has become a foundation for eliciting multi-step reasoning in large language models, but recent studies show that its b

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

TabPFN-3: Technical Report

DGX agent

arXiv:2605.13986v1 Announce Type: new Abstract: Tabular data underpins most high-value prediction problems in science and industry, and TabPFN has driven the foundation model revolution for this modal

model-releasesarxiv-cs-lg
15 May 2026
Model Releases

Beyond Parameter Aggregation: Semantic Consensus for Federated Fine-Tuning of LLMs

DGX agent

arXiv:2605.11857v1 Announce Type: new Abstract: Federated fine-tuning of large language models is commonly formulated as a parameter aggregation problem. However, even parameter-efficient methods requ

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Collective Alignment in LLM Multi-Agent Systems: Disentangling Bias from Cooperation via Statistical Physics

DGX agent

arXiv:2605.10528v1 Announce Type: cross Abstract: We investigate the emergent collective dynamics of LLM-based multi-agent systems on a 2D square lattice and present a model-agnostic statistical-physi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Hint Tuning: Less Data Makes Better Reasoners

DGX agent

arXiv:2605.08665v1 Announce Type: new Abstract: Large reasoning models achieve high accuracy through extended chain-of-thought but generate 5--8 more tokens than necessary, applying verbose reasoning

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering

DGX agent

arXiv:2605.09384v1 Announce Type: cross Abstract: The reasoning gap between large and compact vision-language models (VLMs) limits the deployment of medical AI on portable clinical devices. Compact VL

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Locking Pretrained Weights via Deep Low-Rank Residual Distillation

DGX agent

arXiv:2605.10777v1 Announce Type: new Abstract: The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling the

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces

DGX agent

arXiv:2605.08904v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and tool use. However, the fundamental cognitive faculties essential

model-releasesarxiv-cs-ai
12 May 2026
Safety

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning

DGX agent

arXiv:2605.06241v2 Announce Type: replace Abstract: Reinforcement learning has become the standard for improving reasoning in large language models, yet evidence increasingly suggests that RL does not

safetyarxiv-cs-cl
12 May 2026
← Previous
1…198199200201202…1030
Next →