AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlog
86,428Total entries
1Added by human
86,427Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,016 results
Applications

A Mechanistic Analysis of Adversarial Fine-tuning of Vision Transformers

DGX agent

arXiv:2606.07593v1 Announce Type: cross Abstract: The widespread use of image classification models in high-risk, real-world situations necessitates making these models robust to slight disturbances o

applicationsarxiv-cs-ai
9 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Correcting Mean Bias in Text Embeddings: A Refined Renormalization with Training-Free Improvements on MMTEB

DGX agent

arXiv:2511.11041v2 Announce Type: replace-cross Abstract: We find that current sentence-embedding models produce outputs with a consistent bias: every embedding e decomposes as ilde e + mu, where the

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

DGX agent

arXiv:2606.06622v1 Announce Type: new Abstract: We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation

DGX agent

arXiv:2606.06133v1 Announce Type: cross Abstract: TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequently produ

model-releasesarxiv-cs-ai
6 Jun 2026
Research

Predicting Inference-Time Scaling Gains from Labeled Validation-Set Output Statistics

DGX agent

arXiv:2606.02981v1 Announce Type: new Abstract: Best-of-N inference scaling (drawing N candidate answers from a language model and returning the one a reward model ranks highest) improves accuracy by

researcharxiv-cs-cl
3 Jun 2026
Model Releases

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

DGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

SenseJudge: Human-Centric Preference-Driven Judgment Framework

DGX agent

arXiv:2606.03189v1 Announce Type: new Abstract: Large Language Models (LLMs) as judges across various scenarios such as assessing model responses is becoming an increasingly accepted paradigm. However

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction

DGX agent

arXiv:2606.01850v1 Announce Type: new Abstract: Model compression techniques such as quantization and pruning are widely used to reduce the deployment cost of large language models (LLMs), with existi

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

DGX agent

arXiv:2606.00616v1 Announce Type: cross Abstract: Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos. We introduce pause-

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Physics-Guided Recurrent State-Space Neural Networks for Multi-Step Prediction

DGX agent

arXiv:2606.02278v1 Announce Type: cross Abstract: State-space models are traditionally based on physical knowledge, but multi-step predictions from these physical models can be poor due to model inacc

researcharxiv-cs-lg
2 Jun 2026
Model Releases

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning

DGX agent

arXiv:2606.00148v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) often know the rule but pick the wrong answer: on abstract visual reasoning (AVR) tasks, a model can describe

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

This is actually one of the main advantages startups have over frontier labs, as long as there's a healthy spectrum of open-weight to closed…

DGX agent

This is actually one of the main advantages startups have over frontier labs, as long as there's a healthy spectrum of open-weight to closed-weight models on the cost-performance curve. Building a mod

model-releasesjerry-liu--x
2 Jun 2026
Model Releases

Towards Multidisciplinary Summarization of Hospital Stays: Efficient Sentence-Level Clinical Provenance Categorization

DGX agent

arXiv:2606.02487v1 Announce Type: new Abstract: Effective 'all-team' summarization in high-complexity settings like the Neonatal Intensive Care Unit (NICU) requires aggregating insights from diverse d

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs

DGX agent

arXiv:2512.20732v2 Announce Type: replace-cross Abstract: As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to gene

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Mellum2 Technical Report

DGX agent

arXiv:2605.31268v1 Announce Type: new Abstract: We present Mellum 2, an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters per token. Mellum 2 is a general-p

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Very good advice on self-improving agents. (bookmark it) This is something I am seeing in my own experiments with coding agents and harnesse…

DGX agent

Very good advice on self-improving agents. (bookmark it) This is something I am seeing in my own experiments with coding agents and harnesses for long-horizon tasks. What I have found is that stronger

model-releasesdair-ai--x
1 Jun 2026
Research

Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations

DGX agent

arXiv:2605.29976v1 Announce Type: cross Abstract: We evaluate the climate simulation capabilities of ArchesWeather and ArchesWeatherGen, two machine learning models originally trained for weather fore

researcharxiv-cs-ai
29 May 2026
Model Releases

FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

DGX agent

arXiv:2605.29001v1 Announce Type: cross Abstract: A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o fro

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives

DGX agent

arXiv:2505.21627v4 Announce Type: replace-cross Abstract: State-of-the-art large language models require specialized hardware and substantial energy to operate. As a consequence, cloud-based services

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

DGX agent

arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

ReasonOps: Operator Segmentation for LLM Reasoning Traces

DGX agent

arXiv:2605.29192v1 Announce Type: new Abstract: Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structu

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Adversarial Fine-tuning of Compressed Neural Networks for Joint Improvement of Robustness and Efficiency

DGX agent

arXiv:2403.09441v2 Announce Type: replace Abstract: As deep learning (DL) models are increasingly being integrated into our everyday lives, ensuring their safety by making them robust against adversar

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Do We Really Need Quantum Machine Learning?: A Multidimensional Empirical Study

DGX agent

arXiv:2605.27923v1 Announce Type: cross Abstract: The rapid growth of computer vision and increasingly complex image recognition tasks has exposed fundamental computational limitations of classical ma

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Laguna M.1/XS.2 Technical Report

DGX agent

arXiv:2605.27605v1 Announce Type: new Abstract: We present Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has 225.8B total parameters

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

DGX agent

arXiv:2605.27901v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. Howeve

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness

DGX agent

arXiv:2605.28190v1 Announce Type: new Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding rob

model-releasesarxiv-cs-cl
28 May 2026
Research

HiSpec: Hierarchical Speculative Decoding for LLMs

DGX agent

arXiv:2510.01336v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates LLM inference by using a smaller draft model to speculate tokens that a larger target model verifies. Verific

researcharxiv-cs-ai
27 May 2026
Model Releases

AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting

DGX agent

arXiv:2605.25166v1 Announce Type: cross Abstract: Time series forecasting models are increasingly scaled through large Transformer backbones, yet most existing approaches process all series through a

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

DGX agent

arXiv:2605.25246v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimizat

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

On the Sample Complexity of Robust Binary Hypothesis Testing

DGX agent

arXiv:2605.24741v1 Announce Type: cross Abstract: We study the sample complexity of robust binary hypothesis testing under three standard contamination models: arepsilon-additive (Huber), arepsilon-su

model-releasesarxiv-cs-lg
26 May 2026
Research

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing

DGX agent

arXiv:2605.25902v1 Announce Type: new Abstract: Narrowly finetuned language models memorize implanted content verbatim, but auditing what a deployed model has been taught, without access to its weight

researcharxiv-cs-lg
26 May 2026
Model Releases

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

DGX agent

arXiv:2605.24703v1 Announce Type: cross Abstract: Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA). Unlike text-on

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

DGX agent

arXiv:2605.23281v1 Announce Type: new Abstract: Monocular metric depth estimation has achieved strong progress with large-scale training and universal-camera modeling, yet robust deployment across div

model-releasesarxiv-cs-cv
25 May 2026
Research

Physics Priors Offer Useful Accuracy-Carbon Trade-Offs in Spatio-Temporal Forecasting

DGX agent

arXiv:2509.24517v2 Announce Type: replace Abstract: Development of modern deep learning methods has been driven primarily by the push for improving model efficacy (accuracy metrics). This sole focus o

researcharxiv-cs-lg
23 May 2026
Model Releases

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation

DGX agent

arXiv:2605.21856v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive reasoning abilities across a wide range of tasks, but data contamination undermines the object

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

DrugRAG: Enhancing Pharmacy LLM Performance Through A Novel Retrieval-Augmented Generation Pipeline

DGX agent

arXiv:2512.14896v2 Announce Type: replace Abstract: In our study, we evaluated large language model (LLM) performance on pharmacy licensure-style question-answering tasks and developed an external kno

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

DGX agent

arXiv:2502.12120v3 Announce Type: replace-cross Abstract: Scaling laws guide the development of large language models (LLMs) by offering estimates for the optimal balance of model size, tokens, and co

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

TextSculptor: Training and Benchmarking Scene Text Editing

DGX agent

arXiv:2605.21090v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editin

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

K-Quantization and its Impact on Output Performance

DGX agent

arXiv:2605.19645v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) have shown their remarkable capacities in many NLP tasks. However, their substantial size often pres

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder

DGX agent

arXiv:2605.19568v1 Announce Type: new Abstract: Embedding models are pivotal in industrial information retrieval systems like search and advertising. However, existing pretrained models often exhibit

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

DGX agent

arXiv:2605.18956v1 Announce Type: new Abstract: Recent motion-language models unify tasks like comprehension and generation but operate at a coarse granularity, lacking fine-grained understanding and

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

3DPhysVideo: Consistency-Guided Flow SDE for Video Generation via 3D Scene Reconstruction and Physical Simulation

DGX agent

arXiv:2605.16795v1 Announce Type: cross Abstract: Video generative models have made remarkable progress, yet they often yield visual artifacts that violate grounding in physical dynamics. Recent works

model-releasesarxiv-cs-ai
19 May 2026
Research

E-PMQ: Expert-Guided Post-Merge Quantization with Merged-Weight Anchoring

DGX agent

arXiv:2605.16882v1 Announce Type: new Abstract: Low-resource deployment constraints have made model quantization essential for deploying neural networks while preserving performance. Meanwhile, model

researcharxiv-cs-cl
19 May 2026
Model Releases

Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-worl…

DGX agent

Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-world agentic tasks), but is 5x the cost of Gemini 3 Flash @Goog

model-releasesjeremy-howard--x
19 May 2026
Model Releases

HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools

DGX agent

arXiv:2605.17106v1 Announce Type: new Abstract: Production LLM deployments increasingly maintain heterogeneous model pools spanning order-of-magnitude cost differences. Existing routers make binary st

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning

DGX agent

arXiv:2605.17949v1 Announce Type: new Abstract: Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model reasoni

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

DGX agent

arXiv:2601.12696v2 Announce Type: replace Abstract: Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerab

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

3D Segmentation Using Viewpoint-Dependent Spatial Relationships

DGX agent

arXiv:2605.15708v1 Announce Type: new Abstract: Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentat

model-releasesarxiv-cs-cv
18 May 2026
← Previous
1…251252253254255…1292
Next →