AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

Content type
85,202Total entries
1Added by human
85,201Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
61,046 results
Model Releases

Ask the Right Comparison:Bias-Aware Bayesian Active Top-k Ranking with LLM Judges

DGX agent

arXiv:2607.02104v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise -- to rank responses, select models

model-releasesarxiv-cs-lg
3 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Neural Network-Based Estimation of Time-Dependent Parameters in AR(p) Processes

DGX agent

arXiv:2607.00470v1 Announce Type: cross Abstract: We investigate a forecasting framework based on a simple discrete-time dynamic model with coefficients varying in time. The parameters of the model ar

model-releasesarxiv-cs-lg
2 Jul 2026
Model Releases

Google named a Leader in 2026 Gartner® Magic Quadrant™ for Analytics and Business Intelligence Platforms for third year in a row

DGX agent

For the third consecutive year, Google has been recognized as a Leader in the 2026 Gartner® Magic Quadrant™ for Analytics and Business Intelligence Platforms. This recognition comes on the heels of Go

model-releasesgoogle-cloud-ai
1 Jul 2026
Model Releases

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

DGX agent

arXiv:2606.30059v1 Announce Type: new Abstract: Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

DGX agent

arXiv:2601.17642v2 Announce Type: replace Abstract: Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal o

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

Training-free Truthfulness Detection via Sparse MLP Value Vectors

DGX agent

arXiv:2509.17932v2 Announce Type: replace Abstract: Large language models (LLMs) are prone to generating factually incorrect content, motivating methods for assessing truthfulness from internal model

model-releasesarxiv-cs-cl
29 Jun 2026
Model Releases

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

DGX agent

arXiv:2606.26618v1 Announce Type: new Abstract: Large pretrained text-to-speech (TTS) models sound almost human for well-resourced languages, but much worse for languages that are rare in their traini

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context

DGX agent

arXiv:2606.26654v1 Announce Type: new Abstract: Personalized language-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogu

model-releasesarxiv-cs-cl
26 Jun 2026
Research

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs

DGX agent

arXiv:2511.05933v2 Announce Type: replace Abstract: Reinforcement learning (RL) is often credited with improving language model reasoning at the expense of knowledge. We challenge this narrative by sh

researcharxiv-cs-cl
25 Jun 2026
Model Releases

Not All Invariants Are Equal: Curating Training Data to Accelerate Program Verification with SLMs

DGX agent

arXiv:2603.15510v2 Announce Type: replace Abstract: The synthesis of inductive loop invariants remains a critical bottleneck in automated program verification. While Large Language Models (LLMs) show

model-releasesarxiv-cs-lg
24 Jun 2026
Model Releases

You Don't Need to Run Every Eval

DGX agent

arXiv:2606.24020v1 Announce Type: new Abstract: A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare

model-releasesarxiv-cs-lg
24 Jun 2026
Model Releases

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

DGX agent

arXiv:2606.22918v1 Announce Type: new Abstract: Maintaining physical consistency in video generators and world models increasingly relies on vision-language models (VLMs) as automated judges that prov

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Essential Subspace Merging for Multi-Task Learning

DGX agent

arXiv:2606.19164v2 Announce Type: replace Abstract: Model merging aims to enable multi-task learning by integrating the capabilities of multiple models fine-tuned from the same pre-trained checkpoint

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks

DGX agent

arXiv:2606.22935v1 Announce Type: new Abstract: Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these d

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Physiology-Aware CNN and Zero-Shot Multimodal LLMs for ECG Image Classification: A Comparative Study

DGX agent

arXiv:2606.22889v1 Announce Type: new Abstract: Multimodal large language models (LLMs) are increasingly adopted to interpret 12-lead ECG images, though the interpretations often lack validation. Howe

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology

DGX agent

arXiv:2606.12407v1 Announce Type: new Abstract: General-purpose large language models (LLMs) are routinely used as baselines when evaluating specialized pathology models on whole-slide images (WSIs).

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

DGX agent

arXiv:2606.09864v1 Announce Type: cross Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measu

model-releasesarxiv-cs-ai
10 Jun 2026
Applications

A Mechanistic Analysis of Adversarial Fine-tuning of Vision Transformers

DGX agent

arXiv:2606.07593v1 Announce Type: cross Abstract: The widespread use of image classification models in high-risk, real-world situations necessitates making these models robust to slight disturbances o

applicationsarxiv-cs-ai
9 Jun 2026
Model Releases

Correcting Mean Bias in Text Embeddings: A Refined Renormalization with Training-Free Improvements on MMTEB

DGX agent

arXiv:2511.11041v2 Announce Type: replace-cross Abstract: We find that current sentence-embedding models produce outputs with a consistent bias: every embedding e decomposes as ilde e + mu, where the

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

DGX agent

arXiv:2606.06622v1 Announce Type: new Abstract: We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation

DGX agent

arXiv:2606.06133v1 Announce Type: cross Abstract: TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequently produ

model-releasesarxiv-cs-ai
6 Jun 2026
Research

Predicting Inference-Time Scaling Gains from Labeled Validation-Set Output Statistics

DGX agent

arXiv:2606.02981v1 Announce Type: new Abstract: Best-of-N inference scaling (drawing N candidate answers from a language model and returning the one a reward model ranks highest) improves accuracy by

researcharxiv-cs-cl
3 Jun 2026
Model Releases

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

DGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

SenseJudge: Human-Centric Preference-Driven Judgment Framework

DGX agent

arXiv:2606.03189v1 Announce Type: new Abstract: Large Language Models (LLMs) as judges across various scenarios such as assessing model responses is becoming an increasingly accepted paradigm. However

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction

DGX agent

arXiv:2606.01850v1 Announce Type: new Abstract: Model compression techniques such as quantization and pruning are widely used to reduce the deployment cost of large language models (LLMs), with existi

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

DGX agent

arXiv:2606.00616v1 Announce Type: cross Abstract: Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos. We introduce pause-

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Physics-Guided Recurrent State-Space Neural Networks for Multi-Step Prediction

DGX agent

arXiv:2606.02278v1 Announce Type: cross Abstract: State-space models are traditionally based on physical knowledge, but multi-step predictions from these physical models can be poor due to model inacc

researcharxiv-cs-lg
2 Jun 2026
Model Releases

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning

DGX agent

arXiv:2606.00148v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) often know the rule but pick the wrong answer: on abstract visual reasoning (AVR) tasks, a model can describe

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

This is actually one of the main advantages startups have over frontier labs, as long as there's a healthy spectrum of open-weight to closed…

DGX agent

This is actually one of the main advantages startups have over frontier labs, as long as there's a healthy spectrum of open-weight to closed-weight models on the cost-performance curve. Building a mod

model-releasesjerry-liu--x
2 Jun 2026
Model Releases

Towards Multidisciplinary Summarization of Hospital Stays: Efficient Sentence-Level Clinical Provenance Categorization

DGX agent

arXiv:2606.02487v1 Announce Type: new Abstract: Effective 'all-team' summarization in high-complexity settings like the Neonatal Intensive Care Unit (NICU) requires aggregating insights from diverse d

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs

DGX agent

arXiv:2512.20732v2 Announce Type: replace-cross Abstract: As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to gene

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Mellum2 Technical Report

DGX agent

arXiv:2605.31268v1 Announce Type: new Abstract: We present Mellum 2, an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters per token. Mellum 2 is a general-p

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Very good advice on self-improving agents. (bookmark it) This is something I am seeing in my own experiments with coding agents and harnesse…

DGX agent

Very good advice on self-improving agents. (bookmark it) This is something I am seeing in my own experiments with coding agents and harnesses for long-horizon tasks. What I have found is that stronger

model-releasesdair-ai--x
1 Jun 2026
Research

Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations

DGX agent

arXiv:2605.29976v1 Announce Type: cross Abstract: We evaluate the climate simulation capabilities of ArchesWeather and ArchesWeatherGen, two machine learning models originally trained for weather fore

researcharxiv-cs-ai
29 May 2026
Model Releases

FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

DGX agent

arXiv:2605.29001v1 Announce Type: cross Abstract: A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o fro

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives

DGX agent

arXiv:2505.21627v4 Announce Type: replace-cross Abstract: State-of-the-art large language models require specialized hardware and substantial energy to operate. As a consequence, cloud-based services

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

DGX agent

arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

ReasonOps: Operator Segmentation for LLM Reasoning Traces

DGX agent

arXiv:2605.29192v1 Announce Type: new Abstract: Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structu

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Adversarial Fine-tuning of Compressed Neural Networks for Joint Improvement of Robustness and Efficiency

DGX agent

arXiv:2403.09441v2 Announce Type: replace Abstract: As deep learning (DL) models are increasingly being integrated into our everyday lives, ensuring their safety by making them robust against adversar

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Do We Really Need Quantum Machine Learning?: A Multidimensional Empirical Study

DGX agent

arXiv:2605.27923v1 Announce Type: cross Abstract: The rapid growth of computer vision and increasingly complex image recognition tasks has exposed fundamental computational limitations of classical ma

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Laguna M.1/XS.2 Technical Report

DGX agent

arXiv:2605.27605v1 Announce Type: new Abstract: We present Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has 225.8B total parameters

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

DGX agent

arXiv:2605.27901v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. Howeve

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness

DGX agent

arXiv:2605.28190v1 Announce Type: new Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding rob

model-releasesarxiv-cs-cl
28 May 2026
Research

HiSpec: Hierarchical Speculative Decoding for LLMs

DGX agent

arXiv:2510.01336v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates LLM inference by using a smaller draft model to speculate tokens that a larger target model verifies. Verific

researcharxiv-cs-ai
27 May 2026
Model Releases

AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting

DGX agent

arXiv:2605.25166v1 Announce Type: cross Abstract: Time series forecasting models are increasingly scaled through large Transformer backbones, yet most existing approaches process all series through a

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

DGX agent

arXiv:2605.25246v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimizat

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

On the Sample Complexity of Robust Binary Hypothesis Testing

DGX agent

arXiv:2605.24741v1 Announce Type: cross Abstract: We study the sample complexity of robust binary hypothesis testing under three standard contamination models: arepsilon-additive (Huber), arepsilon-su

model-releasesarxiv-cs-lg
26 May 2026
Research

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing

DGX agent

arXiv:2605.25902v1 Announce Type: new Abstract: Narrowly finetuned language models memorize implanted content verbatim, but auditing what a deployed model has been taught, without access to its weight

researcharxiv-cs-lg
26 May 2026
← Previous
1…247248249250251…1272
Next →