AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
49,435 results
Applications

SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference

DGX agent

arXiv:2605.08151v1 Announce Type: cross Abstract: LLM serving platforms are increasingly deployed as multi-model cloud systems, where user demand is often long-tailed: a few popular large models recei

applicationsarxiv-cs-ai
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

DGX agent

arXiv:2605.08412v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Test-Time Speculation

DGX agent

arXiv:2605.09329v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by using a fast draft model to generate tokens and a more accurate target model to verify them. Its perfo

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations

DGX agent

arXiv:2605.09195v1 Announce Type: new Abstract: Large language models confidently produce outdated answers, and no existing method can detect them. We show this is not an engineering failure but a str

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Mitigating Cognitive Bias in RLHF by Altering Rationality

DGX agent

arXiv:2605.06895v1 Announce Type: new Abstract: How can we make models robust to even imperfect human feedback? In reinforcement learning from human feedback (RLHF), human preferences over model outpu

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation

DGX agent

arXiv:2605.07647v1 Announce Type: cross Abstract: Automated short answer scoring (ASAS) is shifting from discriminative, fine-tuned models to large language models (LLMs) used in few-shot settings. Th

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Scaling Categorical Flow Maps

DGX agent

arXiv:2605.07820v1 Announce Type: new Abstract: Continuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they u

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

DGX agent

arXiv:2602.00933v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) is rapidly becoming the standard interface for Large Language Models (LLMs) to discover and invoke external t

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

DGX agent

arXiv:2605.01630v1 Announce Type: new Abstract: Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scorin

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

DGX agent

arXiv:2605.00674v1 Announce Type: new Abstract: Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference

DGX agent

arXiv:2605.00300v1 Announce Type: cross Abstract: Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the en

model-releasesarxiv-cs-lg
4 May 2026
Model Releases

Low Rank Adaptation for Adversarial Perturbation

DGX agent

arXiv:2604.27487v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA), which leverages the insight that model updates typically reside in a low-dimensional space, has significantly improved the t

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

The Impact of LLM Self-Consistency and Reasoning Effort on Automated Scoring Accuracy and Cost

DGX agent

arXiv:2604.26954v1 Announce Type: cross Abstract: Strategic model selection and reasoning settings are more effective than ensembling for optimizing automated scoring with large language models (LLMs)

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Entropy Centroids as Intrinsic Rewards for Test-Time Scaling

DGX agent

arXiv:2604.26173v1 Announce Type: cross Abstract: An effective way to scale up test-time compute of large language models is to sample multiple responses and then select the best one, as in Grok Heavy

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese

DGX agent

arXiv:2604.25926v1 Announce Type: new Abstract: The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and b

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens

DGX agent

arXiv:2604.26355v1 Announce Type: new Abstract: Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains unde

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

TildeOpen LLM: Leveraging Curriculum Learning to Achieve Equitable Language Representation

DGX agent

arXiv:2603.08182v2 Announce Type: replace-cross Abstract: Large language models often underperform in many European languages due to the dominance of English and a few high-resource languages in train

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation

DGX agent

arXiv:2503.04872v3 Announce Type: replace-cross Abstract: The challenge of reducing the size of Large Language Models (LLMs) while maintaining their performance has gained significant attention. Howev

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Benchmarking and Adapting On-Device LLMs for Clinical Decision Support

DGX agent

arXiv:2601.03266v2 Announce Type: replace Abstract: Large language models (LLMs) have rapidly advanced in clinical decision-making, yet the deployment of proprietary systems is hindered by privacy con

model-releasesarxiv-cs-cl
29 Apr 2026
Safety

Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

DGX agent

arXiv:2604.25891v1 Announce Type: new Abstract: Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavio

safetyarxiv-cs-lg
29 Apr 2026
Model Releases

CRAFT: Grounded Multi-Agent Coordination Under Partial Information

DGX agent

arXiv:2603.25268v2 Announce Type: replace Abstract: We introduce CRAFT, a multi-agent benchmark for evaluating pragmatic communication in large language models under strict partial information. In thi

model-releasesarxiv-cs-cl
29 Apr 2026
Local Ai

Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs

DGX agent

arXiv:2604.25039v1 Announce Type: new Abstract: Large Language Models (LLMs) solve many reasoning tasks via chain-of-thought (CoT) prompting, but smaller models (about 7 to 8B parameters) still strugg

local-aiarxiv-cs-cl
29 Apr 2026
Model Releases

Feasible-First Exploration for Constrained ML Deployment Optimization in Crash-Prone Hierarchical Search Spaces

DGX agent

arXiv:2604.25073v1 Announce Type: new Abstract: Deploying machine learning models under production constraints requires joint optimization over model family, quantization scheme, runtime backend, and

model-releasesarxiv-cs-lg
29 Apr 2026
Model Releases

Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer

DGX agent

arXiv:2604.25409v1 Announce Type: new Abstract: Probabilistic Transformer (PT), a white-box probabilistic model for contextual word representation, has demonstrated substantial similarity to standard

model-releasesarxiv-cs-cl
29 Apr 2026
Tutorials

ViPO: Visual Preference Optimization at Scale

DGX agent

arXiv:2604.24953v1 Announce Type: new Abstract: While preference optimization is crucial for improving visual generative models, how to effectively scale this paradigm remains largely unexplored. Curr

tutorialsarxiv-cs-cv
29 Apr 2026
Model Releases

Domain Fine-Tuning vs. Retrieval-Augmented Generation for Medical Multiple-Choice Question Answering: A Controlled Comparison at the 4B-Parameter Scale

DGX agent

arXiv:2604.23801v1 Announce Type: new Abstract: Practitioners deploying small open-weight large language models (LLMs) for medical question answering face a recurring design choice: invest in a domain

model-releasesarxiv-cs-cl
28 Apr 2026
Tutorials

Generating Verifiable Chain of Thoughts from Exection-Traces

DGX agent

arXiv:2512.00127v3 Announce Type: replace-cross Abstract: Getting language models to reason correctly about code requires training on data where each reasoning step can be checked. Current synthetic C

tutorialsarxiv-cs-ai
28 Apr 2026
Safety

Peer Identity Bias in Multi-Agent LLM Evaluation: An Empirical Study Using the TRUST Democratic Discourse Analysis Pipeline

DGX agent

arXiv:2604.22971v1 Announce Type: cross Abstract: The TRUST democratic discourse analysis pipeline exposes its large language model (LLM) components to peer model identity through multiple structural

safetyarxiv-cs-ai
28 Apr 2026
Model Releases

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns

DGX agent

arXiv:2604.23150v1 Announce Type: cross Abstract: Most recent state-of-the-art (SOTA) large language models (LLMs) use Mixture-of-Experts (MoE) architectures to scale model capacity without proportion

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation

DGX agent

arXiv:2505.16637v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently demonstrated remarkable capabilities in machine translation (MT). However, most advanced MT-specifi

model-releasesarxiv-cs-ai
28 Apr 2026
Research

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

DGX agent

arXiv:2604.24763v1 Announce Type: new Abstract: Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creatin

researcharxiv-cs-cv
28 Apr 2026
Research

From Words to Amino Acids: Does the Curse of Depth Persist?

DGX agent

arXiv:2602.21750v2 Announce Type: replace Abstract: Protein language models (PLMs) have become widely adopted as general-purpose models, demonstrating strong performance in protein engineering and de

researcharxiv-cs-lg
27 Apr 2026
Model Releases

How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining

DGX agent

arXiv:2511.18903v2 Announce Type: replace-cross Abstract: Due to the scarcity of high-quality data, large language models (LLMs) are often trained on mixtures of data with varying quality levels, even

model-releasesarxiv-cs-ai
27 Apr 2026
Model Releases

Toward Automated Robustness Evaluation of Mathematical Reasoning

DGX agent

arXiv:2506.05038v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks. However, these models exhibit unexpecte

model-releasesarxiv-cs-cl
27 Apr 2026
Model Releases

Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs

DGX agent

arXiv:2509.21361v2 Announce Type: replace-cross Abstract: Large language model (LLM) providers boast big numbers for maximum context window sizes. To test the real world use of context windows, we 1)

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Measuring Opinion Bias and Sycophancy via LLM-based Coercion

DGX agent

arXiv:2604.21564v1 Announce Type: new Abstract: Large language models increasingly shape the information people consume: they are embedded in search, consulted for professional advice, deployed as age

model-releasesarxiv-cs-cl
24 Apr 2026
Model Releases

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

DGX agent

arXiv:2604.20051v1 Announce Type: new Abstract: Self-play has recently emerged as a promising paradigm to train Large Language Models (LLMs). In self-play, the target LLM creates the task input (e.g.,

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values

DGX agent

arXiv:2509.03740v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) like CLIP have shown impressive zero-shot and few-shot learning capabilities across diverse applications. Howeve

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs

DGX agent

arXiv:2604.20140v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles with complex

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?

DGX agent

arXiv:2501.03624v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel on many NLP benchmarks, but their behavior on real-world, semi-structured prediction remains underexplored.

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

ReflectMT: Internalizing Reflection for Efficient and High-Quality Machine Translation

DGX agent

arXiv:2604.19144v1 Announce Type: new Abstract: Recent years have witnessed growing interest in applying Large Reasoning Models (LRMs) to Machine Translation (MT). Existing approaches predominantly ad

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images

DGX agent

arXiv:2509.07966v2 Announce Type: replace-cross Abstract: Visual reasoning over structured data such as tables is a critical capability for modern vision-language models (VLMs), yet current benchmarks

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training

DGX agent

arXiv:2510.09354v2 Announce Type: replace Abstract: Large reasoning models exhibit long chain-of-thought reasoning with complex strategies such as backtracking and self-verification. Yet, these capabi

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Predicting LLM Compression Degradation from Spectral Statistics

DGX agent

arXiv:2604.18085v1 Announce Type: new Abstract: Matrix-level low-rank compression is a promising way to reduce the cost of large language models, but running compression and evaluating the resulting m

model-releasesarxiv-cs-lg
21 Apr 2026
Research

PrivaDE: Privacy-preserving Data Evaluation for Blockchain-based Data Marketplaces

DGX agent

arXiv:2510.18109v4 Announce Type: replace-cross Abstract: Evaluating the usefulness of data before purchase is essential when obtaining data for high-quality machine learning models, yet both model bu

researcharxiv-cs-lg
21 Apr 2026
Local Ai

Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement

DGX agent

arXiv:2604.16858v1 Announce Type: new Abstract: Image Quality Assessment (IQA) models are increasingly deployed as perceptual critics to guide generative models and image restoration. This role demand

local-aiarxiv-cs-cv
21 Apr 2026
Research

RefineStat: Efficient Exploration for Probabilistic Program Synthesis

DGX agent

arXiv:2509.01082v3 Announce Type: replace Abstract: Probabilistic programming offers a powerful framework for modeling uncertainty, yet statistical model discovery in this domain entails navigating an

researcharxiv-cs-lg
21 Apr 2026
Research

Toward Efficient Influence Function: Dropout as a Compression Tool

DGX agent

arXiv:2509.15651v2 Announce Type: replace Abstract: Assessing the impact the training data on machine learning models is crucial for understanding the behavior of the model, enhancing the transparency

researcharxiv-cs-lg
21 Apr 2026
← Previous
1…199200201202203…1030
Next →