AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,597 results
12 May 2026

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces

Model ReleasesDGX agent

arXiv:2605.08904v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and tool use. However, the fundamental cognitive faculties essential

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning

SafetyDGX agent

arXiv:2605.06241v2 Announce Type: replace Abstract: Reinforcement learning has become the standard for improving reasoning in large language models, yet evidence increasingly suggests that RL does not

SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference

ApplicationsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.08151v1 Announce Type: cross Abstract: LLM serving platforms are increasingly deployed as multi-model cloud systems, where user demand is often long-tailed: a few popular large models recei

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

Model ReleasesDGX agent

arXiv:2605.08412v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent

Test-Time Speculation

Model ReleasesDGX agent

arXiv:2605.09329v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by using a fast draft model to generate tokens and a more accurate target model to verify them. Its perfo

The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations

Model ReleasesDGX agent

arXiv:2605.09195v1 Announce Type: new Abstract: Large language models confidently produce outdated answers, and no existing method can detect them. We show this is not an engineering failure but a str

11 May 2026

Mitigating Cognitive Bias in RLHF by Altering Rationality

Model ReleasesDGX agent

arXiv:2605.06895v1 Announce Type: new Abstract: How can we make models robust to even imperfect human feedback? In reinforcement learning from human feedback (RLHF), human preferences over model outpu

Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation

Model ReleasesDGX agent

arXiv:2605.07647v1 Announce Type: cross Abstract: Automated short answer scoring (ASAS) is shifting from discriminative, fine-tuned models to large language models (LLMs) used in few-shot settings. Th

Scaling Categorical Flow Maps

Model ReleasesDGX agent

arXiv:2605.07820v1 Announce Type: new Abstract: Continuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they u

7 May 2026

How BASF manages thousands of supply chain decisions with AlphaEvolve’s agentic algorithms

Model ReleasesDGX agent

The agricultural and crop protection supply chain is one of the most intricate networks in the world. It takes up to two years to turn active ingredients into the final products farmers need, and a si

6 May 2026

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Model ReleasesDGX agent

arXiv:2602.00933v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) is rapidly becoming the standard interface for Large Language Models (LLMs) to discover and invoke external t

5 May 2026

Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

Model ReleasesDGX agent

arXiv:2605.01630v1 Announce Type: new Abstract: Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scorin

4 May 2026

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Model ReleasesDGX agent

arXiv:2605.00674v1 Announce Type: new Abstract: Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating

Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference

Model ReleasesDGX agent

arXiv:2605.00300v1 Announce Type: cross Abstract: Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the en

1 May 2026

Low Rank Adaptation for Adversarial Perturbation

Model ReleasesDGX agent

arXiv:2604.27487v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA), which leverages the insight that model updates typically reside in a low-dimensional space, has significantly improved the t

The Impact of LLM Self-Consistency and Reasoning Effort on Automated Scoring Accuracy and Cost

Model ReleasesDGX agent

arXiv:2604.26954v1 Announce Type: cross Abstract: Strategic model selection and reasoning settings are more effective than ensembling for optimizing automated scoring with large language models (LLMs)

30 Apr 2026

Entropy Centroids as Intrinsic Rewards for Test-Time Scaling

Model ReleasesDGX agent

arXiv:2604.26173v1 Announce Type: cross Abstract: An effective way to scale up test-time compute of large language models is to sample multiple responses and then select the best one, as in Grok Heavy

MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese

Model ReleasesDGX agent

arXiv:2604.25926v1 Announce Type: new Abstract: The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and b

Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens

Model ReleasesDGX agent

arXiv:2604.26355v1 Announce Type: new Abstract: Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains unde

TildeOpen LLM: Leveraging Curriculum Learning to Achieve Equitable Language Representation

Model ReleasesDGX agent

arXiv:2603.08182v2 Announce Type: replace-cross Abstract: Large language models often underperform in many European languages due to the dominance of English and a few high-resource languages in train

TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation

Model ReleasesDGX agent

arXiv:2503.04872v3 Announce Type: replace-cross Abstract: The challenge of reducing the size of Large Language Models (LLMs) while maintaining their performance has gained significant attention. Howev

29 Apr 2026

Benchmarking and Adapting On-Device LLMs for Clinical Decision Support

Model ReleasesDGX agent

arXiv:2601.03266v2 Announce Type: replace Abstract: Large language models (LLMs) have rapidly advanced in clinical decision-making, yet the deployment of proprietary systems is hindered by privacy con

Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

SafetyDGX agent

arXiv:2604.25891v1 Announce Type: new Abstract: Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavio

CRAFT: Grounded Multi-Agent Coordination Under Partial Information

Model ReleasesDGX agent

arXiv:2603.25268v2 Announce Type: replace Abstract: We introduce CRAFT, a multi-agent benchmark for evaluating pragmatic communication in large language models under strict partial information. In thi

Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs

Local AiDGX agent

arXiv:2604.25039v1 Announce Type: new Abstract: Large Language Models (LLMs) solve many reasoning tasks via chain-of-thought (CoT) prompting, but smaller models (about 7 to 8B parameters) still strugg

Feasible-First Exploration for Constrained ML Deployment Optimization in Crash-Prone Hierarchical Search Spaces

Model ReleasesDGX agent

arXiv:2604.25073v1 Announce Type: new Abstract: Deploying machine learning models under production constraints requires joint optimization over model family, quantization scheme, runtime backend, and

Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer

Model ReleasesDGX agent

arXiv:2604.25409v1 Announce Type: new Abstract: Probabilistic Transformer (PT), a white-box probabilistic model for contextual word representation, has demonstrated substantial similarity to standard

ViPO: Visual Preference Optimization at Scale

TutorialsDGX agent

arXiv:2604.24953v1 Announce Type: new Abstract: While preference optimization is crucial for improving visual generative models, how to effectively scale this paradigm remains largely unexplored. Curr

28 Apr 2026

Domain Fine-Tuning vs. Retrieval-Augmented Generation for Medical Multiple-Choice Question Answering: A Controlled Comparison at the 4B-Parameter Scale

Model ReleasesDGX agent

arXiv:2604.23801v1 Announce Type: new Abstract: Practitioners deploying small open-weight large language models (LLMs) for medical question answering face a recurring design choice: invest in a domain

Generating Verifiable Chain of Thoughts from Exection-Traces

TutorialsDGX agent

arXiv:2512.00127v3 Announce Type: replace-cross Abstract: Getting language models to reason correctly about code requires training on data where each reasoning step can be checked. Current synthetic C

Peer Identity Bias in Multi-Agent LLM Evaluation: An Empirical Study Using the TRUST Democratic Discourse Analysis Pipeline

SafetyDGX agent

arXiv:2604.22971v1 Announce Type: cross Abstract: The TRUST democratic discourse analysis pipeline exposes its large language model (LLM) components to peer model identity through multiple structural

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns

Model ReleasesDGX agent

arXiv:2604.23150v1 Announce Type: cross Abstract: Most recent state-of-the-art (SOTA) large language models (LLMs) use Mixture-of-Experts (MoE) architectures to scale model capacity without proportion

SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation

Model ReleasesDGX agent

arXiv:2505.16637v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently demonstrated remarkable capabilities in machine translation (MT). However, most advanced MT-specifi

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

ResearchDGX agent

arXiv:2604.24763v1 Announce Type: new Abstract: Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creatin

27 Apr 2026

From Words to Amino Acids: Does the Curse of Depth Persist?

ResearchDGX agent

arXiv:2602.21750v2 Announce Type: replace Abstract: Protein language models (PLMs) have become widely adopted as general-purpose models, demonstrating strong performance in protein engineering and de

How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining

Model ReleasesDGX agent

arXiv:2511.18903v2 Announce Type: replace-cross Abstract: Due to the scarcity of high-quality data, large language models (LLMs) are often trained on mixtures of data with varying quality levels, even

Toward Automated Robustness Evaluation of Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2506.05038v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks. However, these models exhibit unexpecte

24 Apr 2026

Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs

Model ReleasesDGX agent

arXiv:2509.21361v2 Announce Type: replace-cross Abstract: Large language model (LLM) providers boast big numbers for maximum context window sizes. To test the real world use of context windows, we 1)

DeepSeek-V4: a million-token context that agents can actually use

Model ReleasesDGX agent

DeepSeek-V4 is an advanced language model featuring a million-token context window that enables practical agentic applications beyond simple retrieval. The model demonstrates improved efficiency and u

Measuring Opinion Bias and Sycophancy via LLM-based Coercion

Model ReleasesDGX agent

arXiv:2604.21564v1 Announce Type: new Abstract: Large language models increasingly shape the information people consume: they are embedded in search, consulted for professional advice, deployed as age

23 Apr 2026

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

Model ReleasesDGX agent

arXiv:2604.20051v1 Announce Type: new Abstract: Self-play has recently emerged as a promising paradigm to train Large Language Models (LLMs). In self-play, the target LLM creates the task input (e.g.,

CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values

Model ReleasesDGX agent

arXiv:2509.03740v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) like CLIP have shown impressive zero-shot and few-shot learning capabilities across diverse applications. Howeve

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2604.20140v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles with complex

LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?

Model ReleasesDGX agent

arXiv:2501.03624v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel on many NLP benchmarks, but their behavior on real-world, semi-structured prediction remains underexplored.

22 Apr 2026

New innovations in Google Distributed Cloud

Model ReleasesDGX agent

Today at Google Cloud Next, we’re announcing new capabilities in Google Distributed Cloud (GDC) that bring Gemini and our advanced AI stack to wherever your data is, so you don’t need to compromise be

ReflectMT: Internalizing Reflection for Efficient and High-Quality Machine Translation

Model ReleasesDGX agent

arXiv:2604.19144v1 Announce Type: new Abstract: Recent years have witnessed growing interest in applying Large Reasoning Models (LRMs) to Machine Translation (MT). Existing approaches predominantly ad

Right for the Wrong Reasons: Epistemic Regret Minimization for LLM Causal Reasoning

Model ReleasesDGX agent

arXiv:2602.11675v3 Announce Type: replace Abstract: Large language models may answer causal questions correctly for the wrong reasons, substituting associational shortcuts P(Y|X) for the interventiona

Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images

Model ReleasesDGX agent

arXiv:2509.07966v2 Announce Type: replace-cross Abstract: Visual reasoning over structured data such as tables is a critical capability for modern vision-language models (VLMs), yet current benchmarks

21 Apr 2026

📢 Kimi K2.6 API is live • Input Price (Cache Hit): 0.16 / M tokens • Input Price (Cache Miss): 0.95 / M tokens • Output: $4.00 / M tokens…

Model ReleasesDGX agent

📢 Kimi K2.6 API is live • Input Price (Cache Hit): 0.16 / M tokens • Input Price (Cache Miss): 0.95 / M tokens • Output: $4.00 / M tokens Kimi K2.6 is our latest + most intelligent model - stronger lo

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training

Model ReleasesDGX agent

arXiv:2510.09354v2 Announce Type: replace Abstract: Large reasoning models exhibit long chain-of-thought reasoning with complex strategies such as backtracking and self-verification. Yet, these capabi

Predicting LLM Compression Degradation from Spectral Statistics

Model ReleasesDGX agent

arXiv:2604.18085v1 Announce Type: new Abstract: Matrix-level low-rank compression is a promising way to reduce the cost of large language models, but running compression and evaluating the resulting m

PrivaDE: Privacy-preserving Data Evaluation for Blockchain-based Data Marketplaces

ResearchDGX agent

arXiv:2510.18109v4 Announce Type: replace-cross Abstract: Evaluating the usefulness of data before purchase is essential when obtaining data for high-quality machine learning models, yet both model bu

Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement

Local AiDGX agent

arXiv:2604.16858v1 Announce Type: new Abstract: Image Quality Assessment (IQA) models are increasingly deployed as perceptual critics to guide generative models and image restoration. This role demand

RefineStat: Efficient Exploration for Probabilistic Program Synthesis

ResearchDGX agent

arXiv:2509.01082v3 Announce Type: replace Abstract: Probabilistic programming offers a powerful framework for modeling uncertainty, yet statistical model discovery in this domain entails navigating an

Toward Efficient Influence Function: Dropout as a Compression Tool

ResearchDGX agent

arXiv:2509.15651v2 Announce Type: replace Abstract: Assessing the impact the training data on machine learning models is crucial for understanding the behavior of the model, enhancing the transparency

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning

Local AiDGX agent

arXiv:2601.15724v2 Announce Type: replace Abstract: Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning

When Can LLMs Learn to Reason with Weak Supervision?

TutorialsDGX agent

arXiv:2604.18574v1 Announce Type: new Abstract: Large language models have achieved significant reasoning improvements through reinforcement learning with verifiable rewards (RLVR). Yet as model capab

20 Apr 2026

(1D) Ordered Tokens Enable Efficient Test-Time Search

ResearchDGX agent

arXiv:2604.15453v1 Announce Type: cross Abstract: Tokenization is a key component of autoregressive (AR) generative models, converting raw data into more manageable units for modeling. Commonly, token

Aletheia: Gradient-Guided Layer Selection for Efficient LoRA Fine-Tuning Across Architectures

Model ReleasesDGX agent

arXiv:2604.15351v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) has become the dominant parameter-efficient fine-tuning method for large language models, yet standard practice applies LoR

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency

Model ReleasesDGX agent

arXiv:2604.16158v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex tasks. Yet ensuring that the reasoning trace both

← Previous
1…197198199200201…1010
Next →