AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
86,428Total entries
1Added by human
86,427Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,015 results
Model Releases

How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining

DGX agent

arXiv:2511.18903v2 Announce Type: replace-cross Abstract: Due to the scarcity of high-quality data, large language models (LLMs) are often trained on mixtures of data with varying quality levels, even

model-releasesarxiv-cs-ai
27 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Toward Automated Robustness Evaluation of Mathematical Reasoning

DGX agent

arXiv:2506.05038v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks. However, these models exhibit unexpecte

model-releasesarxiv-cs-cl
27 Apr 2026
Model Releases

Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs

DGX agent

arXiv:2509.21361v2 Announce Type: replace-cross Abstract: Large language model (LLM) providers boast big numbers for maximum context window sizes. To test the real world use of context windows, we 1)

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

DeepSeek-V4: a million-token context that agents can actually use

DGX agent

DeepSeek-V4 is an advanced language model featuring a million-token context window that enables practical agentic applications beyond simple retrieval. The model demonstrates improved efficiency and u

model-releaseshugging-face
24 Apr 2026
Model Releases

Measuring Opinion Bias and Sycophancy via LLM-based Coercion

DGX agent

arXiv:2604.21564v1 Announce Type: new Abstract: Large language models increasingly shape the information people consume: they are embedded in search, consulted for professional advice, deployed as age

model-releasesarxiv-cs-cl
24 Apr 2026
Model Releases

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

DGX agent

arXiv:2604.20051v1 Announce Type: new Abstract: Self-play has recently emerged as a promising paradigm to train Large Language Models (LLMs). In self-play, the target LLM creates the task input (e.g.,

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values

DGX agent

arXiv:2509.03740v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) like CLIP have shown impressive zero-shot and few-shot learning capabilities across diverse applications. Howeve

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs

DGX agent

arXiv:2604.20140v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles with complex

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?

DGX agent

arXiv:2501.03624v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel on many NLP benchmarks, but their behavior on real-world, semi-structured prediction remains underexplored.

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

New innovations in Google Distributed Cloud

DGX agent

Today at Google Cloud Next, we’re announcing new capabilities in Google Distributed Cloud (GDC) that bring Gemini and our advanced AI stack to wherever your data is, so you don’t need to compromise be

model-releasesgoogle-cloud-ai
22 Apr 2026
Model Releases

ReflectMT: Internalizing Reflection for Efficient and High-Quality Machine Translation

DGX agent

arXiv:2604.19144v1 Announce Type: new Abstract: Recent years have witnessed growing interest in applying Large Reasoning Models (LRMs) to Machine Translation (MT). Existing approaches predominantly ad

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Right for the Wrong Reasons: Epistemic Regret Minimization for LLM Causal Reasoning

DGX agent

arXiv:2602.11675v3 Announce Type: replace Abstract: Large language models may answer causal questions correctly for the wrong reasons, substituting associational shortcuts P(Y|X) for the interventiona

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images

DGX agent

arXiv:2509.07966v2 Announce Type: replace-cross Abstract: Visual reasoning over structured data such as tables is a critical capability for modern vision-language models (VLMs), yet current benchmarks

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

📢 Kimi K2.6 API is live • Input Price (Cache Hit): 0.16 / M tokens • Input Price (Cache Miss): 0.95 / M tokens • Output: $4.00 / M tokens…

DGX agent

📢 Kimi K2.6 API is live • Input Price (Cache Hit): 0.16 / M tokens • Input Price (Cache Miss): 0.95 / M tokens • Output: $4.00 / M tokens Kimi K2.6 is our latest + most intelligent model - stronger lo

model-releaseskimi-moonshot--x
21 Apr 2026
Model Releases

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training

DGX agent

arXiv:2510.09354v2 Announce Type: replace Abstract: Large reasoning models exhibit long chain-of-thought reasoning with complex strategies such as backtracking and self-verification. Yet, these capabi

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Predicting LLM Compression Degradation from Spectral Statistics

DGX agent

arXiv:2604.18085v1 Announce Type: new Abstract: Matrix-level low-rank compression is a promising way to reduce the cost of large language models, but running compression and evaluating the resulting m

model-releasesarxiv-cs-lg
21 Apr 2026
Research

PrivaDE: Privacy-preserving Data Evaluation for Blockchain-based Data Marketplaces

DGX agent

arXiv:2510.18109v4 Announce Type: replace-cross Abstract: Evaluating the usefulness of data before purchase is essential when obtaining data for high-quality machine learning models, yet both model bu

researcharxiv-cs-lg
21 Apr 2026
Local Ai

Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement

DGX agent

arXiv:2604.16858v1 Announce Type: new Abstract: Image Quality Assessment (IQA) models are increasingly deployed as perceptual critics to guide generative models and image restoration. This role demand

local-aiarxiv-cs-cv
21 Apr 2026
Research

RefineStat: Efficient Exploration for Probabilistic Program Synthesis

DGX agent

arXiv:2509.01082v3 Announce Type: replace Abstract: Probabilistic programming offers a powerful framework for modeling uncertainty, yet statistical model discovery in this domain entails navigating an

researcharxiv-cs-lg
21 Apr 2026
Research

Toward Efficient Influence Function: Dropout as a Compression Tool

DGX agent

arXiv:2509.15651v2 Announce Type: replace Abstract: Assessing the impact the training data on machine learning models is crucial for understanding the behavior of the model, enhancing the transparency

researcharxiv-cs-lg
21 Apr 2026
Local Ai

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning

DGX agent

arXiv:2601.15724v2 Announce Type: replace Abstract: Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning

local-aiarxiv-cs-cv
21 Apr 2026
Tutorials

When Can LLMs Learn to Reason with Weak Supervision?

DGX agent

arXiv:2604.18574v1 Announce Type: new Abstract: Large language models have achieved significant reasoning improvements through reinforcement learning with verifiable rewards (RLVR). Yet as model capab

tutorialsarxiv-cs-lg
21 Apr 2026
Research

(1D) Ordered Tokens Enable Efficient Test-Time Search

DGX agent

arXiv:2604.15453v1 Announce Type: cross Abstract: Tokenization is a key component of autoregressive (AR) generative models, converting raw data into more manageable units for modeling. Commonly, token

researcharxiv-cs-ai
20 Apr 2026
Model Releases

Aletheia: Gradient-Guided Layer Selection for Efficient LoRA Fine-Tuning Across Architectures

DGX agent

arXiv:2604.15351v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) has become the dominant parameter-efficient fine-tuning method for large language models, yet standard practice applies LoR

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency

DGX agent

arXiv:2604.16158v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex tasks. Yet ensuring that the reasoning trace both

model-releasesarxiv-cs-ai
20 Apr 2026
Research

Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms

DGX agent

arXiv:2604.15842v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive capabilities, yet their internal mechanisms for handling reasoning-intensive tasks remain unde

researcharxiv-cs-cl
20 Apr 2026
Model Releases

MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition

DGX agent

arXiv:2604.16009v1 Announce Type: new Abstract: Metacognition, the ability to monitor and regulate one's own reasoning, remains under-evaluated in AI benchmarking. We introduce MEDLEY-BENCH, a benchma

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals

DGX agent

arXiv:2604.15859v1 Announce Type: cross Abstract: Forecasting has become a natural benchmark for reasoning under uncertainty. Yet existing evaluations of large language models remain limited to judgme

model-releasesarxiv-cs-ai
20 Apr 2026
Research

Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes

DGX agent

arXiv:2604.14914v1 Announce Type: new Abstract: Text-driven inversion of generative models is a core paradigm for manipulating 2D or 3D content, unlocking numerous applications such as text-based edit

researcharxiv-cs-cv
17 Apr 2026
Model Releases

Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude. Powered by Claude Opus 4.7, our m…

DGX agent

Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude. Powered by Claude Opus 4.7, our most capable vision model. Available in research preview on t

model-releasesthariq--x
17 Apr 2026
Research

C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions

DGX agent

arXiv:2604.13521v1 Announce Type: new Abstract: Neural network models with latent recurrent processing, where identical layers are recursively applied to the latent state, have gained attention as pro

researcharxiv-cs-lg
16 Apr 2026
Applications

Diagnostics for Individual-Level Prediction Instability in Machine Learning for Healthcare

DGX agent

arXiv:2603.00192v2 Announce Type: replace Abstract: In healthcare, predictive models increasingly inform patient-level decisions, yet little attention is paid to the variability in individual risk est

applicationsarxiv-cs-lg
16 Apr 2026
Research

Structure- and Stability-Preserving Learning of Port-Hamiltonian Systems

DGX agent

arXiv:2604.13297v1 Announce Type: cross Abstract: This paper investigates the problem of data-driven modeling of port-Hamiltonian systems while preserving their intrinsic Hamiltonian structure and sta

researcharxiv-cs-lg
16 Apr 2026
Research

Accelerating Speculative Decoding with Block Diffusion Draft Trees

DGX agent

arXiv:2604.12989v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive language models by using a lightweight drafter to propose multiple future tokens, which the target model

researcharxiv-cs-cl
15 Apr 2026
Model Releases

Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss

DGX agent

arXiv:2604.12911v1 Announce Type: cross Abstract: Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to p

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents

DGX agent

arXiv:2510.10073v2 Announce Type: replace-cross Abstract: Large vision-language model (LVLM)-based web agents are emerging as powerful tools for automating complex online tasks. However, when deployed

model-releasesarxiv-cs-cv
15 Apr 2026
Research

A Mathematical Explanation of Transformers

DGX agent

arXiv:2510.03989v2 Announce Type: replace-cross Abstract: The Transformer architecture has revolutionized the field of sequence modeling and underpins the recent breakthroughs in large language models

researcharxiv-cs-ai
14 Apr 2026
Model Releases

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

DGX agent

arXiv:2601.11044v3 Announce Type: replace Abstract: Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. Howev

model-releasesarxiv-cs-ai
14 Apr 2026
Tutorials

Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation

DGX agent

arXiv:2510.10925v2 Announce Type: replace-cross Abstract: Training student models on synthetic data generated by strong teacher models is a promising way to distilling the capabilities of teachers. Ho

tutorialsarxiv-cs-cl
14 Apr 2026
Model Releases

GIANTS: Generative Insight Anticipation from Scientific Literature

DGX agent

arXiv:2604.09793v1 Announce Type: cross Abstract: Scientific breakthroughs often emerge from synthesizing prior ideas into novel contributions. While language models (LMs) show promise in scientific d

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Incentivizing Honesty among Competitors in Collaborative Learning and Optimization

DGX agent

arXiv:2305.16272v5 Announce Type: replace Abstract: Collaborative learning techniques have the potential to enable training machine learning models that are superior to models trained on a single enti

model-releasesarxiv-cs-lg
14 Apr 2026
Model Releases

Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

DGX agent

arXiv:2604.11666v1 Announce Type: cross Abstract: As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dial

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text

DGX agent

arXiv:2601.17172v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly capable of generating personalized, persuasive text at scale, raising new questions about bias a

model-releasesarxiv-cs-ai
14 Apr 2026
Applications

Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs

DGX agent

arXiv:2604.10495v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed in real-world applications, reliable uncertainty quantification (UQ) becomes critical for safe

applicationsarxiv-cs-cl
14 Apr 2026
Model Releases

A Compact Hybrid Convolution--Frequency State Space Network for Learned Image Compression

DGX agent

arXiv:2511.20151v2 Announce Type: replace Abstract: Learned image compression (LIC) has recently benefited from Transformer- and state space models (SSM)- based backbones for modeling long-range depen

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

Adaptive Simulation Experiment for LLM Policy Optimization

DGX agent

arXiv:2604.08779v1 Announce Type: new Abstract: Large language models (LLMs) have significant potential to improve operational efficiency in operations management. Deploying these models requires spec

model-releasesarxiv-cs-lg
13 Apr 2026
Applications

AI Driven Soccer Analysis Using Computer Vision

DGX agent

arXiv:2604.08722v1 Announce Type: cross Abstract: Sport analysis is crucial for team performance since it provides actionable data that can inform coaching decisions, improve player performance, and e

applicationsarxiv-cs-ai
13 Apr 2026
Research

BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation

DGX agent

arXiv:2604.09497v1 Announce Type: cross Abstract: Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse use cases.

researcharxiv-cs-ai
13 Apr 2026
← Previous
1…253254255256257…1292
Next →