AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlog
86,428Total entries
1Added by human
86,427Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,016 results
Model Releases

Algorithmic Simplification of Neural Networks with Mosaic-of-Motifs

DGX agent

arXiv:2602.14896v2 Announce Type: replace Abstract: Large-scale deep learning models are well-suited for compression. Across a variety of tasks, methods like pruning, quantization, and knowledge disti

model-releasesarxiv-cs-lg
18 May 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Masked Next-Scale Prediction for Self-supervised Scene Text Recognition

DGX agent

arXiv:2605.14885v1 Announce Type: new Abstract: Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine-grained character strokes. Training such models relie

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval

DGX agent

arXiv:2605.06132v2 Announce Type: replace Abstract: In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the 're

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

PEML: Parameter-efficient Multi-Task Learning with Optimized Continuous Prompts

DGX agent

arXiv:2605.14055v1 Announce Type: cross Abstract: Parameter-Efficient Fine-Tuning (PEFT) is widely used for adapting Large Language Models (LLMs) for various tasks. Recently, there has been an increas

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

DGX agent

arXiv:2605.14152v1 Announce Type: cross Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Stateful Reasoning via Insight Replay

DGX agent

arXiv:2605.14457v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has become a foundation for eliciting multi-step reasoning in large language models, but recent studies show that its b

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

TabPFN-3: Technical Report

DGX agent

arXiv:2605.13986v1 Announce Type: new Abstract: Tabular data underpins most high-value prediction problems in science and industry, and TabPFN has driven the foundation model revolution for this modal

model-releasesarxiv-cs-lg
15 May 2026
Model Releases

A lot of people have been wondering about Mythos, Glasswing, and the vulns we / our partners are fixing. Today, I’m excited for us to start …

DGX agent

A lot of people have been wondering about Mythos, Glasswing, and the vulns we / our partners are fixing. Today, I’m excited for us to start sharing more. (For context, I lead Glasswing @AnthropicAI.)

model-releasesboris-cherny--x
13 May 2026
Model Releases

Beyond Parameter Aggregation: Semantic Consensus for Federated Fine-Tuning of LLMs

DGX agent

arXiv:2605.11857v1 Announce Type: new Abstract: Federated fine-tuning of large language models is commonly formulated as a parameter aggregation problem. However, even parameter-efficient methods requ

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Collective Alignment in LLM Multi-Agent Systems: Disentangling Bias from Cooperation via Statistical Physics

DGX agent

arXiv:2605.10528v1 Announce Type: cross Abstract: We investigate the emergent collective dynamics of LLM-based multi-agent systems on a 2D square lattice and present a model-agnostic statistical-physi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Hint Tuning: Less Data Makes Better Reasoners

DGX agent

arXiv:2605.08665v1 Announce Type: new Abstract: Large reasoning models achieve high accuracy through extended chain-of-thought but generate 5--8 more tokens than necessary, applying verbose reasoning

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering

DGX agent

arXiv:2605.09384v1 Announce Type: cross Abstract: The reasoning gap between large and compact vision-language models (VLMs) limits the deployment of medical AI on portable clinical devices. Compact VL

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Locking Pretrained Weights via Deep Low-Rank Residual Distillation

DGX agent

arXiv:2605.10777v1 Announce Type: new Abstract: The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling the

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces

DGX agent

arXiv:2605.08904v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and tool use. However, the fundamental cognitive faculties essential

model-releasesarxiv-cs-ai
12 May 2026
Safety

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning

DGX agent

arXiv:2605.06241v2 Announce Type: replace Abstract: Reinforcement learning has become the standard for improving reasoning in large language models, yet evidence increasingly suggests that RL does not

safetyarxiv-cs-cl
12 May 2026
Applications

SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference

DGX agent

arXiv:2605.08151v1 Announce Type: cross Abstract: LLM serving platforms are increasingly deployed as multi-model cloud systems, where user demand is often long-tailed: a few popular large models recei

applicationsarxiv-cs-ai
12 May 2026
Model Releases

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

DGX agent

arXiv:2605.08412v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Test-Time Speculation

DGX agent

arXiv:2605.09329v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by using a fast draft model to generate tokens and a more accurate target model to verify them. Its perfo

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations

DGX agent

arXiv:2605.09195v1 Announce Type: new Abstract: Large language models confidently produce outdated answers, and no existing method can detect them. We show this is not an engineering failure but a str

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Mitigating Cognitive Bias in RLHF by Altering Rationality

DGX agent

arXiv:2605.06895v1 Announce Type: new Abstract: How can we make models robust to even imperfect human feedback? In reinforcement learning from human feedback (RLHF), human preferences over model outpu

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation

DGX agent

arXiv:2605.07647v1 Announce Type: cross Abstract: Automated short answer scoring (ASAS) is shifting from discriminative, fine-tuned models to large language models (LLMs) used in few-shot settings. Th

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Scaling Categorical Flow Maps

DGX agent

arXiv:2605.07820v1 Announce Type: new Abstract: Continuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they u

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

How BASF manages thousands of supply chain decisions with AlphaEvolve’s agentic algorithms

DGX agent

The agricultural and crop protection supply chain is one of the most intricate networks in the world. It takes up to two years to turn active ingredients into the final products farmers need, and a si

model-releasesgoogle-cloud-ai
7 May 2026
Model Releases

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

DGX agent

arXiv:2602.00933v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) is rapidly becoming the standard interface for Large Language Models (LLMs) to discover and invoke external t

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

DGX agent

arXiv:2605.01630v1 Announce Type: new Abstract: Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scorin

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

DGX agent

arXiv:2605.00674v1 Announce Type: new Abstract: Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference

DGX agent

arXiv:2605.00300v1 Announce Type: cross Abstract: Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the en

model-releasesarxiv-cs-lg
4 May 2026
Model Releases

Low Rank Adaptation for Adversarial Perturbation

DGX agent

arXiv:2604.27487v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA), which leverages the insight that model updates typically reside in a low-dimensional space, has significantly improved the t

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

The Impact of LLM Self-Consistency and Reasoning Effort on Automated Scoring Accuracy and Cost

DGX agent

arXiv:2604.26954v1 Announce Type: cross Abstract: Strategic model selection and reasoning settings are more effective than ensembling for optimizing automated scoring with large language models (LLMs)

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Entropy Centroids as Intrinsic Rewards for Test-Time Scaling

DGX agent

arXiv:2604.26173v1 Announce Type: cross Abstract: An effective way to scale up test-time compute of large language models is to sample multiple responses and then select the best one, as in Grok Heavy

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese

DGX agent

arXiv:2604.25926v1 Announce Type: new Abstract: The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and b

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens

DGX agent

arXiv:2604.26355v1 Announce Type: new Abstract: Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains unde

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

TildeOpen LLM: Leveraging Curriculum Learning to Achieve Equitable Language Representation

DGX agent

arXiv:2603.08182v2 Announce Type: replace-cross Abstract: Large language models often underperform in many European languages due to the dominance of English and a few high-resource languages in train

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation

DGX agent

arXiv:2503.04872v3 Announce Type: replace-cross Abstract: The challenge of reducing the size of Large Language Models (LLMs) while maintaining their performance has gained significant attention. Howev

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Benchmarking and Adapting On-Device LLMs for Clinical Decision Support

DGX agent

arXiv:2601.03266v2 Announce Type: replace Abstract: Large language models (LLMs) have rapidly advanced in clinical decision-making, yet the deployment of proprietary systems is hindered by privacy con

model-releasesarxiv-cs-cl
29 Apr 2026
Safety

Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

DGX agent

arXiv:2604.25891v1 Announce Type: new Abstract: Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavio

safetyarxiv-cs-lg
29 Apr 2026
Model Releases

CRAFT: Grounded Multi-Agent Coordination Under Partial Information

DGX agent

arXiv:2603.25268v2 Announce Type: replace Abstract: We introduce CRAFT, a multi-agent benchmark for evaluating pragmatic communication in large language models under strict partial information. In thi

model-releasesarxiv-cs-cl
29 Apr 2026
Local Ai

Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs

DGX agent

arXiv:2604.25039v1 Announce Type: new Abstract: Large Language Models (LLMs) solve many reasoning tasks via chain-of-thought (CoT) prompting, but smaller models (about 7 to 8B parameters) still strugg

local-aiarxiv-cs-cl
29 Apr 2026
Model Releases

Feasible-First Exploration for Constrained ML Deployment Optimization in Crash-Prone Hierarchical Search Spaces

DGX agent

arXiv:2604.25073v1 Announce Type: new Abstract: Deploying machine learning models under production constraints requires joint optimization over model family, quantization scheme, runtime backend, and

model-releasesarxiv-cs-lg
29 Apr 2026
Model Releases

Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer

DGX agent

arXiv:2604.25409v1 Announce Type: new Abstract: Probabilistic Transformer (PT), a white-box probabilistic model for contextual word representation, has demonstrated substantial similarity to standard

model-releasesarxiv-cs-cl
29 Apr 2026
Tutorials

ViPO: Visual Preference Optimization at Scale

DGX agent

arXiv:2604.24953v1 Announce Type: new Abstract: While preference optimization is crucial for improving visual generative models, how to effectively scale this paradigm remains largely unexplored. Curr

tutorialsarxiv-cs-cv
29 Apr 2026
Model Releases

Domain Fine-Tuning vs. Retrieval-Augmented Generation for Medical Multiple-Choice Question Answering: A Controlled Comparison at the 4B-Parameter Scale

DGX agent

arXiv:2604.23801v1 Announce Type: new Abstract: Practitioners deploying small open-weight large language models (LLMs) for medical question answering face a recurring design choice: invest in a domain

model-releasesarxiv-cs-cl
28 Apr 2026
Tutorials

Generating Verifiable Chain of Thoughts from Exection-Traces

DGX agent

arXiv:2512.00127v3 Announce Type: replace-cross Abstract: Getting language models to reason correctly about code requires training on data where each reasoning step can be checked. Current synthetic C

tutorialsarxiv-cs-ai
28 Apr 2026
Safety

Peer Identity Bias in Multi-Agent LLM Evaluation: An Empirical Study Using the TRUST Democratic Discourse Analysis Pipeline

DGX agent

arXiv:2604.22971v1 Announce Type: cross Abstract: The TRUST democratic discourse analysis pipeline exposes its large language model (LLM) components to peer model identity through multiple structural

safetyarxiv-cs-ai
28 Apr 2026
Model Releases

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns

DGX agent

arXiv:2604.23150v1 Announce Type: cross Abstract: Most recent state-of-the-art (SOTA) large language models (LLMs) use Mixture-of-Experts (MoE) architectures to scale model capacity without proportion

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation

DGX agent

arXiv:2505.16637v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently demonstrated remarkable capabilities in machine translation (MT). However, most advanced MT-specifi

model-releasesarxiv-cs-ai
28 Apr 2026
Research

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

DGX agent

arXiv:2604.24763v1 Announce Type: new Abstract: Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creatin

researcharxiv-cs-cv
28 Apr 2026
Research

From Words to Amino Acids: Does the Curse of Depth Persist?

DGX agent

arXiv:2602.21750v2 Announce Type: replace Abstract: Protein language models (PLMs) have become widely adopted as general-purpose models, demonstrating strong performance in protein engineering and de

researcharxiv-cs-lg
27 Apr 2026
← Previous
1…252253254255256…1292
Next →