AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

Content type
85,202Total entries
1Added by human
85,201Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
61,046 results
Model Releases

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

DGX agent

arXiv:2605.24703v1 Announce Type: cross Abstract: Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA). Unlike text-on

model-releasesarxiv-cs-ai
26 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

DGX agent

arXiv:2605.23281v1 Announce Type: new Abstract: Monocular metric depth estimation has achieved strong progress with large-scale training and universal-camera modeling, yet robust deployment across div

model-releasesarxiv-cs-cv
25 May 2026
Research

Physics Priors Offer Useful Accuracy-Carbon Trade-Offs in Spatio-Temporal Forecasting

DGX agent

arXiv:2509.24517v2 Announce Type: replace Abstract: Development of modern deep learning methods has been driven primarily by the push for improving model efficacy (accuracy metrics). This sole focus o

researcharxiv-cs-lg
23 May 2026
Model Releases

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation

DGX agent

arXiv:2605.21856v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive reasoning abilities across a wide range of tasks, but data contamination undermines the object

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

DrugRAG: Enhancing Pharmacy LLM Performance Through A Novel Retrieval-Augmented Generation Pipeline

DGX agent

arXiv:2512.14896v2 Announce Type: replace Abstract: In our study, we evaluated large language model (LLM) performance on pharmacy licensure-style question-answering tasks and developed an external kno

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws

DGX agent

arXiv:2502.12120v3 Announce Type: replace-cross Abstract: Scaling laws guide the development of large language models (LLMs) by offering estimates for the optimal balance of model size, tokens, and co

model-releasesarxiv-cs-cl
21 May 2026
Model Releases

TextSculptor: Training and Benchmarking Scene Text Editing

DGX agent

arXiv:2605.21090v1 Announce Type: new Abstract: Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editin

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

K-Quantization and its Impact on Output Performance

DGX agent

arXiv:2605.19645v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) have shown their remarkable capacities in many NLP tasks. However, their substantial size often pres

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder

DGX agent

arXiv:2605.19568v1 Announce Type: new Abstract: Embedding models are pivotal in industrial information retrieval systems like search and advertising. However, existing pretrained models often exhibit

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

DGX agent

arXiv:2605.18956v1 Announce Type: new Abstract: Recent motion-language models unify tasks like comprehension and generation but operate at a coarse granularity, lacking fine-grained understanding and

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

3DPhysVideo: Consistency-Guided Flow SDE for Video Generation via 3D Scene Reconstruction and Physical Simulation

DGX agent

arXiv:2605.16795v1 Announce Type: cross Abstract: Video generative models have made remarkable progress, yet they often yield visual artifacts that violate grounding in physical dynamics. Recent works

model-releasesarxiv-cs-ai
19 May 2026
Research

E-PMQ: Expert-Guided Post-Merge Quantization with Merged-Weight Anchoring

DGX agent

arXiv:2605.16882v1 Announce Type: new Abstract: Low-resource deployment constraints have made model quantization essential for deploying neural networks while preserving performance. Meanwhile, model

researcharxiv-cs-cl
19 May 2026
Model Releases

Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-worl…

DGX agent

Google’s new Gemini 3.5 Flash is the clear leader on the Intelligence vs Speed Pareto frontier and makes large gains on GDPval-AA (real-world agentic tasks), but is 5x the cost of Gemini 3 Flash @Goog

model-releasesjeremy-howard--x
19 May 2026
Model Releases

HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools

DGX agent

arXiv:2605.17106v1 Announce Type: new Abstract: Production LLM deployments increasingly maintain heterogeneous model pools spanning order-of-magnitude cost differences. Existing routers make binary st

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning

DGX agent

arXiv:2605.17949v1 Announce Type: new Abstract: Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model reasoni

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

DGX agent

arXiv:2601.12696v2 Announce Type: replace Abstract: Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerab

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

3D Segmentation Using Viewpoint-Dependent Spatial Relationships

DGX agent

arXiv:2605.15708v1 Announce Type: new Abstract: Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentat

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

Algorithmic Simplification of Neural Networks with Mosaic-of-Motifs

DGX agent

arXiv:2602.14896v2 Announce Type: replace Abstract: Large-scale deep learning models are well-suited for compression. Across a variety of tasks, methods like pruning, quantization, and knowledge disti

model-releasesarxiv-cs-lg
18 May 2026
Model Releases

Masked Next-Scale Prediction for Self-supervised Scene Text Recognition

DGX agent

arXiv:2605.14885v1 Announce Type: new Abstract: Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine-grained character strokes. Training such models relie

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval

DGX agent

arXiv:2605.06132v2 Announce Type: replace Abstract: In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the 're

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

PEML: Parameter-efficient Multi-Task Learning with Optimized Continuous Prompts

DGX agent

arXiv:2605.14055v1 Announce Type: cross Abstract: Parameter-Efficient Fine-Tuning (PEFT) is widely used for adapting Large Language Models (LLMs) for various tasks. Recently, there has been an increas

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

DGX agent

arXiv:2605.14152v1 Announce Type: cross Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Stateful Reasoning via Insight Replay

DGX agent

arXiv:2605.14457v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has become a foundation for eliciting multi-step reasoning in large language models, but recent studies show that its b

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

TabPFN-3: Technical Report

DGX agent

arXiv:2605.13986v1 Announce Type: new Abstract: Tabular data underpins most high-value prediction problems in science and industry, and TabPFN has driven the foundation model revolution for this modal

model-releasesarxiv-cs-lg
15 May 2026
Model Releases

A lot of people have been wondering about Mythos, Glasswing, and the vulns we / our partners are fixing. Today, I’m excited for us to start …

DGX agent

A lot of people have been wondering about Mythos, Glasswing, and the vulns we / our partners are fixing. Today, I’m excited for us to start sharing more. (For context, I lead Glasswing @AnthropicAI.)

model-releasesboris-cherny--x
13 May 2026
Model Releases

Beyond Parameter Aggregation: Semantic Consensus for Federated Fine-Tuning of LLMs

DGX agent

arXiv:2605.11857v1 Announce Type: new Abstract: Federated fine-tuning of large language models is commonly formulated as a parameter aggregation problem. However, even parameter-efficient methods requ

model-releasesarxiv-cs-lg
13 May 2026
Model Releases

Collective Alignment in LLM Multi-Agent Systems: Disentangling Bias from Cooperation via Statistical Physics

DGX agent

arXiv:2605.10528v1 Announce Type: cross Abstract: We investigate the emergent collective dynamics of LLM-based multi-agent systems on a 2D square lattice and present a model-agnostic statistical-physi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Hint Tuning: Less Data Makes Better Reasoners

DGX agent

arXiv:2605.08665v1 Announce Type: new Abstract: Large reasoning models achieve high accuracy through extended chain-of-thought but generate 5--8 more tokens than necessary, applying verbose reasoning

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering

DGX agent

arXiv:2605.09384v1 Announce Type: cross Abstract: The reasoning gap between large and compact vision-language models (VLMs) limits the deployment of medical AI on portable clinical devices. Compact VL

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Locking Pretrained Weights via Deep Low-Rank Residual Distillation

DGX agent

arXiv:2605.10777v1 Announce Type: new Abstract: The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling the

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces

DGX agent

arXiv:2605.08904v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and tool use. However, the fundamental cognitive faculties essential

model-releasesarxiv-cs-ai
12 May 2026
Safety

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning

DGX agent

arXiv:2605.06241v2 Announce Type: replace Abstract: Reinforcement learning has become the standard for improving reasoning in large language models, yet evidence increasingly suggests that RL does not

safetyarxiv-cs-cl
12 May 2026
Applications

SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference

DGX agent

arXiv:2605.08151v1 Announce Type: cross Abstract: LLM serving platforms are increasingly deployed as multi-model cloud systems, where user demand is often long-tailed: a few popular large models recei

applicationsarxiv-cs-ai
12 May 2026
Model Releases

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

DGX agent

arXiv:2605.08412v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Test-Time Speculation

DGX agent

arXiv:2605.09329v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by using a fast draft model to generate tokens and a more accurate target model to verify them. Its perfo

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations

DGX agent

arXiv:2605.09195v1 Announce Type: new Abstract: Large language models confidently produce outdated answers, and no existing method can detect them. We show this is not an engineering failure but a str

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Mitigating Cognitive Bias in RLHF by Altering Rationality

DGX agent

arXiv:2605.06895v1 Announce Type: new Abstract: How can we make models robust to even imperfect human feedback? In reinforcement learning from human feedback (RLHF), human preferences over model outpu

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation

DGX agent

arXiv:2605.07647v1 Announce Type: cross Abstract: Automated short answer scoring (ASAS) is shifting from discriminative, fine-tuned models to large language models (LLMs) used in few-shot settings. Th

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Scaling Categorical Flow Maps

DGX agent

arXiv:2605.07820v1 Announce Type: new Abstract: Continuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they u

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

How BASF manages thousands of supply chain decisions with AlphaEvolve’s agentic algorithms

DGX agent

The agricultural and crop protection supply chain is one of the most intricate networks in the world. It takes up to two years to turn active ingredients into the final products farmers need, and a si

model-releasesgoogle-cloud-ai
7 May 2026
Model Releases

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

DGX agent

arXiv:2602.00933v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) is rapidly becoming the standard interface for Large Language Models (LLMs) to discover and invoke external t

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

DGX agent

arXiv:2605.01630v1 Announce Type: new Abstract: Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scorin

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

DGX agent

arXiv:2605.00674v1 Announce Type: new Abstract: Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference

DGX agent

arXiv:2605.00300v1 Announce Type: cross Abstract: Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the en

model-releasesarxiv-cs-lg
4 May 2026
Model Releases

Low Rank Adaptation for Adversarial Perturbation

DGX agent

arXiv:2604.27487v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA), which leverages the insight that model updates typically reside in a low-dimensional space, has significantly improved the t

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

The Impact of LLM Self-Consistency and Reasoning Effort on Automated Scoring Accuracy and Cost

DGX agent

arXiv:2604.26954v1 Announce Type: cross Abstract: Strategic model selection and reasoning settings are more effective than ensembling for optimizing automated scoring with large language models (LLMs)

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Entropy Centroids as Intrinsic Rewards for Test-Time Scaling

DGX agent

arXiv:2604.26173v1 Announce Type: cross Abstract: An effective way to scale up test-time compute of large language models is to sample multiple responses and then select the best one, as in Grok Heavy

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese

DGX agent

arXiv:2604.25926v1 Announce Type: new Abstract: The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and b

model-releasesarxiv-cs-cl
30 Apr 2026
← Previous
1…248249250251252…1272
Next →