AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,597 results
30 Jul 2026

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

Model ReleasesDGX agent

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell

Memory bandwidth, not VRAM size, sets your tokens/sec — here's the arithmetic

Model ReleasesDGX agent

Every week someone asks which card to buy and the thread turns into people naming GPUs they happen to own. There's an actual calculation behind it, it takes two numbers off the spec sheet, and it pred

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2509.22768v3 Announce Type: replace Abstract: We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models

Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach

Model ReleasesDGX agent

arXiv:2607.26608v1 Announce Type: new Abstract: Training-free fusion of heterogeneous multimodal large language models (MLLMs) provides a direct route for cross-scale capability transfer, yet improvem

29 Jul 2026

Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization

Model ReleasesDGX agent

arXiv:2607.25451v1 Announce Type: new Abstract: Language models are almost always quantized before they are deployed, and a growing line of work asks whether quantization also lowers their privacy ris

Emergent Latent-State Computation under Stochastic Volatility

Model ReleasesDGX agent

arXiv:2607.25459v1 Announce Type: cross Abstract: Mechanistic interpretability has largely focused on language models and deterministic toy tasks. Much less is known about how sequence models internal

HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising

Model ReleasesDGX agent

arXiv:2607.24779v1 Announce Type: new Abstract: Online advertising bidding systems typically deploy multiple offline-trained expert models (e.g., PID controllers, model predictive control, offline RL

28 Jul 2026

Auditing Alignment Controllability in LLMs via Political Axes

Model ReleasesDGX agent

arXiv:2607.23519v1 Announce Type: cross Abstract: Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in dep

CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

Model ReleasesDGX agent

arXiv:2607.23159v1 Announce Type: new Abstract: Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although most are discard

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

Model ReleasesDGX agent

arXiv:2607.23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Model ReleasesDGX agent

arXiv:2607.23802v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-sc

Generative Video Compression with Adaptive Score Distillation

Local AiDGX agent

arXiv:2607.22772v1 Announce Type: cross Abstract: Diffusion models provide strong generative capabilities for video compression at ultra-low bitrates. Existing diffusion-based video codecs adapt base

LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning

Model ReleasesDGX agent

arXiv:2607.22777v1 Announce Type: cross Abstract: Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along amino-acid se

MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion

Model ReleasesDGX agent

arXiv:2607.22696v1 Announce Type: cross Abstract: High-resolution video diffusion models built on Diffusion Transformers (DiTs) deliver strong fidelity but quickly exhaust the memory budget of a singl

PReSS: An Automated Black-Box Framework for Evaluating Political Stance Stability in LLMs

SafetyDGX agent

arXiv:2504.17052v4 Announce Type: replace Abstract: Existing evaluations of political bias in large language models (LLMs) typically classify outputs as left- or right-leaning. We extend this perspect

Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

Model ReleasesDGX agent

arXiv:2607.22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existi

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

Model ReleasesDGX agent

arXiv:2507.21134v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their d

TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.23838v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base at query t

27 Jul 2026

The Hard Decision Layer: Evidence for Committed Inference in Transformers

Model ReleasesDGX agent

arXiv:2607.21613v1 Announce Type: cross Abstract: We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Deci

WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics

Model ReleasesDGX agent

arXiv:2604.00024v2 Announce Type: replace Abstract: Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Wom

26 Jul 2026

90 agentic bakeoff runs: ThinkingCap vs Fable Fusion vs stock Qwen3.6-27B

Model ReleasesDGX agent

Last week someone here said ThinkingCap and Fable Fusion 'really do beat the OG' for agentic work, so I ran it: 6 self-grading tasks, 5 reps, 3 models, 90 isolated runs. Tooling, since that's half the

We compared different LLMs on IMO 2026 [R]

Model ReleasesDGX agent

There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math

24 Jul 2026

A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction

Model ReleasesDGX agent

arXiv:2607.20453v1 Announce Type: cross Abstract: Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain knowledge,

23 Jul 2026

It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief

Model ReleasesDGX agent

arXiv:2607.18232v2 Announce Type: replace Abstract: Users frequently express their beliefs to large language models (LLMs). In some situations, the LLM should accept these contextual beliefs as true.

22 Jul 2026

NuExtract3 is now available on Ollama: 4B VLM for document-to-Markdown and structured JSON extraction

Model ReleasesDGX agent

Disclosure: I work at NuMind, the team that trained NuExtract3. NuExtract3 is an Apache-2.0, open-weight 4B VLM based on Qwen3.5-4B. It is specialized for document understanding rather than general ch

OpenCode + Ollama + MCP

Local AiDGX agent

I installed OpenCode and an Ollama model (qwen3.5) sucessfully connected the model respond in OpenCode but doesn't find My MCP server, i Made one using fastMCP other models like bigPickle and openai m

15 Jul 2026

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

Model ReleasesDGX agent

arXiv:2511.04689v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly

MobileSAM2: Lightweight Segment Anything for Spatial Intelligence

ResearchDGX agent

arXiv:2607.12297v1 Announce Type: new Abstract: The recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applicat

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

Model ReleasesDGX agent

arXiv:2607.12477v1 Announce Type: new Abstract: Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenar

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

Model ReleasesDGX agent

arXiv:2607.12963v1 Announce Type: new Abstract: As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by lo

14 Jul 2026

Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning

Model ReleasesDGX agent

The NVIDIA Nemotron Model Reasoning Challenge on Kaggle attracted over 5,000 participants who all began from the same open model, benchmark, and infrastructure. The strongest solutions treated reasoni

10 Jul 2026

LUMI: Tokenizer-Agnostic LLM-Based Lossless Image Compression

Model ReleasesDGX agent

arXiv:2607.08221v1 Announce Type: new Abstract: Large language model (LLM)-based lossless image compression methods typically represent pixel data through the native text interface of a pretrained mod

Validity of LLMs as data annotators: AMALIA on authority

Model ReleasesDGX agent

arXiv:2607.08731v1 Announce Type: cross Abstract: A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicl

9 Jul 2026

Can Reinforcement Learning Efficiently Discover Price Manipulation?

Model ReleasesDGX agent

arXiv:2607.06121v1 Announce Type: cross Abstract: In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditio

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas

SafetyDGX agent

arXiv:2601.21433v2 Announce Type: replace Abstract: Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change in framin

Predicting LLM Safety Before Release by Simulating Deployment

Model ReleasesDGX agent

arXiv:2607.07184v1 Announce Type: cross Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about

7 Jul 2026

Enhancement of E-commerce Sponsored Search Relevancy with LLM

Model ReleasesDGX agent

arXiv:2607.03886v1 Announce Type: cross Abstract: Sponsored search plays a crucial role as a revenue stream for search engines, wherein advertisers competitively bid on keywords that align with the us

ICR-RL: Deep Reinforcement Learning via In-Context Regression

Model ReleasesDGX agent

arXiv:2509.11259v2 Announce Type: replace-cross Abstract: Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling them

NEST: Nascent Encoded Steganographic Thoughts

Model ReleasesDGX agent

arXiv:2602.14095v2 Announce Type: replace Abstract: Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model agents; however, this oversight is compromis

OrthoReg: Orthogonal Regularization for Hybrid Symbolic-Neural Dynamical Systems

Model ReleasesDGX agent

arXiv:2606.19145v2 Announce Type: replace-cross Abstract: Dynamical systems are fundamental to modeling the natural world, yet modeling them involves a persistent trade-off: manually prescribed mechan

6 Jul 2026

tencent/Hy3

Model ReleasesDGX agent

tencent/Hy3 New Apache 2.0 licensed model from Tencent in China: Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tence

5 Jul 2026

sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

Model ReleasesDGX agent

I wrote about the sqlite-utils 4.0rc1 release a couple of weeks ago. Since we only have Claude Fable on our Max subscriptions for a few more days, I decided to see if it could help me get to a 4.0 sta

3 Jul 2026

Ask the Right Comparison:Bias-Aware Bayesian Active Top-k Ranking with LLM Judges

Model ReleasesDGX agent

arXiv:2607.02104v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise -- to rank responses, select models

2 Jul 2026

Neural Network-Based Estimation of Time-Dependent Parameters in AR(p) Processes

Model ReleasesDGX agent

arXiv:2607.00470v1 Announce Type: cross Abstract: We investigate a forecasting framework based on a simple discrete-time dynamic model with coefficients varying in time. The parameters of the model ar

1 Jul 2026

Google named a Leader in 2026 Gartner® Magic Quadrant™ for Analytics and Business Intelligence Platforms for third year in a row

Model ReleasesDGX agent

For the third consecutive year, Google has been recognized as a Leader in the 2026 Gartner® Magic Quadrant™ for Analytics and Business Intelligence Platforms. This recognition comes on the heels of Go

30 Jun 2026

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

Model ReleasesDGX agent

arXiv:2606.30059v1 Announce Type: new Abstract: Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external

29 Jun 2026

Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

Model ReleasesDGX agent

arXiv:2601.17642v2 Announce Type: replace Abstract: Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal o

Training-free Truthfulness Detection via Sparse MLP Value Vectors

Model ReleasesDGX agent

arXiv:2509.17932v2 Announce Type: replace Abstract: Large language models (LLMs) are prone to generating factually incorrect content, motivating methods for assessing truthfulness from internal model

26 Jun 2026

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

Model ReleasesDGX agent

arXiv:2606.26618v1 Announce Type: new Abstract: Large pretrained text-to-speech (TTS) models sound almost human for well-resourced languages, but much worse for languages that are rare in their traini

SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context

Model ReleasesDGX agent

arXiv:2606.26654v1 Announce Type: new Abstract: Personalized language-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogu

25 Jun 2026

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs

ResearchDGX agent

arXiv:2511.05933v2 Announce Type: replace Abstract: Reinforcement learning (RL) is often credited with improving language model reasoning at the expense of knowledge. We challenge this narrative by sh

24 Jun 2026

Not All Invariants Are Equal: Curating Training Data to Accelerate Program Verification with SLMs

Model ReleasesDGX agent

arXiv:2603.15510v2 Announce Type: replace Abstract: The synthesis of inductive loop invariants remains a critical bottleneck in automated program verification. While Large Language Models (LLMs) show

You Don't Need to Run Every Eval

Model ReleasesDGX agent

arXiv:2606.24020v1 Announce Type: new Abstract: A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare

23 Jun 2026

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

Model ReleasesDGX agent

arXiv:2606.22918v1 Announce Type: new Abstract: Maintaining physical consistency in video generators and world models increasingly relies on vision-language models (VLMs) as automated judges that prov

Essential Subspace Merging for Multi-Task Learning

Model ReleasesDGX agent

arXiv:2606.19164v2 Announce Type: replace Abstract: Model merging aims to enable multi-task learning by integrating the capabilities of multiple models fine-tuned from the same pre-trained checkpoint

Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks

Model ReleasesDGX agent

arXiv:2606.22935v1 Announce Type: new Abstract: Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these d

Physiology-Aware CNN and Zero-Shot Multimodal LLMs for ECG Image Classification: A Comparative Study

Model ReleasesDGX agent

arXiv:2606.22889v1 Announce Type: new Abstract: Multimodal large language models (LLMs) are increasingly adopted to interpret 12-lead ECG images, though the interpretations often lack validation. Howe

11 Jun 2026

How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology

Model ReleasesDGX agent

arXiv:2606.12407v1 Announce Type: new Abstract: General-purpose large language models (LLMs) are routinely used as baselines when evaluating specialized pathology models on whole-slide images (WSIs).

10 Jun 2026

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

Model ReleasesDGX agent

arXiv:2606.09864v1 Announce Type: cross Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measu

9 Jun 2026

A Mechanistic Analysis of Adversarial Fine-tuning of Vision Transformers

ApplicationsDGX agent

arXiv:2606.07593v1 Announce Type: cross Abstract: The widespread use of image classification models in high-risk, real-world situations necessitates making these models robust to slight disturbances o

← Previous
1…195196197198199…1010
Next →