AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

Content type
AllBlog
85,202Total entries
1Added by human
85,201Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
61,047 results
Model Releases

Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction

DGX agent

arXiv:2607.28079v1 Announce Type: new Abstract: Chemical property prediction plays a critical role in accelerating scientific discovery in chemistry, materials science, and drug development. However,

model-releasesarxiv-cs-lg
31 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Is it just me, or are current LLM benchmarks failing to capture actual usability? (Gemma 4 vs. Gemini/Claude Opus)

DGX agent

Disclaimer, this was kinda written with AI (Gemma 4 again) but it also did really well here, it outputted what I wanted, when I asked it to refine stuff or improve on certain areas it did that without

model-releasesr-localllama
31 Jul 2026
Model Releases

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

DGX agent

arXiv:2607.27919v1 Announce Type: new Abstract: Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independent

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

DGX agent

arXiv:2607.27378v1 Announce Type: new Abstract: Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-g

model-releasesarxiv-cs-cv
31 Jul 2026
Model Releases

APEX-Accounting

DGX agent

arXiv:2607.27189v1 Announce Type: new Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountant

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

GPT-Red: Automated Red Teaming via Self-Play at Scale

DGX agent

arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce extbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

DGX agent

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Memory bandwidth, not VRAM size, sets your tokens/sec — here's the arithmetic

DGX agent

Every week someone asks which card to buy and the thread turns into people naming GPUs they happen to own. There's an actual calculation behind it, it takes two numbers off the spec sheet, and it pred

model-releasesr-ollama
30 Jul 2026
Model Releases

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation

DGX agent

arXiv:2509.22768v3 Announce Type: replace Abstract: We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach

DGX agent

arXiv:2607.26608v1 Announce Type: new Abstract: Training-free fusion of heterogeneous multimodal large language models (MLLMs) provides a direct route for cross-scale capability transfer, yet improvem

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

Bits and Memories: Measuring Verbatim Extraction Across LLM Quantization

DGX agent

arXiv:2607.25451v1 Announce Type: new Abstract: Language models are almost always quantized before they are deployed, and a growing line of work asks whether quantization also lowers their privacy ris

model-releasesarxiv-cs-lg
29 Jul 2026
Model Releases

Emergent Latent-State Computation under Stochastic Volatility

DGX agent

arXiv:2607.25459v1 Announce Type: cross Abstract: Mechanistic interpretability has largely focused on language models and deterministic toy tasks. Much less is known about how sequence models internal

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising

DGX agent

arXiv:2607.24779v1 Announce Type: new Abstract: Online advertising bidding systems typically deploy multiple offline-trained expert models (e.g., PID controllers, model predictive control, offline RL

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

Auditing Alignment Controllability in LLMs via Political Axes

DGX agent

arXiv:2607.23519v1 Announce Type: cross Abstract: Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in dep

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

DGX agent

arXiv:2607.23159v1 Announce Type: new Abstract: Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although most are discard

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

DGX agent

arXiv:2607.23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

DGX agent

arXiv:2607.23802v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-sc

model-releasesarxiv-cs-ai
28 Jul 2026
Local Ai

Generative Video Compression with Adaptive Score Distillation

DGX agent

arXiv:2607.22772v1 Announce Type: cross Abstract: Diffusion models provide strong generative capabilities for video compression at ultra-low bitrates. Existing diffusion-based video codecs adapt base

local-aiarxiv-cs-cv
28 Jul 2026
Model Releases

LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning

DGX agent

arXiv:2607.22777v1 Announce Type: cross Abstract: Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along amino-acid se

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion

DGX agent

arXiv:2607.22696v1 Announce Type: cross Abstract: High-resolution video diffusion models built on Diffusion Transformers (DiTs) deliver strong fidelity but quickly exhaust the memory budget of a singl

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

PReSS: An Automated Black-Box Framework for Evaluating Political Stance Stability in LLMs

DGX agent

arXiv:2504.17052v4 Announce Type: replace Abstract: Existing evaluations of political bias in large language models (LLMs) typically classify outputs as left- or right-leaning. We extend this perspect

safetyarxiv-cs-cl
28 Jul 2026
Model Releases

Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

DGX agent

arXiv:2607.22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existi

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

DGX agent

arXiv:2507.21134v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their d

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation

DGX agent

arXiv:2607.23838v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base at query t

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

The Hard Decision Layer: Evidence for Committed Inference in Transformers

DGX agent

arXiv:2607.21613v1 Announce Type: cross Abstract: We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Deci

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics

DGX agent

arXiv:2604.00024v2 Announce Type: replace Abstract: Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Wom

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

90 agentic bakeoff runs: ThinkingCap vs Fable Fusion vs stock Qwen3.6-27B

DGX agent

Last week someone here said ThinkingCap and Fable Fusion 'really do beat the OG' for agentic work, so I ran it: 6 self-grading tasks, 5 reps, 3 models, 90 isolated runs. Tooling, since that's half the

model-releasesr-localllama
26 Jul 2026
Model Releases

We compared different LLMs on IMO 2026 [R]

DGX agent

There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math

model-releasesr-machinelearning
26 Jul 2026
Model Releases

A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction

DGX agent

arXiv:2607.20453v1 Announce Type: cross Abstract: Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain knowledge,

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief

DGX agent

arXiv:2607.18232v2 Announce Type: replace Abstract: Users frequently express their beliefs to large language models (LLMs). In some situations, the LLM should accept these contextual beliefs as true.

model-releasesarxiv-cs-cl
23 Jul 2026
Model Releases

NuExtract3 is now available on Ollama: 4B VLM for document-to-Markdown and structured JSON extraction

DGX agent

Disclosure: I work at NuMind, the team that trained NuExtract3. NuExtract3 is an Apache-2.0, open-weight 4B VLM based on Qwen3.5-4B. It is specialized for document understanding rather than general ch

model-releasesr-ollama
22 Jul 2026
Local Ai

OpenCode + Ollama + MCP

DGX agent

I installed OpenCode and an Ollama model (qwen3.5) sucessfully connected the model respond in OpenCode but doesn't find My MCP server, i Made one using fastMCP other models like bigPickle and openai m

local-air-ollama
22 Jul 2026
Model Releases

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

DGX agent

arXiv:2511.04689v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly

model-releasesarxiv-cs-ai
15 Jul 2026
Research

MobileSAM2: Lightweight Segment Anything for Spatial Intelligence

DGX agent

arXiv:2607.12297v1 Announce Type: new Abstract: The recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applicat

researcharxiv-cs-cv
15 Jul 2026
Model Releases

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

DGX agent

arXiv:2607.12477v1 Announce Type: new Abstract: Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenar

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

DGX agent

arXiv:2607.12963v1 Announce Type: new Abstract: As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by lo

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning

DGX agent

The NVIDIA Nemotron Model Reasoning Challenge on Kaggle attracted over 5,000 participants who all began from the same open model, benchmark, and infrastructure. The strongest solutions treated reasoni

model-releasesnvidia-developer
14 Jul 2026
Model Releases

LUMI: Tokenizer-Agnostic LLM-Based Lossless Image Compression

DGX agent

arXiv:2607.08221v1 Announce Type: new Abstract: Large language model (LLM)-based lossless image compression methods typically represent pixel data through the native text interface of a pretrained mod

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Validity of LLMs as data annotators: AMALIA on authority

DGX agent

arXiv:2607.08731v1 Announce Type: cross Abstract: A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicl

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

Can Reinforcement Learning Efficiently Discover Price Manipulation?

DGX agent

arXiv:2607.06121v1 Announce Type: cross Abstract: In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditio

model-releasesarxiv-cs-ai
9 Jul 2026
Safety

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas

DGX agent

arXiv:2601.21433v2 Announce Type: replace Abstract: Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change in framin

safetyarxiv-cs-ai
9 Jul 2026
Model Releases

Predicting LLM Safety Before Release by Simulating Deployment

DGX agent

arXiv:2607.07184v1 Announce Type: cross Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Enhancement of E-commerce Sponsored Search Relevancy with LLM

DGX agent

arXiv:2607.03886v1 Announce Type: cross Abstract: Sponsored search plays a crucial role as a revenue stream for search engines, wherein advertisers competitively bid on keywords that align with the us

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

ICR-RL: Deep Reinforcement Learning via In-Context Regression

DGX agent

arXiv:2509.11259v2 Announce Type: replace-cross Abstract: Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling them

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

NEST: Nascent Encoded Steganographic Thoughts

DGX agent

arXiv:2602.14095v2 Announce Type: replace Abstract: Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model agents; however, this oversight is compromis

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

OrthoReg: Orthogonal Regularization for Hybrid Symbolic-Neural Dynamical Systems

DGX agent

arXiv:2606.19145v2 Announce Type: replace-cross Abstract: Dynamical systems are fundamental to modeling the natural world, yet modeling them involves a persistent trade-off: manually prescribed mechan

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

tencent/Hy3

DGX agent

tencent/Hy3 New Apache 2.0 licensed model from Tencent in China: Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tence

model-releasessimon-willison
6 Jul 2026
Model Releases

sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

DGX agent

I wrote about the sqlite-utils 4.0rc1 release a couple of weeks ago. Since we only have Claude Fable on our Max subscriptions for a few more days, I decided to see if it could help me get to a 4.0 sta

model-releasessimon-willison
5 Jul 2026
← Previous
1…246247248249250…1272
Next →