AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlog
86,428Total entries
1Added by human
86,427Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,016 results
Local Ai

Generative Video Compression with Adaptive Score Distillation

DGX agent

arXiv:2607.22772v1 Announce Type: cross Abstract: Diffusion models provide strong generative capabilities for video compression at ultra-low bitrates. Existing diffusion-based video codecs adapt base

local-aiarxiv-cs-cv
28 Jul 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning

DGX agent

arXiv:2607.22777v1 Announce Type: cross Abstract: Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along amino-acid se

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion

DGX agent

arXiv:2607.22696v1 Announce Type: cross Abstract: High-resolution video diffusion models built on Diffusion Transformers (DiTs) deliver strong fidelity but quickly exhaust the memory budget of a singl

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

PReSS: An Automated Black-Box Framework for Evaluating Political Stance Stability in LLMs

DGX agent

arXiv:2504.17052v4 Announce Type: replace Abstract: Existing evaluations of political bias in large language models (LLMs) typically classify outputs as left- or right-leaning. We extend this perspect

safetyarxiv-cs-cl
28 Jul 2026
Model Releases

Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

DGX agent

arXiv:2607.22692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existi

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

DGX agent

arXiv:2507.21134v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their d

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation

DGX agent

arXiv:2607.23838v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base at query t

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

The Hard Decision Layer: Evidence for Committed Inference in Transformers

DGX agent

arXiv:2607.21613v1 Announce Type: cross Abstract: We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Deci

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics

DGX agent

arXiv:2604.00024v2 Announce Type: replace Abstract: Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Wom

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

90 agentic bakeoff runs: ThinkingCap vs Fable Fusion vs stock Qwen3.6-27B

DGX agent

Last week someone here said ThinkingCap and Fable Fusion 'really do beat the OG' for agentic work, so I ran it: 6 self-grading tasks, 5 reps, 3 models, 90 isolated runs. Tooling, since that's half the

model-releasesr-localllama
26 Jul 2026
Model Releases

We compared different LLMs on IMO 2026 [R]

DGX agent

There are a few reasons why problems from International Mathematical Olympiad function as a good benchmark for LLMs: - The problems are new, not included in the training data of any model - Hard math

model-releasesr-machinelearning
26 Jul 2026
Model Releases

A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction

DGX agent

arXiv:2607.20453v1 Announce Type: cross Abstract: Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain knowledge,

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief

DGX agent

arXiv:2607.18232v2 Announce Type: replace Abstract: Users frequently express their beliefs to large language models (LLMs). In some situations, the LLM should accept these contextual beliefs as true.

model-releasesarxiv-cs-cl
23 Jul 2026
Model Releases

NuExtract3 is now available on Ollama: 4B VLM for document-to-Markdown and structured JSON extraction

DGX agent

Disclosure: I work at NuMind, the team that trained NuExtract3. NuExtract3 is an Apache-2.0, open-weight 4B VLM based on Qwen3.5-4B. It is specialized for document understanding rather than general ch

model-releasesr-ollama
22 Jul 2026
Local Ai

OpenCode + Ollama + MCP

DGX agent

I installed OpenCode and an Ollama model (qwen3.5) sucessfully connected the model respond in OpenCode but doesn't find My MCP server, i Made one using fastMCP other models like bigPickle and openai m

local-air-ollama
22 Jul 2026
Model Releases

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

DGX agent

arXiv:2511.04689v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly

model-releasesarxiv-cs-ai
15 Jul 2026
Research

MobileSAM2: Lightweight Segment Anything for Spatial Intelligence

DGX agent

arXiv:2607.12297v1 Announce Type: new Abstract: The recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applicat

researcharxiv-cs-cv
15 Jul 2026
Model Releases

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

DGX agent

arXiv:2607.12477v1 Announce Type: new Abstract: Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenar

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

DGX agent

arXiv:2607.12963v1 Announce Type: new Abstract: As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by lo

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning

DGX agent

The NVIDIA Nemotron Model Reasoning Challenge on Kaggle attracted over 5,000 participants who all began from the same open model, benchmark, and infrastructure. The strongest solutions treated reasoni

model-releasesnvidia-developer
14 Jul 2026
Model Releases

LUMI: Tokenizer-Agnostic LLM-Based Lossless Image Compression

DGX agent

arXiv:2607.08221v1 Announce Type: new Abstract: Large language model (LLM)-based lossless image compression methods typically represent pixel data through the native text interface of a pretrained mod

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Validity of LLMs as data annotators: AMALIA on authority

DGX agent

arXiv:2607.08731v1 Announce Type: cross Abstract: A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicl

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

Can Reinforcement Learning Efficiently Discover Price Manipulation?

DGX agent

arXiv:2607.06121v1 Announce Type: cross Abstract: In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditio

model-releasesarxiv-cs-ai
9 Jul 2026
Safety

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas

DGX agent

arXiv:2601.21433v2 Announce Type: replace Abstract: Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change in framin

safetyarxiv-cs-ai
9 Jul 2026
Model Releases

Predicting LLM Safety Before Release by Simulating Deployment

DGX agent

arXiv:2607.07184v1 Announce Type: cross Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Enhancement of E-commerce Sponsored Search Relevancy with LLM

DGX agent

arXiv:2607.03886v1 Announce Type: cross Abstract: Sponsored search plays a crucial role as a revenue stream for search engines, wherein advertisers competitively bid on keywords that align with the us

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

ICR-RL: Deep Reinforcement Learning via In-Context Regression

DGX agent

arXiv:2509.11259v2 Announce Type: replace-cross Abstract: Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling them

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

NEST: Nascent Encoded Steganographic Thoughts

DGX agent

arXiv:2602.14095v2 Announce Type: replace Abstract: Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model agents; however, this oversight is compromis

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

OrthoReg: Orthogonal Regularization for Hybrid Symbolic-Neural Dynamical Systems

DGX agent

arXiv:2606.19145v2 Announce Type: replace-cross Abstract: Dynamical systems are fundamental to modeling the natural world, yet modeling them involves a persistent trade-off: manually prescribed mechan

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

tencent/Hy3

DGX agent

tencent/Hy3 New Apache 2.0 licensed model from Tencent in China: Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tence

model-releasessimon-willison
6 Jul 2026
Model Releases

sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

DGX agent

I wrote about the sqlite-utils 4.0rc1 release a couple of weeks ago. Since we only have Claude Fable on our Max subscriptions for a few more days, I decided to see if it could help me get to a 4.0 sta

model-releasessimon-willison
5 Jul 2026
Model Releases

Ask the Right Comparison:Bias-Aware Bayesian Active Top-k Ranking with LLM Judges

DGX agent

arXiv:2607.02104v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise -- to rank responses, select models

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Neural Network-Based Estimation of Time-Dependent Parameters in AR(p) Processes

DGX agent

arXiv:2607.00470v1 Announce Type: cross Abstract: We investigate a forecasting framework based on a simple discrete-time dynamic model with coefficients varying in time. The parameters of the model ar

model-releasesarxiv-cs-lg
2 Jul 2026
Model Releases

Google named a Leader in 2026 Gartner® Magic Quadrant™ for Analytics and Business Intelligence Platforms for third year in a row

DGX agent

For the third consecutive year, Google has been recognized as a Leader in the 2026 Gartner® Magic Quadrant™ for Analytics and Business Intelligence Platforms. This recognition comes on the heels of Go

model-releasesgoogle-cloud-ai
1 Jul 2026
Model Releases

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

DGX agent

arXiv:2606.30059v1 Announce Type: new Abstract: Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

DGX agent

arXiv:2601.17642v2 Announce Type: replace Abstract: Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal o

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

Training-free Truthfulness Detection via Sparse MLP Value Vectors

DGX agent

arXiv:2509.17932v2 Announce Type: replace Abstract: Large language models (LLMs) are prone to generating factually incorrect content, motivating methods for assessing truthfulness from internal model

model-releasesarxiv-cs-cl
29 Jun 2026
Model Releases

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

DGX agent

arXiv:2606.26618v1 Announce Type: new Abstract: Large pretrained text-to-speech (TTS) models sound almost human for well-resourced languages, but much worse for languages that are rare in their traini

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context

DGX agent

arXiv:2606.26654v1 Announce Type: new Abstract: Personalized language-model assistants are often evaluated through a memory lens: can a model recall preferences users have explicitly stated in dialogu

model-releasesarxiv-cs-cl
26 Jun 2026
Research

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs

DGX agent

arXiv:2511.05933v2 Announce Type: replace Abstract: Reinforcement learning (RL) is often credited with improving language model reasoning at the expense of knowledge. We challenge this narrative by sh

researcharxiv-cs-cl
25 Jun 2026
Model Releases

Not All Invariants Are Equal: Curating Training Data to Accelerate Program Verification with SLMs

DGX agent

arXiv:2603.15510v2 Announce Type: replace Abstract: The synthesis of inductive loop invariants remains a critical bottleneck in automated program verification. While Large Language Models (LLMs) show

model-releasesarxiv-cs-lg
24 Jun 2026
Model Releases

You Don't Need to Run Every Eval

DGX agent

arXiv:2606.24020v1 Announce Type: new Abstract: A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare

model-releasesarxiv-cs-lg
24 Jun 2026
Model Releases

Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation

DGX agent

arXiv:2606.22918v1 Announce Type: new Abstract: Maintaining physical consistency in video generators and world models increasingly relies on vision-language models (VLMs) as automated judges that prov

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Essential Subspace Merging for Multi-Task Learning

DGX agent

arXiv:2606.19164v2 Announce Type: replace Abstract: Model merging aims to enable multi-task learning by integrating the capabilities of multiple models fine-tuned from the same pre-trained checkpoint

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

Hybrid Compression: Integrating Pruning and Quantization for Optimized Neural Networks

DGX agent

arXiv:2606.22935v1 Announce Type: new Abstract: Deep neural networks have witnessed remarkable advancements in recent years and have become integral to various applications. However, alongside these d

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Physiology-Aware CNN and Zero-Shot Multimodal LLMs for ECG Image Classification: A Comparative Study

DGX agent

arXiv:2606.22889v1 Announce Type: new Abstract: Multimodal large language models (LLMs) are increasingly adopted to interpret 12-lead ECG images, though the interpretations often lack validation. Howe

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology

DGX agent

arXiv:2606.12407v1 Announce Type: new Abstract: General-purpose large language models (LLMs) are routinely used as baselines when evaluating specialized pathology models on whole-slide images (WSIs).

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

DGX agent

arXiv:2606.09864v1 Announce Type: cross Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measu

model-releasesarxiv-cs-ai
10 Jun 2026
← Previous
1…250251252253254…1292
Next →