AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

86,993Total entries
1Added by human
86,992Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,477 results
18 Aug 2026

LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks

Model ReleasesDGX agent

arXiv:2608.14927v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collabora

MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration

Model ReleasesDGX agent

arXiv:2601.23049v2 Announce Type: replace Abstract: Medical calculators are fundamental to quantitative, evidence-based clinical practice. However, their real-world use is an adaptive, multi-stage pro

MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter

Research
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2608.16709v1 Announce Type: cross Abstract: A radiologist reading a model's output faces two problems. The model returns a number and no reason, and any system that turns that number into readab

Misconception Diagnosis From Student-Tutor Dialogue: Generate, Retrieve, Rerank

Model ReleasesDGX agent

arXiv:2602.02414v2 Announce Type: replace Abstract: Timely and accurate identification of student misconceptions is key to improving learning outcomes and pre-empting the compounding of student errors

MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios

Model ReleasesDGX agent

arXiv:2505.21333v3 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, the

Muse Glimmer 30B is now available on Fireworks' Dedicated Training API for both LoRA and Full-Parameter fine-tuning. This is a U.S.-develope…

Model ReleasesDGX agent

Muse Glimmer 30B is now available on Fireworks' Dedicated Training API for both LoRA and Full-Parameter fine-tuning. This is a U.S.-developed, open-weight model and one of the strongest of its size fo

Neurosymbolic Embodied Agents

SafetyDGX agent

arXiv:2608.16794v1 Announce Type: cross Abstract: Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dyn

Perspective-Invariant Attack with Enhanced Transferability of Adversarial Examples

Local AiDGX agent

arXiv:2608.15115v1 Announce Type: new Abstract: Adversarial examples generated on a surrogate deep neural network (DNN) can often successfully fool other black-box DNN models. This cross-model transfe

Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

Model ReleasesDGX agent

arXiv:2608.16831v1 Announce Type: new Abstract: Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a f

Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency

ResearchDGX agent

arXiv:2608.16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring

PRISM: Streaming Human Motion Generation with Per-Joint Latent Decomposition

Model ReleasesDGX agent

arXiv:2603.08590v3 Announce Type: replace Abstract: Text-to-motion generation has advanced with larger corpora and stronger generators, yet many models still rely on holistic frame- or clip-level late

Privacy-Preserving Decentralized Federated Learning via Explainable Adaptive Differential Privacy

ApplicationsDGX agent

arXiv:2509.10691v3 Announce Type: replace-cross Abstract: Decentralized federated learning enables collaborative model training without a central server, but shared model updates can still leak sensit

Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026

Model ReleasesDGX agent

arXiv:2608.16318v1 Announce Type: cross Abstract: Recent advances in Generative Artificial Intelligence (GenAI) have substantially improved the ability of large language models (LLMs) to generate and

ROC-n-reroll: How verifier imperfection affects test-time scaling

Model ReleasesDGX agent

arXiv:2507.12399v3 Announce Type: replace Abstract: Test-time scaling aims to improve language model performance by leveraging additional compute during inference. Many works have empirically studied

Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB

Model ReleasesDGX agent

I managed to run the 143–144 GiB DeepSeek-V4-Flash-0731 UD-Q4_K_XL GGUF on four RTX 3060 12GB cards while keeping a 360k–376k context window. Hardware: CPU: Intel Core i9-10920X, 12C/24T RAM: 128 GB D

S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices

Model ReleasesDGX agent

arXiv:2608.15018v1 Announce Type: new Abstract: Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While speculative de

Score Attack: A Lower Bound Technique for Optimal Differentially Private Learning

Model ReleasesDGX agent

arXiv:2303.07152v3 Announce Type: replace-cross Abstract: Achieving optimal statistical performance while ensuring the privacy of personal data is a challenging yet crucial objective in modern data an

Securing AI-Generated Code: A Just-in-Time Vulnerability Detection and Remediation Pipeline

Model ReleasesDGX agent

arXiv:2608.16187v1 Announce Type: cross Abstract: AI-assisted development tools generate vulnerable code at significant rates, yet few automated mechanisms exist to detect, enrich, fix, and verify sec

Sharing a new way to work with Stable Audio

Model ReleasesDGX agent

Stable Audio 3.0 now offers a DAW‑plugin that brings real‑time audio generation directly into your favorite digital audio workstation, letting you arrange, edit, and mix generated tracks like any othe

StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

Model ReleasesDGX agent

arXiv:2608.15089v1 Announce Type: new Abstract: Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate

Sterilizable Scene Graph Generation for Operating Rooms

Model ReleasesDGX agent

arXiv:2608.16469v1 Announce Type: new Abstract: Scene graph generation from surgical video enables a holistic and structured understanding of surgical scenes by modeling objects and their semantic rel

SubZero+: Efficient Zeroth-Order LLM Fine-Tuning via Large Learning Rates

Model ReleasesDGX agent

arXiv:2608.15665v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables backpropagation-free fine-tuning of large language models, but existing ZO methods suffer from high-variance grad

SUGFW+: An Uncertainty-guided Feature Weighting Framework for Cold Start Active Adaptation of SAM in Medical Image Segmentation

TutorialsDGX agent

arXiv:2608.16110v1 Announce Type: new Abstract: Cold Start Active Learning (CSAL) is important in improving the performance of a medical image segmentation model with low annotation budget by querying

TDD-Agent: Test-Driven Reasoning for Code Generation

Model ReleasesDGX agent

arXiv:2608.16742v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level tasks remains

The Commercial Tax: Rent-vs-Own Blind Spots in Multi-Hop Retrieval Benchmarks

Model ReleasesDGX agent

arXiv:2608.16096v1 Announce Type: cross Abstract: Enterprises connect language models to their own data through retrieval. The benchmarks that rank multi-hop retrieval systems leave out two facts a bu

Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis

AgentsDGX agent

arXiv:2608.16775v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly being deployed in cybersecurity operations to assist cybersecurity analysts with rapid decision-making a

TRACE-CASH: Trial-History-Conditioned Reinforcement Learning for Adaptive Configuration Exploration in Time-Series CASH

Local AiDGX agent

arXiv:2608.16410v1 Announce Type: new Abstract: Combined algorithm selection and hyperparameter optimization (CASH) searches a conditional space in which the selected model determines which hyperparam

Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis

Model ReleasesDGX agent

arXiv:2608.16379v1 Announce Type: new Abstract: Evaluating speech recognition for a Kurdish variety written in a Latin field orthography, using a model that outputs Arabic script, creates a measuremen

ViTaR: Visuo-Tactile Residual Adaptation for Foundation VLA Manipulation

Model ReleasesDGX agent

arXiv:2608.15816v1 Announce Type: new Abstract: As Vision-Language-Action (VLA) models scale toward real-world deployment, contact-rich manipulation exposes a critical blind spot: these policies encod

When Do Explanations Help In-Context Learning? A Comparative Study of Natural Language Explanation Types and Faithfulness

SafetyDGX agent

arXiv:2608.16627v1 Announce Type: cross Abstract: Natural language explanations (NLEs) are increasingly used as inputs, for example, as few-shot rationales that influence model behavior in in-context

When Less Is Enough: Context Selection and Prompting Strategies for Bengali News Headline Generation

Model ReleasesDGX agent

arXiv:2608.15879v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance in text generation tasks, yet their effectiveness on headline generation remains sensitive to

17 Aug 2026

Adaptive Stopping for Multi-Turn LLM Reasoning

ApplicationsDGX agent

arXiv:2604.01413v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly rely on multi-turn reasoning and interaction, such as adaptive retrieval-augmented generation (RAG)

Attention Capture Is Not Detection: A Two-Stage Account of How Humans Miss Localized AI Image Edits

Model ReleasesDGX agent

arXiv:2608.13865v1 Announce Type: new Abstract: As AI-generated image edits proliferate, the platforms meant to curb the resulting disinformation treat detectability as a single, undifferentiated prop

Buy the Rumor, Sell the News: When Is News Priced In?

Model ReleasesDGX agent

arXiv:2608.14014v1 Announce Type: new Abstract: Two old market sayings hold that news is already priced in by the time it is published, and that the rumor is bought while the news is sold. Both place

CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing

SafetyDGX agent

arXiv:2608.13925v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a single forward pass. However, existing dLLMs

CVT-Bench: Probing Spatial-State Integrity through Counterfactual Viewpoint Transformations

Model ReleasesDGX agent

arXiv:2603.21114v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) perform strongly on isolated spatial tasks, but whether their predictions remain persistent and mutually co

Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis

Model ReleasesDGX agent

arXiv:2608.13608v1 Announce Type: new Abstract: Agentic 'Continual Learning Harnesses', systems that pair an LLM with retrieval or memory to improve from feedback without retraining, have shown growin

LightTeaNet: A Weakly Supervised Lightweight CNN for Multi-Label Tea Leaf Disease Detection and Localization

Model ReleasesDGX agent

arXiv:2608.14178v1 Announce Type: new Abstract: Tea is known as an important crop in many parts of South and Southeast Asia, yet the production of tea is still hampered by the multiple diseases that d

Ling 3.0 support merged into llama.cpp

Model ReleasesDGX agent

Support for the new ling 3.0 models has been merged into llama.cpp: https://github.com/ggml-org/llama.cpp/pull/26608#event-29549472828 Ling tiny 8b1b - https://huggingface.co/inclusionAI/Ling-3.0-tiny

MACS: A Hybrid Multi-Agent Framework for Reliable Conversational E-Commerce Recommendation

Model ReleasesDGX agent

arXiv:2608.14068v1 Announce Type: cross Abstract: Conversational recommendation for e-commerce is increasingly mediated by large language models (LLMs), yet many real-world deployments operate under a

Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems

Local AiDGX agent

arXiv:2608.13571v1 Announce Type: cross Abstract: When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time. This retry overh

On the Brittleness of Maximum Likelihood Estimation for Gaussian Process Hyperparameter Optimization

ResearchDGX agent

arXiv:2608.13793v1 Announce Type: cross Abstract: Machine learning (ML) has become an indispensable part of modern engineering design workflows. A crucial step in training an ML model is the selection

Responsiveness Verification: Will Predictions Change? How Much? How Often?

SafetyDGX agent

arXiv:2507.02169v2 Announce Type: replace Abstract: Machine learning models are often used in applications where their inputs change due to routine interactions, strategic manipulation, or noise. In s

Unpopular opinion : Qwen 3.8 27b is not an overthinker

Model ReleasesDGX agent

Yes it uses a ton more reasoning tokens than 3.6 did But test in on the same tasks with the other chinese models, glm 5.3, deepseek v4 flash and pro, etc it's really similar, and they are needed The r

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

Model ReleasesDGX agent

arXiv:2608.14375v1 Announce Type: new Abstract: Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering

16 Aug 2026

Dear Dario, 1. If Claude can cure cancer to save people like your dad, why should we 'pace the progress'? Does that mean more people with He…

Model ReleasesDGX agent

Dear Dario, 1. If Claude can cure cancer to save people like your dad, why should we 'pace the progress'? Does that mean more people with Hepatitis C will die? 2. If Fable is so cyber-capable that it

Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72

Model ReleasesDGX agent

https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/ 4k tokens per second per GPU of which there are 72. 350 tokens per s

Qwen3.8 27B reasoning effort low/medium/xhigh comparison

Model ReleasesDGX agent

I did a short test of the different reasoning efforts, since on default xhigh the model thinks a lot. Not very scientific, just a quick 'generate an SVG of a pelican on a bicycle' prompt with 3 differ

There are two separate questions regarding Dario Amodei’s post about AI and biology: motivation and reality. Motivation. Dario started in bi…

Model ReleasesDGX agent

There are two separate questions regarding Dario Amodei’s post about AI and biology: motivation and reality. Motivation. Dario started in biology (was a PhD student of the great Bill Bialek). It might

15 Aug 2026

Qwen3.8-27B is now part of your everyday life — from smartphones to vehicles. ⚡️🚀Appreciate your work! @MediaTek

Model ReleasesDGX agent

Qwen3.8-27B is now part of your everyday life — from smartphones to vehicles. ⚡️🚀Appreciate your work! @MediaTek Congratulations to the @Alibaba_Qwen on the launch of Qwen 3.8! MediaTek continues our

The new DeepSeek-V4-Pro (0813) is now fully rolled out on Ollama's cloud and included in Pro and Max subscriptions. ollama run deepseek-v4-p…

Model ReleasesDGX agent

The new DeepSeek-V4-Pro (0813) is now fully rolled out on Ollama's cloud and included in Pro and Max subscriptions. ollama run deepseek-v4-pro:cloud Hosted in the US with Zero Data Retention (ZDR) and

14 Aug 2026

A Generative Approach for Improving Multi-Label Defect Classification in Photovoltaic Modules

Model ReleasesDGX agent

arXiv:2608.12725v1 Announce Type: new Abstract: This paper addresses the challenge of multi-label defect classification in electroluminescence (EL) images of photovoltaic (PV) cells. Training models o

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

ResearchDGX agent

arXiv:2608.13430v1 Announce Type: cross Abstract: Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbaliz

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Model ReleasesDGX agent

arXiv:2608.13560v1 Announce Type: cross Abstract: Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process cent

CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

Model ReleasesDGX agent

arXiv:2608.12805v1 Announce Type: new Abstract: Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constra

Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues

ResearchDGX agent

arXiv:2608.12599v1 Announce Type: new Abstract: Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enacting withdrawn

EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding

ResearchDGX agent

arXiv:2608.13072v1 Announce Type: new Abstract: Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition protocols and indi

From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion

Local AiDGX agent

arXiv:2608.13043v1 Announce Type: new Abstract: Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based acceleration

Geometric and Behavioral Stratification in Transformer Residual Streams

ResearchDGX agent

arXiv:2608.12447v1 Announce Type: cross Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of di

Incremental Evaluation and Training in Relational Deep Learning

ApplicationsDGX agent

arXiv:2608.13023v1 Announce Type: new Abstract: Relational Deep Learning (RDL) models multi-tabular databases as temporal heterogeneous graphs to enable end-to-end representation learning. However, pr

← Previous
1…325326327328329…1042
Next →