AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,617
  • Agents7,497
  • Applications5,365
  • Concepts5
  • Hardware1,816
  • Industry6,151
  • Local Ai4,900
  • Model Releases23,593
  • Research19,967
  • Safety13,267
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,617
  • Agents7,497
  • Applications5,365
  • Concepts5
  • Hardware1,816
  • Industry6,151
  • Local Ai4,900
  • Model Releases23,593
  • Research19,967
  • Safety13,267
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

87,617Total entries
1Added by human
87,616Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,987 results
6 Aug 2026

In addition to the upgrade in intelligence with GPT-5.6 Luna, Free and Go users can now use the “Think” button for more reasoning on harder …

Model ReleasesDGX agent

OpenAI has released GPT‑5.6 Sol, which powers both instant and deep‑reasoning modes for ChatGPT Plus and Pro customers, delivering fact‑centric responses. Free and Go tier users will receive unlimited

KV cache quantization benchmarks: 413 pairs tested on Qwen 3.6 27B, Gemma 4 31B. KLD with BeeLlama.cpp v0.4.0: KVarN 6-bit beats q8_0, precision tail 1024 dominates

Model ReleasesDGX agent

Link to the article: KV Cache Quantization Benchmarks: KVarN, Precision Tail KLD benchmarks with BeeLlama.cpp v0.4.0, fork of llama.cpp with more KV cache quantization options. Models: Qwen 3.6 27B Q5

LaPrune: Controllable Differentiable Sparsity at Million Scale

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2608.04057v1 Announce Type: cross Abstract: Top-k selection determines which components of a sparse model remain active. Hard selection blocks gradients, while continuous relaxations often coupl

MediRec: Enhancing Chinese Medication Recommendation with Explainable Clinical Reasoning

Model ReleasesDGX agent

arXiv:2510.21084v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong potential for clinical decision support through their advanced language understanding and reaso

nvidia/NVIDIA-Nemotron-Parse-2.0 · Hugging Face

Model ReleasesDGX agent

NVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information. Given a Red, Green, Blu

OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing

Model ReleasesDGX agent

arXiv:2608.04434v1 Announce Type: new Abstract: Recent large language models (LLMs) have demonstrated remarkable progress in constraint-aware navigation, maze reasoning, and graph reasoning. However,

PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning

Model ReleasesDGX agent

arXiv:2608.04575v1 Announce Type: cross Abstract: Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language model

PSI3D: Plug-and-Play 3D Stochastic Inference with Slice-wise Latent Diffusion Prior

ResearchDGX agent

arXiv:2512.18367v2 Announce Type: replace-cross Abstract: Diffusion models are highly expressive image priors for Bayesian inverse problems. However, most diffusion models cannot operate on large-scal

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

Model ReleasesDGX agent

arXiv:2608.04514v1 Announce Type: new Abstract: Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-co

REZE: Recognition-Based Zero-Shot Extraction for Video Temporal Grounding

ResearchDGX agent

arXiv:2608.04480v1 Announce Type: new Abstract: Video temporal grounding (VTG) refers to the task of identifying the time interval in a video that corresponds to a given natural-language query. A comm

Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load

Model ReleasesDGX agent

arXiv:2608.05018v1 Announce Type: new Abstract: Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law designate as

Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays

Model ReleasesDGX agent

arXiv:2608.04043v1 Announce Type: new Abstract: Resistive pressure arrays are the cheapest and most widely shipped tactile sensors, yet tactile representation learning has concentrated on optical sens

The Loss Does Not See the Basis, but Adam Does

Model ReleasesDGX agent

arXiv:2608.05136v1 Announce Type: new Abstract: Gradient descent on a factored model W = UV^op is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization,

This webinar is happening in 30 minutes, and that means there's still time to register! Following the discussion will be an open Q&A with ou…

Model ReleasesDGX agent

This webinar is happening in 30 minutes, and that means there's still time to register! Following the discussion will be an open Q&A with our Head of AI Education @Prof_OZ, and the @arizeai team. See

Transfer Learning for Named Entity Recognition of Classical Latin through LLM Prompting

Model ReleasesDGX agent

arXiv:2608.04015v1 Announce Type: new Abstract: With the increase in digitized resources of Classical Latin texts and modern breakthroughs of Large Language Models (LLMs), I contribute to ancient lang

Uber burned through its 2026 AI coding budget in four months. Microsoft canceled most of its Claude Code licenses six months after rolling t…

Model ReleasesDGX agent

Uber burned through its 2026 AI coding budget in four months. Microsoft canceled most of its Claude Code licenses six months after rolling them out. The mechanics are simple: per-token cost keeps fall

5 Aug 2026

A Physics-Flavored Transformer Network for Parametrizing Contraction Dynamics of Engineered Skeletal Muscle Tissues

SafetyDGX agent

arXiv:2608.03927v1 Announce Type: new Abstract: Engineered Skeletal Muscle Tissues (ESMs) have become a key structure for biomedical disease modeling and pharmacological screening, yet their functiona

A ultra-lightweight mini agent - zero framework and local/ollama first

Model ReleasesDGX agent

https://github.com/mohsinkaleem/agent-mini.git A minimal 3k lines, local-first AI agent you can actually understand and extend. Optimized for smaller local models like qwen 3.6 4b or 9b pip install ag

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2608.03744v1 Announce Type: new Abstract: Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be

Automated Visualization Code Synthesis via Multi-Path Reasoning and Feedback-Driven Optimization

Model ReleasesDGX agent

arXiv:2502.11140v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have become a cornerstone for automated visualization code generation, enabling users to create charts through na

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

Model ReleasesDGX agent

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding

Model ReleasesDGX agent

arXiv:2607.11844v2 Announce Type: replace Abstract: Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos inv

Building a Fully Local PDF Read-Aloud & PDF-to-Audiobook Desktop App with Kokoro 82M, Qwen, and llama.cpp

Model ReleasesDGX agent

Hey everyone, I’ve been building Speechfony - a desktop app for reading PDFs (and EPUBs) with offline text-to-speech. Open a document, listen sentence-by-sentence with highlighting, or export selected

CARE-Bench: Benchmarking Patient-Facing LLM Triage

Model ReleasesDGX agent

arXiv:2608.03731v1 Announce Type: new Abstract: Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the

ConlangBench: Exploring Language Knowledge and Learning in LLMs through Diverse Constructed Languages

Model ReleasesDGX agent

arXiv:2608.03505v1 Announce Type: new Abstract: Constructed languages (conlangs) are intentionally created human languages with a rich tradition of linguistic creativity. Despite their potential for s

DeepSeek V4 Flash 0731 at 10–17 t/s (nothink) on MacBook M5 Pro **64GB***, partly via SSD streaming

Model ReleasesDGX agent

Inspired by a post from u/giveen I motivated claude (no patinence on my side to work through everything myself) to help me get DS running on my MacBook M5 Pro 64GB and it exceeded my expectations.. be

EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation

ResearchDGX agent

arXiv:2608.02990v1 Announce Type: new Abstract: Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. Howev

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks

ResearchDGX agent

arXiv:2608.02966v1 Announce Type: new Abstract: Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scorin

Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release (Ivan Mehta/TechCrunch)

Model ReleasesDGX agent

Ivan Mehta / TechCrunch: Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release — Hark, a startup

HyVIC: A Metric-Driven Spatio-Spectral Hyperspectral Image Compression Architecture Based on Variational Autoencoders

Model ReleasesDGX agent

arXiv:2603.26468v2 Announce Type: replace Abstract: The rapid growth of hyperspectral data archives in remote sensing (RS) necessitates effective compression methods for storage and transmission. Rece

ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization

SafetyDGX agent

arXiv:2608.03210v1 Announce Type: new Abstract: Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift

Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation

Model ReleasesDGX agent

arXiv:2608.02639v1 Announce Type: cross Abstract: Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed tone at th

Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning

Model ReleasesDGX agent

arXiv:2608.03138v1 Announce Type: cross Abstract: Generating a rigorous paper introduction with large language models (LLMs) remains challenging, since it requires coordinating background, gap identif

Inverted Detection and Control in Steering Vectors

Model ReleasesDGX agent

arXiv:2608.02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.g., truthfulness) in large language model outputs. A key assumption un

Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks

Model ReleasesDGX agent

arXiv:2608.02621v1 Announce Type: cross Abstract: Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy fo

MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

Model ReleasesDGX agent

arXiv:2608.02613v1 Announce Type: cross Abstract: Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory bench

MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation

Model ReleasesDGX agent

arXiv:2608.03275v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capa

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification

Model ReleasesDGX agent

arXiv:2608.03474v1 Announce Type: new Abstract: Recent advances in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in web UI generation. However, existing benchmarks pre

MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding

Model ReleasesDGX agent

arXiv:2608.03708v1 Announce Type: new Abstract: Text-to-image diffusion models enable personalization of specific visual concepts from a small number of reference images. However, generating a single

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

Model ReleasesDGX agent

arXiv:2608.03885v1 Announce Type: new Abstract: Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-ti

Particle-based Generalised Stochastic Optimisation

Model ReleasesDGX agent

arXiv:2608.02844v1 Announce Type: cross Abstract: We develop a class of diffusion-based stochastic particle optimisation methods for loss functions with intractable gradients. Specifically, we conside

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud

SafetyDGX agent

arXiv:2608.03682v1 Announce Type: new Abstract: Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, a

Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds

Model ReleasesDGX agent

arXiv:2608.03135v1 Announce Type: cross Abstract: Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple concepts. We tra

S^3: Improving Agent Safety through Multi-Stage Defense

Model ReleasesDGX agent

arXiv:2608.02683v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish compl

Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC Generation

Model ReleasesDGX agent

arXiv:2608.02672v1 Announce Type: cross Abstract: Cloud misconfiguration remains a leading cause of security incidents, yet whether LLMs and SLMs can generate security-compliant Infrastructure-as-Code

SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA

Model ReleasesDGX agent

arXiv:2509.25459v2 Announce Type: replace Abstract: Large Language Models (LLMs) show promise in generating long-form scientific explanations that synthesize evidence and connect multiple factors. How

Skill libraries are shipping in agent harnesses on the assumption that writing skills down compounds. A new benchmark tests that directly. C…

Model ReleasesDGX agent

Skill libraries are shipping in agent harnesses on the assumption that writing skills down compounds. A new benchmark tests that directly. ContinualSkillBench covers five domains, each with 100 interc

Stable Diffusion might actually be remembered in the history books, and I don’t think that’s an overstatement

Model ReleasesDGX agent

Hear me out before you roll your eyes. We tend to only recognize turning points in hindsight. Nobody in 1993 thought the Mosaic browser would be a history book moment, but the web is. I think Stable D

The Geometric Nature and a Free Proxy for Flow-Matching Uncertainty

ResearchDGX agent

arXiv:2607.27933v2 Announce Type: replace Abstract: Flow matching (FM) has become a popular action head paradigm for modern embodied models. However, as a conditional generative model, it does not exp

The Ignition Is Real, and It Lives at the Readout: Latent composition, difficulty-clocked ignition, and the interface-constituted commit in a recurrent-depth reasoner

Model ReleasesDGX agent

arXiv:2608.03263v1 Announce Type: cross Abstract: We test whether the 'compositional ignition' reported in latent-reasoning models is real computation, an instrument artifact, or inherited from verbal

Thinking of buying more DRAM right now...

Model ReleasesDGX agent

So I'm looking at https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF and I realize my 128GB of DRAM just isn't cutting it for this (incredibly powerful) model. If only I had another 64GB, I th

Traceable Multi-Agent System for Knowledge-Based Forecasting

AgentsDGX agent

arXiv:2608.03339v1 Announce Type: new Abstract: Enterprise forecasting increasingly relies on autonomous agents that interpret documents, search for data, generate code, and revise models. While this

TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows

Model ReleasesDGX agent

arXiv:2608.02680v1 Announce Type: cross Abstract: Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure with retrie

UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks

Model ReleasesDGX agent

arXiv:2608.03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragment

VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space

Model ReleasesDGX agent

arXiv:2608.02878v1 Announce Type: new Abstract: Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% accuracy on stan

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

Model ReleasesDGX agent

arXiv:2608.03994v1 Announce Type: new Abstract: We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes

WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament

Model ReleasesDGX agent

arXiv:2608.04008v1 Announce Type: new Abstract: Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewher

Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure

Model ReleasesDGX agent

arXiv:2608.02657v1 Announce Type: cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts

4 Aug 2026

A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use

Model ReleasesDGX agent

arXiv:2608.00218v1 Announce Type: new Abstract: Agentic LLMs exhibit three consequential tool-use failures: invalid arguments (validity), unnecessary calls (over-calling), and omitted calls when tools

A Large-Scale Multi-Dimensional Empirical Study of LLMs for Conversation Summarization

Model ReleasesDGX agent

arXiv:2606.15974v2 Announce Type: replace Abstract: Despite the significant advancement of LLMs in conversation summarization, their evaluation remains limited by insufficient scenarios, input lengths

← Previous
1…379380381382383…1050
Next →