AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
AllBlog
90,914Total entries
1Added by human
90,913Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,676 results
Research

GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks

DGX agent

Systematic failures of vision models on semantically coherent subsets, known as error slices, reveal limitations in robustness and evaluation. Existing slice discovery approaches largely model slices

researchapple-ml-research
27 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

I ran the 35B agentic comparison someone asked for (stock vs Ornith vs KAT-Coder, 120 runs)

DGX agent

Someone in the comments of my 27B post-train bakeoff asked for the 35B version, so I ran it. Same setup as last time: fresh Coder workspaces on my k8s cluster, each driving my own agent (Hermes) headl

model-releasesr-localllama
27 Jul 2026
Model Releases

IR275K: A Benchmark for Infrared Multi-Frame Super-Resolution Toward Efficient Remote Sensing

DGX agent

arXiv:2607.22380v1 Announce Type: new Abstract: Efficient processing is becoming increasingly important in infrared remote sensing, where satellite constellations produce large volumes of observations

model-releasesarxiv-cs-cv
27 Jul 2026
Model Releases

Kimi K3 weights drop today. We're deploying on A100s, H200s and B300s this week and the A100 math is already rough

DGX agent

tldr; we are going to host K3 on A100s (yes, thats correct, we'll try to see if it holds up), H200s & B300s - expect results for A100s & H200s this week while we setup the B300 cluster this weekend &

model-releasesr-localllama
27 Jul 2026
Model Releases

Ling-3.0-flash weights: SGLang says day-0, vLLM says when they land, llama.cpp closed the 2.6 request as not_planned

DGX agent

Some Ling-3.0-flash threads here last week ended on the same two questions with no real answer, so I went through the repos. State as of writing, with links so you can check instead of taking my word

model-releasesr-localllama
27 Jul 2026
Research

Moving Beyond Diversity: Visual Token Pruning as Subspace Reconstruction for Efficient VLMs

DGX agent

arXiv:2606.18681v2 Announce Type: replace Abstract: Despite their remarkable performance, Vision Language Models (VLMs) incur substantial computational overhead due to the large number of visual token

researcharxiv-cs-cv
27 Jul 2026
Safety

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode

DGX agent

arXiv:2607.22083v1 Announce Type: cross Abstract: We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-

safetyarxiv-cs-cl
27 Jul 2026
Model Releases

NWaaS: A Non-Intrusive and Privacy-Preserving Watermarking-as-a-Service System with Adaptive Resource Scheduling

DGX agent

arXiv:2507.18036v2 Announce Type: replace-cross Abstract: Securing intellectual property (IP) in Machine Learning as a Service is critical yet challenging. While deep neural network watermarking serve

model-releasesarxiv-cs-cv
27 Jul 2026
Model Releases

One Hand Watches The Other: Dynamic Multi-Agent Cooperation for Sample-Efficient Bimanual Manipulation in Dynamic Environments

DGX agent

arXiv:2607.22119v1 Announce Type: cross Abstract: Multi-stream robot manipulation policies achieve unparalleled sample efficiency and generalization by modeling actions relative to environmental refer

model-releasesarxiv-cs-lg
27 Jul 2026
Safety

Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning

DGX agent

arXiv:2507.01551v3 Announce Type: replace-cross Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~

safetyarxiv-cs-cl
27 Jul 2026
Model Releases

Has anyone compared pre-training, SFT/LoRA and reinforcement post-training on Qwen3.6-27B?

DGX agent

Qwen3.6-27B: SFT vs continued pre-training vs RL? I’m interested in adapting Qwen3.6-27B, but I’m increasingly unsure whether conventional SFT/LoRA is the best route if the goal is to add a capability

model-releasesr-localllama
26 Jul 2026
Model Releases

5700, 48GB RAM and a 3090 24Gb. Best OS and framework/model?

DGX agent

Hi everyone, I have a system which I have been using for gaming, R7 5700X, 48GB DDR4, RTX3090 24GB. But I want to use it for Local AI to reduce my reliance on cloud AI providers (mainly usage limits -

model-releasesr-localllama
25 Jul 2026
Model Releases

Getting a second GPU in addition to my RTX3090

DGX agent

Hello, I've been learning how to use local LLMs for a year or so on my workstation, using a RTX3090. Current setup : - i5 12400 - 64gb RAM - RTX 3090 - OS : Fedora KDE workstation I'm using LMStudio t

model-releasesr-localllama
25 Jul 2026
Model Releases

The Opus 5 system card itself is a fun PDF to parse. It's 193 pages and stacked with labeled and unlabeled charts 📊 LlamaParse does a surpr…

DGX agent

The Opus 5 system card itself is a fun PDF to parse. It's 193 pages and stacked with labeled and unlabeled charts 📊 LlamaParse does a surprisingly good job on agentic (1.25c per page) and agentic plus

model-releasesjerry-liu--x
25 Jul 2026
Model Releases

A reminder this is 20% off via the Nous Portal <3

DGX agent

mr‑r0b0t announced that Claude Opus 5 is available at a 20 % discount through the Nous Portal. The model can be accessed via the Hermes Agent on the Nous Portal, as well as through OpenRouter and Anth

model-releasesnous-research--x
24 Jul 2026
Model Releases

AI Assistants Overassist

DGX agent

arXiv:2607.21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance from AI assis

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs

DGX agent

arXiv:2607.20479v1 Announce Type: new Abstract: Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail

model-releasesarxiv-cs-ai
24 Jul 2026
Safety

Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

DGX agent

arXiv:2607.21558v1 Announce Type: new Abstract: Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy

safetyarxiv-cs-ai
24 Jul 2026
Model Releases

CAMeR: Keyword-Gated Hybrid Activation for Adaptive Memory Retention in LLM Agents

DGX agent

arXiv:2607.20458v1 Announce Type: cross Abstract: Large language model (LLM) agents operating over extended dialogues accumulate vast amounts of information, yet existing memory systems either retain

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

DGX agent

arXiv:2607.21155v1 Announce Type: cross Abstract: Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiri

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making

DGX agent

arXiv:2607.20491v1 Announce Type: new Abstract: Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time. We i

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention

DGX agent

arXiv:2607.20457v1 Announce Type: cross Abstract: Inference with large language models (LLMs) on long sequences is computationally expensive due to the quadratic complexity of self-attention. Distribu

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1

DGX agent

arXiv:2607.20589v1 Announce Type: new Abstract: Persona simulation involves utilizing large language models (LLMs) to anticipate human choices or interactions based on specific characteristic informat

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries

DGX agent

arXiv:2607.20757v1 Announce Type: cross Abstract: Transformers are known to have internal continuous symmetries that leave outputs invariant, while modifying quantization. GaugeQuant leverages this in

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Geometric Configurations of Perturbed Jailbreak Prompts

DGX agent

arXiv:2607.20581v1 Announce Type: cross Abstract: Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat

model-releasesarxiv-cs-ai
24 Jul 2026
Safety

GOAT

DGX agent

GOAT For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safe

safetyjim-fan--x
24 Jul 2026
Model Releases

Gradient Concentration, Not Weight Saliency, Explains Representation-Level Class Unlearning

DGX agent

arXiv:2607.21353v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of specific training data while preserving model utility. Many state-of-the-art approaches pursue this g

model-releasesarxiv-cs-lg
24 Jul 2026
Agents

having an event stream like activegraph is the base of how we move to next level and that will probably be composition, which can let you ac…

DGX agent

having an event stream like activegraph is the base of how we move to next level and that will probably be composition, which can let you achieve better results with smaller models. - get the core eve

agentsyohei-nakajima--x
24 Jul 2026
Model Releases

Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought

DGX agent

arXiv:2607.20427v1 Announce Type: cross Abstract: Mixture-of-Experts architectures have revolutionized scaling, yet the underlying logic of their routing remains a black box. In this paper, we uncover

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation

DGX agent

arXiv:2607.20494v1 Announce Type: new Abstract: Production LLM applications commonly stack a regex filter in front of model-side alignment; prior work found no measurable coverage gain from adding a l

model-releasesarxiv-cs-ai
24 Jul 2026
Research

LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning

DGX agent

Long-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algorithmic puzzles, we demonstrate that while decomposit

researchapple-ml-research
24 Jul 2026
Research

Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts

DGX agent

arXiv:2607.20462v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated output

researcharxiv-cs-ai
24 Jul 2026
Hardware

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

DGX agent

arXiv:2607.20709v1 Announce Type: new Abstract: Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agen

hardwarearxiv-cs-ai
24 Jul 2026
Model Releases

Open Source Tax Engine outperforming fable 5 and gpt sol

DGX agent

This is an open source and free tax engine which scored 96% on TaxCalcBench [highest ever recorded score till date] surpassing fable 5 and sol with just sonnet 5. The only 2 cases where it missed, it

model-releasesr-ollama
24 Jul 2026
Model Releases

Opus 5 now available in Hermes Agent

DGX agent

Claude Opus 5 is now released in the Hermes Agent, a product of Nous Research and Teknium. Users can access the model through multiple gateways, including the Nous Portal, OpenRouter, and Anthropic Di

model-releasesnous-research--x
24 Jul 2026
Model Releases

PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

DGX agent

arXiv:2607.20482v1 Announce Type: new Abstract: Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspeci

model-releasesarxiv-cs-ai
24 Jul 2026
Safety

PILD: Physics-Informed Learning via Diffusion

DGX agent

arXiv:2601.21284v2 Announce Type: replace-cross Abstract: Diffusion models have emerged as powerful generative tools for modeling complex data distributions, yet their purely data-driven nature limits

safetyarxiv-cs-ai
24 Jul 2026
Model Releases

PromptPack: Scaling LLM Annotation Agents for Online Recommendation

DGX agent

arXiv:2607.20528v1 Announce Type: new Abstract: Online recommendation platforms increasingly use Large Language Models (LLMs) to extract structured features from ad creatives. While deploying a single

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

RE-AD: Real-Time Requirement Adherence for Data Labeling

DGX agent

arXiv:2607.20455v1 Announce Type: cross Abstract: Human-annotated data remains fundamental to training frontier Large Language Models (LLMs). However, crowd-sourced annotations often suffer from quali

model-releasesarxiv-cs-ai
24 Jul 2026
Safety

Robust Critics: Defending LLMs Against Multi-Turn Attacks

DGX agent

arXiv:2607.20472v1 Announce Type: new Abstract: When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the c

safetyarxiv-cs-ai
24 Jul 2026
Local Ai

Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating

DGX agent

arXiv:2607.20481v1 Announce Type: new Abstract: Local-cloud collaboration is a practical way to deploy large language models under resource constraints, but existing methods often rely on trained rout

local-aiarxiv-cs-ai
24 Jul 2026
Model Releases

Scene Parameter Saliency via Differentiable Light Transport

DGX agent

arXiv:2607.21562v1 Announce Type: new Abstract: Gradient-based saliency methods reveal which input features most influence a neural network's output, and are a standard tool for model interpretability

model-releasesarxiv-cs-cv
24 Jul 2026
Tutorials

SCoPE: Shift-Aware Speaker-Conditioned Priors for Emotion Recognition in Conversations

DGX agent

arXiv:2607.20445v1 Announce Type: new Abstract: In conversations, human emotions are transient; however, they tend to persist across multiple utterances. For example, we rarely switch instantly betwee

tutorialsarxiv-cs-cl
24 Jul 2026
Applications

SoccerSynth Field: enhancing field detection with synthetic data from virtual soccer simulator

DGX agent

arXiv:2503.13969v2 Announce Type: replace Abstract: Field detection in team sports is an essential task in sports video analysis. However, collecting large-scale and diverse real-world datasets for tr

applicationsarxiv-cs-cv
24 Jul 2026
Local Ai

Spent two weeks on a kernel that benchmarked 29x faster. End to end it's maybe 6-10%, and it's not even wired in yet.

DGX agent

I've been building a C99 inference engine from scratch (no Python, no BLAS, just gcc and make) that runs BitNet's ternary models on CPU. A few weeks ago I got obsessed with the matmul kernel - wrote a

local-air-localllama
24 Jul 2026
Model Releases

StabilityBench: Benchmarking Instability in LLMs

DGX agent

arXiv:2607.20558v1 Announce Type: cross Abstract: AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poor

model-releasesarxiv-cs-ai
24 Jul 2026
Safety

Stay hungry, stay foolish. Absolute legend

DGX agent

Stay hungry, stay foolish. Absolute legend For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by ever

safetyjim-fan--x
24 Jul 2026
Model Releases

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs

DGX agent

arXiv:2509.14257v3 Announce Type: replace-cross Abstract: Large Language Model agents achieve strong performance on multi-step reasoning and tool-use tasks, but their impressive capabilities typically

model-releasesarxiv-cs-ai
24 Jul 2026
← Previous
1…504505506507508…1369
Next →