AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

86,457Total entries
1Added by human
86,456Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,039 results
5 Aug 2026

Route-Align-Verify for Functional Correctness in Code Generation

Model ReleasesDGX agent

arXiv:2608.03341v1 Announce Type: cross Abstract: Large language models (LLMs) have substantially improved code generation, yet achieving strong functional correctness remains difficult, especially fo

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

Model ReleasesDGX agent

arXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is f

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference

SafetyDGX agent

arXiv:2608.03335v1 Announce Type: new Abstract: Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. Th

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics

Local AiDGX agent

arXiv:2608.03291v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model's reasoning

UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space

ResearchDGX agent

arXiv:2608.03817v1 Announce Type: cross Abstract: Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions

Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't

Model ReleasesDGX agent

arXiv:2608.02829v1 Announce Type: new Abstract: Model families train every size from scratch. Can a pretrained large model be converted into a smaller sibling? We characterize the 1.4B->410M conversio

4 Aug 2026

A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2 & Nano Banana2

Model ReleasesDGX agent

arXiv:2608.01258v1 Announce Type: new Abstract: The realism of images generated by multimodal large language models (MLLMs), such as GPT Image2 and Nano Banana2, has improved rapidly in recent years.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction

Model ReleasesDGX agent

arXiv:2608.00434v1 Announce Type: new Abstract: Multi-Token Prediction (MTP) has emerged as an effective paradigm that augments a shared Large Language Model backbone with auxiliary heads, training th

Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

SafetyDGX agent

arXiv:2608.02446v1 Announce Type: cross Abstract: Relevance evaluation plays a crucial role in personalized search systems, serving as a guardrail alongside user engagement metrics to ensure that sear

Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study

Model ReleasesDGX agent

arXiv:2608.01865v1 Announce Type: new Abstract: Automatic speech recognition (ASR) performance degrades sharply on dysarthric speech, yet how disordered articulation reshapes a model's internal repres

Beyond Global Latents: Chunk-Based Sparse Grid VAE for Scalable 3D Modeling

Local AiDGX agent

arXiv:2608.02016v1 Announce Type: new Abstract: Sparse voxel grids preserve the spatial structure needed for detailed 3D reconstruction, but their memory still grows rapidly with resolution as active

CADENA: Stepwise CAD Reverse Engineering

Model ReleasesDGX agent

arXiv:2608.00799v1 Announce Type: new Abstract: Computer-Aided Design (CAD) underpins modern engineering, yet converting existing shapes into editable models still demands substantial expert effort. M

Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking

ResearchDGX agent

arXiv:2608.01021v1 Announce Type: cross Abstract: In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples coul

Can Language Models Identify Shadow Trading Targets? An NLP Evaluation of SEC Enforcement Theory

ResearchDGX agent

arXiv:2608.01322v1 Announce Type: new Abstract: Shadow trading -- trading in a peer firm's securities on the basis of material nonpublic information (MNPI) about an 'economically linked' company -- is

Don't Judge a Book by its Cover: Testing LLMs' Robustness Under Logical Obfuscation

Model ReleasesDGX agent

arXiv:2602.01132v2 Announce Type: replace Abstract: Tasks such as solving arithmetic equations, evaluating truth tables, and completing syllogisms are handled well by large language models (LLMs) in t

EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next

Model ReleasesDGX agent

arXiv:2603.12147v2 Announce Type: replace Abstract: Egocentric video provides a natural modality for studying human behavior, but conventional visual understanding captures mainly observable scenes, o

FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing?

Model ReleasesDGX agent

arXiv:2608.00909v1 Announce Type: new Abstract: Can large language models generate not just correct, but fast hardware? This paper investigates the question in financial FPGA design, where 5-10 nanose

From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States

Model ReleasesDGX agent

arXiv:2607.09842v2 Announce Type: replace-cross Abstract: We investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the hidden-state tra

Generate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 Super

HardwareDGX agent

NVIDIA Alpamayo 2 Super is a publicly available 34‑billion‑parameter vision–language–action model that merges a 32‑B Cosmos 3 Super Reasoner with a 2‑B action‑expert diffusion network. It produces uni

GIFT: Geometry-Invariant Fine-Tuning for Non-Lambertian Monocular Depth Estimation

Model ReleasesDGX agent

arXiv:2608.02068v1 Announce Type: new Abstract: Monocular depth foundation models, benefiting from large-scale synthetic training data, have demonstrated strong generalization. However, they often hal

HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing

SafetyDGX agent

arXiv:2603.15257v2 Announce Type: replace Abstract: Tactile sensing is a crucial capability for Vision-Language-Action (VLA) architectures, as it enables dexterous and safe manipulation in contact-ric

HyperODE: Zero-Shot Surrogate for Simulation and Inference of Dynamical Systems

Model ReleasesDGX agent

arXiv:2608.00852v1 Announce Type: new Abstract: Understanding and controlling complex dynamical systems often requires executing thousands of numerical simulations across vast parametric landscapes, w

I built a DwarfStar-inspired Vulkan/Metal inference engine for Qwen3.6-35B-A3B on 16 GB machines

Model ReleasesDGX agent

Disclosure: I’m the author and maintainer of QuarkStar. I built QuarkStar, a small native inference engine inspired by Antirez’s DwarfStar. QuarkStar currently supports: Qwen3.6-35B-A3B, using the sam

It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling

Model ReleasesDGX agent

arXiv:2608.01207v1 Announce Type: new Abstract: Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe transfers poor

Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention

Model ReleasesDGX agent

arXiv:2602.04789v4 Announce Type: replace Abstract: Advanced autoregressive (AR) video generation models have improved visual fidelity and interactivity, but the quadratic complexity of attention rema

MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving

Model ReleasesDGX agent

arXiv:2608.02449v1 Announce Type: new Abstract: Deploying vision-language models (VLMs) for safety-critical spatial reasoning on resource-constrained autonomous driving platforms requires both compact

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

Model ReleasesDGX agent

arXiv:2608.02372v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in task-oriented dialogue systems that support multi-step decision-making in high-stakes domains

Proteus: A Truncation-Robust Entropy Model for Progressive LiDAR Compression

SafetyDGX agent

arXiv:2608.00687v1 Announce Type: new Abstract: LiDAR point clouds provide explicit, deterministic physical boundaries critical for collaborative safety-critical perception. However, wireless channels

Qwen3.8-Max is available in Hermes Agent now! Let's build! 🚀🚀

Model ReleasesDGX agent

Qwen 3.8‑Max, Alibaba.Qwen’s latest large‑language model, has been added to Hermes Agent and can currently be accessed at a 20 % discount. The update aims to streamline integration of the model for de

Refine Drugs, Don't Complete Them: Uniform-Source Discrete Flows for Fragment-Based Drug Discovery

Model ReleasesDGX agent

arXiv:2509.26405v2 Announce Type: replace Abstract: We introduce InVirtuoGen, a discrete flow generative model for fragmented SMILES for de novo and fragment-constrained generation, and target-propert

Shieldstral is available under Apache 2.0. Try it: https://huggingface.co/mistralai/Shieldstral-1.0

Model ReleasesDGX agent

**Shieldstral** is a 3‑billion‑parameter open‑weights model developed by Mistral AI for content safety. It can be deployed on-device and is distributed under the Apache 2.0 license. The model is avail

Similarity-Aware Machine Unlearning

Model ReleasesDGX agent

arXiv:2608.00246v1 Announce Type: new Abstract: Machine unlearning removes the influence of user-specified training examples from a trained model, avoiding the need to retrain it from scratch. Localiz

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

Model ReleasesDGX agent

arXiv:2603.08091v2 Announce Type: replace Abstract: Large language model (LLM)-based judges are widely adopted for automated evaluation and reward modeling, yet their judgments are often affected by j

TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

Model ReleasesDGX agent

arXiv:2608.00640v1 Announce Type: new Abstract: Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high

Who Should Be Generated? Justifying Demographic Targets in Open-Ended Generation

SafetyDGX agent

arXiv:2608.02551v1 Announce Type: cross Abstract: Fairness evaluation concerns not only what a model produces, but also what its outputs ought to be compared against. When a model generates 'a CEO in

3 Aug 2026

A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation

ResearchDGX agent

arXiv:2607.29077v1 Announce Type: new Abstract: Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input required to obt

Agentic Harness for Real-World Compilers

Model ReleasesDGX agent

arXiv:2603.20075v2 Announce Type: replace-cross Abstract: Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements enable

b10237

Model ReleasesDGX agent

llama : MTP support for DeepSeek V3.2 (#26457) llama : MTP support for DeepSeek V3.2 model : no need to include MTP layers during DeepSeek V3.2 model type discovery Co-authored-by: Stanisław Szymczyk

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

Model ReleasesDGX agent

arXiv:2607.28801v1 Announce Type: cross Abstract: Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial

Benchmarks Are Not Validation: A System-Level View of Financial LLM Applications

Model ReleasesDGX agent

arXiv:2607.28840v1 Announce Type: new Abstract: Large language models are increasingly deployed in financial applications that combine retrieval, proprietary data, tool use, orchestration logic, monit

Can Synthetic Data Overcome the Generalization Limits of AI-Based Flower and Pod Detection Across Cowpea Breeding Genotypes and Environments?

ResearchDGX agent

arXiv:2607.28796v1 Announce Type: new Abstract: High-throughput phenotyping requires AI-enabled computer vision models that generalize across genotypes, locations, and growing seasons, yet such models

HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution

AgentsDGX agent

arXiv:2607.13683v2 Announce Type: replace Abstract: Large Language Models (LLMs) have enabled capable agents across diverse applications. Beyond the foundation model, the performance of an agent is go

I gave five different local LLMs a town. They invented Facebook and a duck-based credit bureau. (MIT, self-hosted, you don't play it — you watch it)

Model ReleasesDGX agent

Each villager in Pepperton is a different model — a mistral, a qwen3, a qwen2.5, a phi4-mini, a llama3.2 — because model families have genuinely different temperaments, and the friction between them i

LARA: Lightweight Adapters in the Residual Stream for Composable Adaptation and Alignment

Model ReleasesDGX agent

arXiv:2607.28669v1 Announce Type: new Abstract: We present LARA (Lightweight Additive Residual Adaptation), a method for efficient adaptation that operates in the residual stream of a frozen model rat

🔥Let's talk about Qwen! #AMA

Model ReleasesDGX agent

The post announces an “Ask Me Anything” (AMA) about Qwen, the Alibaba‑developed foundation model. It introduces the Qwen Foundation Model Team and directs readers to their GitHub repository @QwenDevs,

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

Local AiDGX agent

arXiv:2512.03438v3 Announce Type: replace Abstract: Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally opt

Pay for The Second-Best Service: A Game-Theoretic Approach Against Dishonest LLM Providers

ApplicationsDGX agent

arXiv:2511.00847v5 Announce Type: replace-cross Abstract: The widespread adoption of Large Language Models (LLMs) through Application Programming Interfaces (APIs) induces a critical vulnerability: th

Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters

Local AiDGX agent

arXiv:2607.29238v1 Announce Type: cross Abstract: InMyStyle is a privacy first, single user system that adapts small language models to rewrite AI-edited text towards an individual user's writing styl

So-Fake: Benchmarking and Explaining Social Media Image Forgery Detection

Model ReleasesDGX agent

arXiv:2505.18660v5 Announce Type: replace Abstract: Recent advances in AI-powered generative models have enabled the creation of increasingly realistic synthetic images, posing significant risks to in

WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics

Model ReleasesDGX agent

arXiv:2601.02430v3 Announce Type: replace-cross Abstract: Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and com

2 Aug 2026

Parlor v2: best-effort fully local GPT-Live clone on an M3 Pro

Model ReleasesDGX agent

GPT-Live is so good that I use it almost every day. I've been wanting to replicate it since it was released. My first attempt was to fine-tune Gemma 4 12B to behave like a full-duplex model. Something

Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG

Model ReleasesDGX agent

Hey y'all. I'll be concise. TL;DR: DS V4-Flash-0731 @ UD-IQ2_M running fully in VRAM on 3xMI50s (90.9 GB model, 96 GB VRAM). Actual speed on llama-server is: - Text Generation: ~15-16 tokens/second st

Try handling complex tasks to your local models with GraphARC, graph engineering yes !

Local AiDGX agent

🚀 We just built our first real-time implementation of Graph Engineering, inspired by our experience building graph tooling used by 4,000+ developers. 🔗 Repo: https://github.com/CodeGraphContext/grapha

Why are almost all new benchmarks and leaderboards coding focused?

Model ReleasesDGX agent

I know in in this community LLM's are generally used for coding but there are other usecases besides coding and those usecases should be tested too. I also know benchmarks can sometimes be benchmaxxed

31 Jul 2026

ARES: Anomaly Recognition Model For Edge Streams

ApplicationsDGX agent

arXiv:2511.22078v2 Announce Type: replace Abstract: Many real-world scenarios involving streaming information can be represented as temporal graphs, where data flows through dynamic changes in edges o

Benchmarking LLM Competence on Logical Inference over Probability Operators

Model ReleasesDGX agent

arXiv:2607.27405v1 Announce Type: new Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

Model ReleasesDGX agent

arXiv:2607.27816v1 Announce Type: new Abstract: Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversation

Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs

ResearchDGX agent

arXiv:2607.27700v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) suffer from prohibitive inference overhead due to long sequences of visual tokens. However, existing visual token re

Challenges in annotations by humans and LLMs: A case study of evaluative language

ApplicationsDGX agent

arXiv:2607.28119v1 Announce Type: new Abstract: In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find

Critical attention scaling in long-context transformers

Model ReleasesDGX agent

arXiv:2510.05554v2 Announce Type: replace Abstract: As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity

← Previous
1…271272273274275…1034
Next →