AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,993
  • Agents7,449
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,136
  • Local Ai4,859
  • Model Releases23,375
  • Research19,835
  • Safety13,176
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

86,993Total entries
1Added by human
86,992Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,477 results
11 Aug 2026

VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging

Model ReleasesDGX agent

arXiv:2511.18121v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visu

XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment

Model ReleasesDGX agent

arXiv:2608.09892v1 Announce Type: new Abstract: Robot policy evaluation and deployment remain fragmented by model-specific software dependencies, data representations, and runtime interfaces, so that

Your Prompt Is Not the Only Prompt: How Much Do LLMs Weight Structured-Output Schema Descriptions?

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.08254v1 Announce Type: new Abstract: Structured output, where an LLM populates a predefined JSON schema, has become a default mechanism for data labeling and information extraction, but it

ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling

Local AiDGX agent

arXiv:2608.07974v1 Announce Type: new Abstract: Large language model (LLM) fine-tuning at the edge adapts the model to scenario-specific data while preserving privacy. Although existing studies propos

10 Aug 2026

Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory

Model ReleasesDGX agent

arXiv:2608.07169v1 Announce Type: new Abstract: Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which strug

An End-to-End Agent Auditing Engine

AgentsDGX agent

arXiv:2608.07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of d

AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation

Model ReleasesDGX agent

arXiv:2603.20986v2 Announce Type: replace Abstract: Phase-field modeling links thermodynamics and kinetics to microstructural evolution, but multiphysics frameworks such as MOOSE require expertise to

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

Model ReleasesDGX agent

arXiv:2608.06930v1 Announce Type: new Abstract: Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main

Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation

Model ReleasesDGX agent

arXiv:2608.06751v1 Announce Type: cross Abstract: Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical

Coupling Planning with Episodic Memory in LLM Agents for Software Issue Resolution

Model ReleasesDGX agent

arXiv:2608.06811v1 Announce Type: cross Abstract: Resolving a real software issue with a large language model (LLM) agent is a long repair episode, often tens to hundreds of steps spanning exploration

CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training

Model ReleasesDGX agent

arXiv:2608.06471v1 Announce Type: cross Abstract: Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in real-world s

DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX Sparks

Model ReleasesDGX agent

Having a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getti

FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding

Local AiDGX agent

arXiv:2608.06819v1 Announce Type: cross Abstract: Token-level collaboration allows a large language model (LLM) to assist a small language model (SLM) when their predictions diverge. Existing methods

Google named a Leader in The Forrester Wave™: AI Platforms, Q3 2026

Model ReleasesDGX agent

At Google Cloud, we help organizations of all sizes build and operationalize complex agentic workflows with total confidence. By combining world-class AI research with an open, fully integrated AI pla

Human-AI Perceptual Alignment by Playing Hues and Cues

Local AiDGX agent

arXiv:2608.07141v1 Announce Type: new Abstract: Evaluating the perceptual alignment between Contrastive Vision-Language Models (CVLMs) and humans is typically constrained by traditional benchmarks tha

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

Model ReleasesDGX agent

arXiv:2608.07417v1 Announce Type: cross Abstract: Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simplified video-

IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents

Model ReleasesDGX agent

arXiv:2608.06735v1 Announce Type: new Abstract: Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as

inclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face

Model ReleasesDGX agent

Looks like the Ling team open weighted a much smaller version of the Ling-3.0-flash they open weighted a few days ago. It's 8B params with 1.3B active, and seems to fall between the 4B and 8-12B Qwen

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin

TutorialsDGX agent

arXiv:2608.06411v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of

Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework

Model ReleasesDGX agent

arXiv:2608.06909v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, a

Modular TTT: Rethinking Test-Time Training as Composable Modules

Model ReleasesDGX agent

arXiv:2608.07110v1 Announce Type: cross Abstract: Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite

MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs

Model ReleasesDGX agent

arXiv:2607.08970v2 Announce Type: replace-cross Abstract: Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate observa

PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks

Model ReleasesDGX agent

arXiv:2505.04397v3 Announce Type: replace-cross Abstract: Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexplore

Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation

SafetyDGX agent

arXiv:2608.07176v1 Announce Type: cross Abstract: Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training

Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation

SafetyDGX agent

arXiv:2608.07154v1 Announce Type: cross Abstract: Open-source robotics and foundation models have lowered the barrier to embodied AI, yet language-guided laboratory automation still requires reliable

RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs

Model ReleasesDGX agent

arXiv:2608.07088v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing trai

Simple-OPD: Demystifying Warm-up for On-policy Distillation

Model ReleasesDGX agent

arXiv:2608.06802v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend str

Skaling: Chinchilla's Exponents Meet Kaplan's Coupling

ResearchDGX agent

arXiv:2608.07222v1 Announce Type: new Abstract: Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarc

Suppress and Diversify: Refining Robust Pathways for Corruption Robustness

Model ReleasesDGX agent

arXiv:2608.06712v1 Announce Type: new Abstract: Model robustness against natural image corruptions is essential for safety-critical applications. While existing methods primarily focus on implicit rep

YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family

Model ReleasesDGX agent

arXiv:2608.07051v1 Announce Type: new Abstract: Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-time detectors, whose heterogeneous op

9 Aug 2026

DeepSeek-V4-Flash-0731 Q8_K_XL sometimes stops mid-task in OpenCode - anyone else seeing this?

Model ReleasesDGX agent

Hey everyone, I've been experimenting with the new DeepSeek-V4-Flash-0731 release locally using the Unsloth Studio Q8_K_XL GGUF with OpenCode. Overall, it's been working really well, but I've noticed

8 Aug 2026

enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think

Model ReleasesDGX agent

Disclaimer - no LLM was used to write this post/note As larger post about my setup will come later, want to give heads-up to folks who use VLLM and >= 2 GPUs. So I have pretty meaty server (8 channel

Qwen3.6 27B + 35B on vLLM, single R9700 (gfx1201)

Model ReleasesDGX agent

I've been tuning my new Radeon AI Pro R9700, and figured that this would be useful information for people who are trying to optimise their setups. I'm pretty happy with these results and looking forwa

7 Aug 2026

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval

Model ReleasesDGX agent

arXiv:2608.05260v1 Announce Type: new Abstract: Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from d

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset

Model ReleasesDGX agent

arXiv:2608.06165v1 Announce Type: cross Abstract: Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

Model ReleasesDGX agent

arXiv:2608.06312v1 Announce Type: new Abstract: Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficie

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation

ResearchDGX agent

arXiv:2608.05999v1 Announce Type: new Abstract: Vision-language-action (VLA) models have demonstrated remarkable capabilities in robotic manipulation by leveraging pretrained vision-language models. H

ChronoVision: Temporal Reasoning via Latent State Reconstruction

Model ReleasesDGX agent

arXiv:2608.05631v1 Announce Type: new Abstract: Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. T

CNM-BERT: A Drop-In Structural Embedding for Chinese Characters via Ideographic Description Sequences

Model ReleasesDGX agent

arXiv:2608.05167v1 Announce Type: new Abstract: Token-based encoders like BERT treat Chinese characters as atomic identifiers, ignoring their recursive orthographic structure. Consequently, models rel

DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimization

Model ReleasesDGX agent

arXiv:2608.00641v2 Announce Type: replace Abstract: Bayesian optimization (BO) relies on a surrogate model and an acquisition function, yet the most suitable choices vary across tasks and optimization

Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case

Model ReleasesDGX agent

arXiv:2608.06075v1 Announce Type: cross Abstract: Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural

DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory Graph

Model ReleasesDGX agent

arXiv:2608.05170v1 Announce Type: cross Abstract: Role-playing agents (RPAs) have emerged as a key application of large language models, enabling immersive and high-fidelity character simulation. Accu

Energy-Guided Flow Matching

Model ReleasesDGX agent

arXiv:2608.05811v1 Announce Type: new Abstract: Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dim

EschaLabs/Qwen3.6-35B-A3B-Escha-W2 · Hugging Face

Model ReleasesDGX agent

Hey peeps. I know you're tired of low quants giving hard to believe numbers. I'm quite skeptical too and from what I tried I'm often left with the impression that the claims fall short. So this model

Evidential Rule Learning for Interpretable Classification with Abstention

Model ReleasesDGX agent

arXiv:2608.05859v1 Announce Type: cross Abstract: Interpretable classification often requires more than accurate predictions for real-life deployment: models should be transparent about the evidence b

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Model ReleasesDGX agent

arXiv:2608.05747v1 Announce Type: new Abstract: Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overloo

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

Model ReleasesDGX agent

arXiv:2608.06301v1 Announce Type: new Abstract: As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts,

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging

Model ReleasesDGX agent

arXiv:2608.05541v1 Announce Type: new Abstract: Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning. However,

Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation

Model ReleasesDGX agent

arXiv:2608.05210v1 Announce Type: cross Abstract: Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by

IPV-Bench: Benchmarking Image Protection Methods under Diverse Image-to-Video Generation Scenarios

Model ReleasesDGX agent

arXiv:2603.26154v2 Announce Type: replace Abstract: Image-to-video (I2V) generation models can be misused to animate a single image into a convincing fake video, motivating perturbation-based image pr

MoCA: Implicit Social Context Analysis

Model ReleasesDGX agent

arXiv:2608.05825v1 Announce Type: new Abstract: Human social communication, such as affection and intent, is often conveyed in highly implicit ways, where underlying meanings are expressed through ind

NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering

Model ReleasesDGX agent

arXiv:2608.06292v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves question answering by grounding large language models (LLMs) in external knowledge such as text corpora. H

nnMIL: A generalizable multiple instance learning framework for computational pathology

ApplicationsDGX agent

arXiv:2511.14907v2 Announce Type: replace Abstract: Computational pathology holds substantial promise for improving diagnosis and guiding treatment decisions. Recent pathology foundation models enable

PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

Model ReleasesDGX agent

arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden

Robust Native Language Identification through Agentic Decomposition

Model ReleasesDGX agent

arXiv:2509.16666v2 Announce Type: replace Abstract: Large language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual

Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI

Model ReleasesDGX agent

arXiv:2608.06167v1 Announce Type: new Abstract: We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by aut

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

Model ReleasesDGX agent

arXiv:2608.05785v1 Announce Type: cross Abstract: Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fund

Unifying Structured and Unstructured Data Insights with BQ Search Innovations

Model ReleasesDGX agent

Modern enterprises possess a vast amount of unstructured data, yet they frequently encounter significant challenges in managing and extracting value from it. Historically, unlocking the insights hidde

6 Aug 2026

Advancing Utility Pole and Sign Detection Through Deep Learning

Model ReleasesDGX agent

arXiv:2608.04061v1 Announce Type: new Abstract: Utility poles are an essential part of the infrastructure used to support power distribution systems and other critical public services. Their regular i

An active-learning framework for real-time depth perception from monocular vision streams

Model ReleasesDGX agent

arXiv:2608.04917v1 Announce Type: new Abstract: Biological visual systems can perceive depth from monocular vision flow, continuously integrating temporal visual cues while maintaining a balance betwe

← Previous
1…328329330331332…1042
Next →