AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,284 results
6 Aug 2026

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

Model ReleasesDGX agent

arXiv:2608.04657v1 Announce Type: new Abstract: World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mob

MoCA: Multi-modal Cross-masked Autoencoder for Digital Health Measurements

Model ReleasesDGX agent

arXiv:2506.02260v4 Announce Type: replace-cross Abstract: Wearable devices enable continuous multi-modal physiological and behavioral monitoring, yet analysis of these data streams faces fundamental c

Modality Agreement- and Conflict-Aware Prototype Hypergraph Learning for Multimodal Intent Understanding

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2608.04054v1 Announce Type: cross Abstract: Multimodal intent recognition requires understanding not only what textual, acoustic, and visual signals share, but also how they disagree. Such disag

Monte Carlo Tree Search for Table-to-Multimodal Report Generation

Model ReleasesDGX agent

arXiv:2608.04071v1 Announce Type: new Abstract: Automatically generating professional multimodal reports comprising both textual analysis and visual charts from structured tabular data is a critical c

MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding

Model ReleasesDGX agent

arXiv:2604.00513v3 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention. Alth

Neural Diversity Regularizes Hallucinations in Language Models

Model ReleasesDGX agent

arXiv:2510.20690v3 Announce Type: replace-cross Abstract: Language models continue to hallucinate despite increases in parameters, compute, and data. We propose neural diversity -- decorrelated parall

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

Model ReleasesDGX agent

arXiv:2608.04397v1 Announce Type: new Abstract: We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 p

Non-asymptotic implicit bias of logistic regression at early-stage gradient descent dynamics

Model ReleasesDGX agent

arXiv:2608.04382v1 Announce Type: new Abstract: Gradient descent has been of particular interest in modern machine learning beyond sole focus on optimization. Implicit bias emerging from optimization,

Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

Model ReleasesDGX agent

arXiv:2606.17188v3 Announce Type: replace-cross Abstract: Current multilingual evaluations for Vision-Language Models (VLMs) assume a one-to-one mapping between language and orthography, overlooking b

NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts

Model ReleasesDGX agent

arXiv:2608.04030v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) has transformed text-to-image synthesis, yet its ability to represent specialized engineering domains remains

nvidia/NVIDIA-Nemotron-Parse-2.0 · Hugging Face

Model ReleasesDGX agent

NVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information. Given a Red, Green, Blu

nvidias nemotron omni only loads its text half on a mac, so i wrote the vision and audio towers in mlx

Model ReleasesDGX agent

nvidias nemotron omni is open weights and it sees, hears and reasons. theres already a 4bit mlx quant on hugging face but only the text backbone loads with standard mlx tooling. the model card says it

🟩 NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp

Model ReleasesDGX agent

🐦‍⬛ Magpie-TTS Multilingual 🦜 Nemotron Speech Streaming EN 0.6B 🦜 Nemotron-3.5 ASR Streaming 🦜 Parakeet CTC 1.1B 🦜 Parakeet TDT 0.6B v3 🥦 NanoCodec Merged PR https://huggingface.co/nvidia/magpie_tts_m

OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

Model ReleasesDGX agent

arXiv:2608.05049v1 Announce Type: new Abstract: Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmar

OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing

Model ReleasesDGX agent

arXiv:2608.04434v1 Announce Type: new Abstract: Recent large language models (LLMs) have demonstrated remarkable progress in constraint-aware navigation, maze reasoning, and graph reasoning. However,

OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films

Model ReleasesDGX agent

arXiv:2608.04224v1 Announce Type: new Abstract: Historical films suffer from co-occurring visual and audio degradations---blur, noise, flicker, hiss, clipping, and dropout---yet existing methods resto

On the Effectiveness of Adaptation Strategies for VLM-Based Federated Learning in Remote Sensing

Model ReleasesDGX agent

arXiv:2608.04791v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative training of deep learning models across decentralized image archives without requiring data centralization

One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model f…

Model ReleasesDGX agent

One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements in a single model family Here's Meta AI's Spark (8th April), Spark 1.1 (9th Jul

One Surrogate to Fool Them All: Universal, Transferable, and Targeted Adversarial Attacks with CLIP

Model ReleasesDGX agent

arXiv:2505.19840v3 Announce Type: replace-cross Abstract: Deep Neural Networks (DNNs) have achieved widespread success yet remain prone to adversarial attacks. Typically, such attacks either involve f

Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planni…

Model ReleasesDGX agent

Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planning, tool calls, retries, and long contexts compound token us

OpenAI debuts Agent Plugins, an open standard for bundling skills and MCP servers, and says Amazon, Cursor, Microsoft, and Vercel are on its steering committee (Zac Hall/9to5Mac)

Model ReleasesDGX agent

Zac Hall / 9to5Mac: OpenAI debuts Agent Plugins, an open standard for bundling skills and MCP servers, and says Amazon, Cursor, Microsoft, and Vercel are on its steering committee — OpenAI's GPT-5 tur

OpenAI updates the default model for free users to GPT-5.6 Luna, adds unlimited text chats for free users, rolls out an improved GPT-5.6 Sol version, and more (Herb Scribner/Axios)

Model ReleasesDGX agent

Herb Scribner / Axios: OpenAI updates the default model for free users to GPT-5.6 Luna, adds unlimited text chats for free users, rolls out an improved GPT-5.6 Sol version, and more — OpenAI on Thursd

PADFormer: Pose-agnostic Anomaly Detection from Sparse View Images

Model ReleasesDGX agent

arXiv:2608.04210v1 Announce Type: new Abstract: Pose-agnostic Anomaly Detection (PAD) remains challenging as anomalies can appear under arbitrary viewpoints, requiring methods to handle significant po

Persistent Object Narratives for Token-Efficient Video Language Models

Model ReleasesDGX agent

arXiv:2608.04866v1 Announce Type: new Abstract: Video large language models (Video-LLMs) have made strong progress in open-ended video understanding. However, their visual interfaces remain token-inte

Personalized Federated Sparse Adaptation of Time-Series Foundation Models

Model ReleasesDGX agent

arXiv:2608.04695v1 Announce Type: cross Abstract: Federated adaptation of time-series foundation models (TSFMs) is attractive for building energy forecasting because meter data are private, distribute

Physics-informed reduced-order modelling with equivariant spectral submanifolds

Model ReleasesDGX agent

arXiv:2608.04239v1 Announce Type: new Abstract: Spectral submanifold (SSM) reduction has emerged as a mathematically principled route to reliable nonlinear reduced-order models, capturing dynamics bey

PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning

Model ReleasesDGX agent

arXiv:2608.04575v1 Announce Type: cross Abstract: Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language model

PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation

Model ReleasesDGX agent

arXiv:2608.01791v2 Announce Type: replace-cross Abstract: The rapid development of photonic integrated circuits (PICs) is shifting the design flow from traditional graphical user interface (GUI)-based

Plus and Pro users also now have a slider to choose how much reasoning effort ChatGPT puts into each response. We think it’s easier to use, …

Model ReleasesDGX agent

OpenAI announced that Plus and Pro subscribers now have a slider to adjust the amount of reasoning effort ChatGPT applies to each response. The update employs GPT‑5.6 Sol for both Instant and deep rea

Plus and Pro users can access the updated version of GPT‑5.6 Sol in ChatGPT along with the new slider starting today. This version of GPT-5.…

Model ReleasesDGX agent

Plus and Pro users can access the updated version of GPT‑5.6 Sol in ChatGPT along with the new slider starting today. This version of GPT-5.6 Sol is for everyday chats, so it will only be available in

Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes

Model ReleasesDGX agent

arXiv:2608.04426v1 Announce Type: cross Abstract: We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future

Privacy-Preserving Action Recognition: Taxonomy, Methods, and Privacy-Utility Trade-offs

Model ReleasesDGX agent

arXiv:2608.04501v1 Announce Type: new Abstract: Video surveillance in public safety, healthcare, and smart environments has made continuous human monitoring routine, raising real risks to personal ide

Protoreasoning in Tiny Transformers

Model ReleasesDGX agent

arXiv:2608.04980v1 Announce Type: cross Abstract: We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study step-by-ste

Radar4D-VLM: Proposal-Grounded Temporal 4D Radar Reasoning Across Frozen Language Models

Model ReleasesDGX agent

arXiv:2608.04130v1 Announce Type: new Abstract: Vision-language models for autonomous driving primarily rely on cameras and LiDAR, leaving 4D radar largely unexplored as a standalone perceptual modali

Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos

Model ReleasesDGX agent

arXiv:2608.04939v1 Announce Type: new Abstract: Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or

ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination

Model ReleasesDGX agent

arXiv:2608.04385v1 Announce Type: new Abstract: Vision-Language Models (VLMs) often lose visual grounding during multi-step reasoning: as reasoning chains grow longer, later inference steps rely incre

Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression

Model ReleasesDGX agent

arXiv:2608.04569v1 Announce Type: new Abstract: Hard prompt compression reduces long-context inference cost by independently scoring tokens, sentences, or chunks and retaining the highest-scoring unit

RepairFormer: Automated Repair of Structured Inputs Using Transformers

Model ReleasesDGX agent

arXiv:2608.05060v1 Announce Type: cross Abstract: Structured input files such as JSON, DOT, OBJ, INI, S-expression, and TinyC are widely used in software systems, but small corruptions can cause parse

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists

Model ReleasesDGX agent

arXiv:2608.04783v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repository-scale ass

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

Model ReleasesDGX agent

arXiv:2608.04514v1 Announce Type: new Abstract: Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-co

ResPlan: A Large-Scale Vector-Graph Dataset of 17,000 Residential Floor Plans

Model ReleasesDGX agent

arXiv:2508.14006v2 Announce Type: replace Abstract: We introduce ResPlan, a dataset of 17,000 residential floor plans with vector geometry, room-connectivity graphs, and metric-scale coordinates. Each

Retrieve in Time, Correct in Frequency

Model ReleasesDGX agent

arXiv:2608.04527v1 Announce Type: new Abstract: Frozen vision-language-action (VLA) policies generate temporally extended action chunks, but long-horizon manipulation remains vulnerable to accumulated

RIP to every 'which framework should I use' thread. harness on top, framework in the middle, runtime at the floor. one stack, three heights,…

Model ReleasesDGX agent

RIP to every 'which framework should I use' thread. harness on top, framework in the middle, runtime at the floor. one stack, three heights, fully composable. paste these four images into claude and a

Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition

Model ReleasesDGX agent

arXiv:2608.05115v1 Announce Type: cross Abstract: Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom incident rec

Robust and Personalized Federated Learning for Aircraft-Engine Prognostics under Benign and Adversarial Client Heterogeneity

Model ReleasesDGX agent

arXiv:2608.04045v1 Announce Type: cross Abstract: Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry without shar

Robust Control under Stationary Ambiguity

Model ReleasesDGX agent

arXiv:2608.04832v1 Announce Type: new Abstract: Control policies optimized in simulation can perform poorly in the real system when the parameters x of the simulator are estimated from limited data bu

Robustness Emerges Early in Training Dynamics, but Is Not Preserved

Model ReleasesDGX agent

arXiv:2608.04442v1 Announce Type: cross Abstract: Robustness to natural corruptions remains a fundamental challenge for deep neural networks. In this paper, we identify a robustness fading phenomenon

RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis

Model ReleasesDGX agent

arXiv:2602.11506v4 Announce Type: replace-cross Abstract: The transition toward localized intelligence through Small Language Models (SLMs) has intensified the need for rigorous performance characteri

Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?

Model ReleasesDGX agent

arXiv:2608.05097v1 Announce Type: new Abstract: Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same

SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models

Model ReleasesDGX agent

arXiv:2608.04975v1 Announce Type: cross Abstract: SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theo

Scotoma-2: Gemma4, but with less annoying slop and better writing.

Model ReleasesDGX agent

GGUFs here: https://huggingface.co/ReadyArt/gemma-4-31B-it-scotoma-2-GGUF Disclaimer: By slop, we are specifically talking about specific tics with the model(sentence structures), but this doesn't inc

Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First

Model ReleasesDGX agent

arXiv:2608.04804v1 Announce Type: cross Abstract: Frontier language models can resolve repository-level software issues, but each attempt is expensive, and existing routers select a model from the iss

SEAR: Simple and Efficient Adaptation of Visual Geometric Transformers for Unpaired RGB+Thermal 3D Reconstruction

Model ReleasesDGX agent

arXiv:2603.18774v2 Announce Type: replace Abstract: Foundational feed-forward visual geometry models enable accurate and efficient camera pose estimation and scene reconstruction by learning strong sc

Semantic Frame Interpolation

Model ReleasesDGX agent

arXiv:2507.05173v2 Announce Type: replace Abstract: Generating intermediate video content of varying lengths based on given first and last frames, along with text prompt information, offers significan

Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load

Model ReleasesDGX agent

arXiv:2608.05018v1 Announce Type: new Abstract: Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law designate as

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2608.04244v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks

SimMOF: AI agent for Automated MOF Simulations

Model ReleasesDGX agent

arXiv:2603.29152v2 Announce Type: replace Abstract: Metal-organic frameworks (MOFs) offer a vast design space, and as such, computational simulations play a critical role in predicting their structura

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Model ReleasesDGX agent

arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

Model ReleasesDGX agent

arXiv:2608.05137v1 Announce Type: new Abstract: Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, incl

Sources: Alibaba plans to ask heavy commercial users of its next Qwen open model for a share of revenue; Moonshot's Kimi K3 requires up to a 30% revenue share (Reuters)

Model ReleasesDGX agent

Reuters: Sources: Alibaba plans to ask heavy commercial users of its next Qwen open model for a share of revenue; Moonshot's Kimi K3 requires up to a 30% revenue share — Chinese technology giant Aliba

← Previous
1…2526272829…372
Next →