AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

86,457Total entries
1Added by human
86,456Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,039 results
30 Jun 2026

Can OCR-VLMs Read Devanagari? A Stress-Test Benchmark and Post-Correction Study

Model ReleasesDGX agent

arXiv:2606.29213v1 Announce Type: new Abstract: OCR systems, ranging from classical engines to specialised OCR vision-language models (OCR-VLMs) and frontier multimodal LLMs, report strong results on

Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models

ResearchDGX agent

arXiv:2606.29897v1 Announce Type: cross Abstract: Voice anonymization aims to protect speaker identity while preserving linguistic content and speech usability. However, most anonymization systems are

Demonstration-Free Robotic Control via LLM Agents

Model ReleasesDGX agent

arXiv:2601.20334v2 Announce Type: replace-cross Abstract: Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong performance but typically require task

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

DiffRGD: An Inference-Time Diffusion Guidance Through Riemannian Gradient Descent

ResearchDGX agent

arXiv:2606.28417v1 Announce Type: new Abstract: Recently, diffusion models have been widely adopted in generative modeling and have served as foundational models for many image generation tasks. To co

Diversity is the Strength of the AI Crowd

Model ReleasesDGX agent

arXiv:2606.29661v1 Announce Type: new Abstract: Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combine

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks

Model ReleasesDGX agent

arXiv:2510.14207v3 Announce Type: replace Abstract: Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. Prior jail

Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers

Model ReleasesDGX agent

arXiv:2510.25013v2 Announce Type: replace-cross Abstract: Mechanistic interpretability aims to reverse-engineer large language models (LLMs) into human-understandable computational circuits. However,

Enhancing Automatic Chord Recognition via Pseudo-Labeling and Knowledge Distillation

ResearchDGX agent

arXiv:2602.19778v4 Announce Type: replace-cross Abstract: Automatic Chord Recognition (ACR) is constrained by the scarcity of aligned chord labels, as well-aligned annotations are costly to acquire. A

Entropy-Regularized Reinforcement Learning for Linear-Quadratic Stackelberg Differential Games in Regime-Switching Diffusion Models

ResearchDGX agent

arXiv:2606.28671v1 Announce Type: new Abstract: Stackelberg differential games (SDGs) provide a powerful framework for hierarchical decision-making in stochastic and continuous-time environments, yet

Explaining Attention with Program Synthesis

Model ReleasesDGX agent

arXiv:2606.19317v2 Announce Type: replace-cross Abstract: A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descrip

HEARTS: Benchmarking LLM Reasoning on Health Time Series

Model ReleasesDGX agent

arXiv:2603.06638v3 Announce Type: replace-cross Abstract: The rise of large language models (LLMs) has shifted time series analysis from narrow analytics to general-purpose reasoning. Yet, existing be

Hierarchical Experimentalist Agents

Model ReleasesDGX agent

arXiv:2606.29315v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametr

How Far Do On-Prem Open LLMs Get on Text-to-SQL? A Cross-Family Size x Technique Frontier on BIRD

Model ReleasesDGX agent

arXiv:2606.29733v1 Announce Type: new Abstract: Organizations that cannot send data to a cloud API increasingly ask: how good is Text-to-SQL if the model must run on-premises on open weights, and whic

How Outpost VFX Uses AWS to Accelerate AI Model Training for Visual Effects

HardwareDGX agent

In this post, we explore how Outpost VFX achieved 8x faster training speeds using AWS infrastructure to transform their face replacement workflow, the technical architecture they implemented to overco

Large and Deep Factor Models

ResearchDGX agent

arXiv:2402.06635v3 Announce Type: replace-cross Abstract: We show that a deep neural network (DNN) trained to construct a stochastic discount factor (SDF) admits an additive decomposition separating n

Learning from Reliable Latent Prompts for Visual Recognition with Missing Modalities

Model ReleasesDGX agent

arXiv:2606.30597v1 Announce Type: new Abstract: Large-scale multimodal models (LMMs) have achieved superior performance in visual recognition by synergizing information across diverse, massive-scale p

MirrorCode: AI can rebuild entire programs from behavior alone

Model ReleasesDGX agent

arXiv:2606.30182v1 Announce Type: new Abstract: AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C compiler. Ho

MUSE: Unlocking Timestep as Native Task Steering for One-Step Dense Prediction

Model ReleasesDGX agent

arXiv:2606.30370v1 Announce Type: new Abstract: Monocular dense prediction has recently seen remarkable success by repurposing pre-trained diffusion models. This opens a promising yet challenging aven

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

Model ReleasesDGX agent

arXiv:2606.29814v1 Announce Type: new Abstract: We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared

Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating

Local AiDGX agent

arXiv:2603.12598v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) have shown remarkable potential across a wide array of vision-language tasks, leading to their adoption in crit

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

Model ReleasesDGX agent

arXiv:2512.09066v2 Announce Type: replace-cross Abstract: Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks rapi

Ornith-1.0-35B is now available in claude code through hf-claude

Model ReleasesDGX agent

Ornith-1.0-35B, a 35-billion parameter model, has been made available for use through Claude Code via Hugging Face integration. This announcement indicates expanded model availability and integration

Room 2016 for those attending @aiDotEngineer 2:25pm. Will also cover Galactica, early Llama reasoning efforts and more - think this is the f…

Model ReleasesDGX agent

This post announces a conference session at @aiDotEngineer scheduled for 2:25pm in Room 2016, covering topics including the Galactica model and early reasoning efforts in Llama models, with additional

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation

Model ReleasesDGX agent

arXiv:2606.30124v1 Announce Type: new Abstract: While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semant

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

Model ReleasesDGX agent

arXiv:2606.28589v1 Announce Type: new Abstract: Current approaches to enhance Large Language Model (LLM) reasoning, such as Chain-of-Thought and 'Wait' prompts, primarily encourage models to think mor

The Contagion Tensor: A Framework for Measuring Output-Distribution Coupling in Multi-Agent LLM Systems -- and Auditing the Claims It Enables

Model ReleasesDGX agent

arXiv:2606.28839v1 Announce Type: new Abstract: We introduce the Contagion Tensor, a measurement framework for quantifying how large language model (LLM) output distributions couple across modalities,

The Hidden Cost of Structured Generation in LLMs: Draft-Conditioned Constrained Decoding

Model ReleasesDGX agent

arXiv:2603.03305v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to generate executable outputs, JSON objects, and API calls, where a single syntax error ca

The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis

SafetyDGX agent

arXiv:2606.29581v1 Announce Type: cross Abstract: Modern LLM deployments routinely compress models and raise sampling temperature to reduce cost, latency, or repetition, yet safety evaluations usually

The NTNU System at the S&I Challenge 2025 SLA Open Track

Model ReleasesDGX agent

arXiv:2506.05121v3 Announce Type: replace Abstract: A recent line of research on spoken language assessment (SLA) employs neural models such as BERT and wav2vec 2.0 (W2V) to evaluate speaking proficie

VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake Detection

Model ReleasesDGX agent

arXiv:2603.21526v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) offer a promising path toward interpretable deepfake detection by generating textual explanations. However,

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

Model ReleasesDGX agent

arXiv:2606.28332v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains p

29 Jun 2026

A Unified Framework for Vision Transformers Equivariant to Discrete Subgroups of O(2)

Model ReleasesDGX agent

arXiv:2606.27864v1 Announce Type: new Abstract: Vision transformers have become a dominant architecture for visual recognition. However, standard models do not explicitly encode the planar symmetries

Bridging Ab Initio Symmetries and Global Nuclear Masses with Interpretable Neural Networks

Model ReleasesDGX agent

arXiv:2606.28287v1 Announce Type: cross Abstract: Ab initio modeling has established Wigner's SU(4) and Elliott's SU(3) as dominant symmetries of the nuclear force in light and intermediate-mass nucle

Physics-constrained neural networks for surrogate modeling of lossless periodic structures

ResearchDGX agent

arXiv:2606.28119v1 Announce Type: cross Abstract: We introduce a physics-constrained neural network (PCNN) for the rapid prediction of rigorous coupled-wave analysis (RCWA) outputs in the form of Jone

Qwen-Image-2.0-RL Technical Report

Model ReleasesDGX agent

arXiv:2606.27608v1 Announce Type: new Abstract: We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF) and on-policy distillation (OPD) t

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

SafetyDGX agent

arXiv:2606.28153v1 Announce Type: cross Abstract: Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not comprehensively

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

Model ReleasesDGX agent

arXiv:2606.27828v1 Announce Type: new Abstract: Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely r

26 Jun 2026

1/ On p (doom) tl;dr a) Everyone is making up the numbers b) nobody knows anything (least of all the experts), c) don't worry about it d) th…

Model ReleasesDGX agent

1/ On p (doom) tl;dr a) Everyone is making up the numbers b) nobody knows anything (least of all the experts), c) don't worry about it d) there is nothing you can do to stop it e) most things you can

Boundary-Aware Context Grounding for A Low-Channel EEG Agent

Model ReleasesDGX agent

arXiv:2606.26519v1 Announce Type: new Abstract: Large language models (LLMs) can make scientific software easier to use. However, a general model does not automatically know which measurements a parti

Escaping Iterative Parameter-Space Noise: Differentially Private Learning with a Hypernetwork

Model ReleasesDGX agent

arXiv:2606.26772v1 Announce Type: new Abstract: Differentially private (DP) training of neural networks is often hindered by the large amount of noise required by gradient-based methods such as DP-SGD

LearniBridge: Learnable Calibration of Feature Caching for Diffusion Models Acceleration

ResearchDGX agent

arXiv:2606.26778v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have driven substantial progress in image and video generation but suffer from prohibitive computational costs. Feature ca

MetaboNet-Bench: A Multi-modal Benchmark for Glucose Forecasting in Type 1 Diabetes

Model ReleasesDGX agent

arXiv:2606.18640v2 Announce Type: replace Abstract: Glucose forecasting algorithms are an important aspect of glycemic control management in type 1 diabetes. So far, the research community has develop

TMP: Tree-structured Mixed-policy Pruning for Large-scale Image Generation and Editing

Model ReleasesDGX agent

arXiv:2606.27089v1 Announce Type: new Abstract: Modern image generation model rapidly grows their sizes to meet high-fidelity image synthesis. However, they gradually become unaffordable for their eno

25 Jun 2026

Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge

SafetyDGX agent

arXiv:2602.02219v2 Announce Type: replace Abstract: Large language models are widely employed as evaluators, a paradigm commonly referred to as LLM-as-a-judge. Prior research has predominantly examine

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation

ResearchDGX agent

arXiv:2606.25432v1 Announce Type: cross Abstract: Inference efficiency is typically pursued by shrinking the model: distillation, pruning, quantization, and sparse routing each lower per-token cost wh

Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration

ResearchDGX agent

arXiv:2606.24970v1 Announce Type: new Abstract: Pruning Large Language Models (LLMs) reduces memory and inference costs by removing parts of the network, producing smaller models that retain most of t

Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs

Model ReleasesDGX agent

arXiv:2606.25758v1 Announce Type: new Abstract: While test-time adaptation (TTA) empowers vision-language models to adapt without costly retraining, it remains highly vulnerable to out-of-distribution

Falcon: Functional Assembly and Language for Compositional Reasoning in X-ray

Model ReleasesDGX agent

arXiv:2606.25701v1 Announce Type: new Abstract: Conventional vision-language models are largely object-centric, focusing on detecting and describing individual entities. In safety-critical X-ray bagga

Knowledge Cascade: Reverse Knowledge Distillation on Nonparametric Multivariate Functional Estimation

TutorialsDGX agent

arXiv:2606.25927v1 Announce Type: cross Abstract: As machine learning models and datasets continue to grow, developing complex models has become increasingly computationally demanding. Knowledge disti

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining

Model ReleasesDGX agent

arXiv:2606.26050v1 Announce Type: cross Abstract: Midway through an ordinary pretraining run, a small language model learns the pronoun-gender rule: cued with a girl's name ('Sue cried because'), it r

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity

Model ReleasesDGX agent

arXiv:2606.25634v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable progress in single-image perception, yet their ability to reason about complex cross-view

Velocity Prediction in Automatic Guitar Transcription

ResearchDGX agent

arXiv:2606.24912v1 Announce Type: cross Abstract: Automatic Music Transcription (AMT) models have achieved a high level of success in polyphonic transcription of various instruments. Velocity, typical

What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics

Model ReleasesDGX agent

arXiv:2606.25182v1 Announce Type: new Abstract: Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit policy-violating responses despite

WOLF-VLA: Whole-Body Humanoid Optimal Locomotion Framework for Vision-Language-Action Learning

Model ReleasesDGX agent

arXiv:2606.25591v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently demonstrated strong generalization in robotic manipulation, yet their applicability to whole-body, con

24 Jun 2026

Deep Learning Approaches for 3D Medical Scene Completion: From Geometric Modeling to Generative Paradigms

AgentsDGX agent

arXiv:2606.24180v1 Announce Type: cross Abstract: Three-dimensional scene completion has evolved as a major problem in computer vision and robotics, and its applications are diverse, including autonom

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

Model ReleasesDGX agent

arXiv:2606.24797v1 Announce Type: cross Abstract: Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Nevertheless, ex

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding

Model ReleasesDGX agent

arXiv:2606.24422v1 Announce Type: new Abstract: We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of mo

ESBMC-GraphPLC: Formal Verification of Graphical PLCopen XML Ladder Diagram Programs Using SMT-Based Model Checking

Local AiDGX agent

arXiv:2606.18941v3 Announce Type: replace-cross Abstract: PLCopen XML defines two encoding formats for IEC 61131-3 Ladder Diagram programs: a textual encoding using elements, and a graphical encoding

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO

Model ReleasesDGX agent

arXiv:2606.23032v2 Announce Type: replace Abstract: Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language model

Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice

Model ReleasesDGX agent

arXiv:2606.23716v1 Announce Type: cross Abstract: Legal AI benchmark research frequently invokes the assumption that large language models can improve access to justice, including for people who canno

← Previous
1…276277278279280…1034
Next →