AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
Human
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
59,295 results
11 Aug 2026

Understanding Reasoning from Pretraining to Post-Training

SafetyDGX agent

arXiv:2607.16097v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is l

UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation

TutorialsDGX agent

arXiv:2608.09287v1 Announce Type: cross Abstract: Data-Free Knowledge Distillation (DFKD) transfers knowledge from a pretrained teacher model to a compact student model by synthesizing semantically in

Unified Hallucination Fuzzing for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2608.07525v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applicat

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Unimodality-Promoting Regularized Learning for Ordinal Regression

SafetyDGX agent

arXiv:2608.08359v1 Announce Type: new Abstract: Ordinal regression, also called ordinal classification, is classification of ordinal data, in which the underlying target variable is categorical and co

UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation

Local AiDGX agent

arXiv:2608.09143v1 Announce Type: new Abstract: Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodifi

UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

Model ReleasesDGX agent

arXiv:2608.08627v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes

UniScale: Arbitrary-Scale Industrial Anomaly Generation

TutorialsDGX agent

arXiv:2608.07864v1 Announce Type: new Abstract: Industrial anomaly inspection faces a major challenge due to the lack of real-world anomaly samples. While generative models are used to create anomaly

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

ResearchDGX agent

arXiv:2608.08676v1 Announce Type: cross Abstract: Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, t

Universal or Language-Family-Specific Script Unification for Cross-Lingual Transfer? A Case Study on Turkic Languages

ApplicationsDGX agent

arXiv:2608.09356v1 Announce Type: new Abstract: Closely related languages written in different scripts expose little surface overlap to multilingual models, limiting cross-lingual transfer. We compare

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

Model ReleasesDGX agent

arXiv:2608.09209v1 Announce Type: new Abstract: Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic

UnsDrive: Towards Robust End-to-End Autonomous Driving in Unstructured Scenes

SafetyDGX agent

arXiv:2608.09098v1 Announce Type: new Abstract: End-to-end planning has shown strong promise for autonomous driving, but most existing methods are designed for structured urban roads and generalize po

UNSPECIFIC: General Constraint Synthesis for Breaking Copy-and-Paste Shortcut in LLM Instruction Following

Model ReleasesDGX agent

arXiv:2608.09154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly expected to follow long lists of constraints in complex instructions, and synthesizing instructions from a

Unsupervised Domain Adaptation for Multitask Image Analysis in Realistic Context with Extreme Label Shift; Application to the CTAO first Large Sized Telescope

ResearchDGX agent

arXiv:2608.09630v1 Announce Type: cross Abstract: Unsupervised domain adaptation is a widespread set of methods that leverages the knowledge of a labeled source domain to train a model to perform well

Unsupervised Point Cloud Registration with Self-Distillation

Model ReleasesDGX agent

arXiv:2409.07558v2 Announce Type: replace Abstract: Rigid point cloud registration is a fundamental problem and highly relevant in robotics and autonomous driving. Nowadays deep learning methods can b

Unsure but Certain: Uncovering the Representation-Confidence Gap in Diffusion Language Models

ResearchDGX agent

arXiv:2608.08791v1 Announce Type: new Abstract: Diffusion language models use broad context to create text, suggesting they might handle input noise better than standard models. Testing reveals this i

Unveiling the Secret of AdaLN-Zero in Diffusion Transformer

ResearchDGX agent

arXiv:2608.09438v1 Announce Type: new Abstract: Diffusion transformer (DiT), a rapidly emerging architecture for image generation, has gained much attention. However, despite ongoing efforts to improv

UPolarSQ: Polar Representation Learning for Optic Disc and Peripapillary Atrophy Segmentation and Quantification in Fundus Photographs

ResearchDGX agent

arXiv:2608.08771v1 Announce Type: new Abstract: Myopia-induced posterior-pole remodeling is frequently accompanied by Optic Disc (OD) deformation and Peripapillary Atrophy (PPA), both of which provide

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control

ApplicationsDGX agent

arXiv:2608.07870v1 Announce Type: new Abstract: Improving sample efficiency remains a core challenge in reinforcement learning (RL), especially in real-world settings like robotics, where data collect

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models

SafetyDGX agent

arXiv:2608.08622v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses u

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

SafetyDGX agent

arXiv:2608.09448v1 Announce Type: cross Abstract: Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains d

Variance reduction in lattice QCD observables via normalizing flows

ResearchDGX agent

arXiv:2603.02984v2 Announce Type: replace-cross Abstract: Normalizing flows can be used to construct unbiased, reduced-variance estimators for lattice field theory observables that are defined by a de

VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging

Model ReleasesDGX agent

arXiv:2511.18121v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visu

VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge

AgentsDGX agent

arXiv:2608.07994v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex product documen

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

Model ReleasesDGX agent

arXiv:2608.08477v1 Announce Type: new Abstract: We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to

VeinCast: Physics-Guided Dynamic Field Graphs with Graph-Conditioned Fusion for Global Medium-Range Weather Forecasting

Model ReleasesDGX agent

arXiv:2608.09286v1 Announce Type: cross Abstract: Global medium-range weather forecasting requires modeling structured yet state-dependent interactions among heterogeneous atmospheric fields. Existing

verdi: retrieval is not transfer for continual world model optimization

HardwareDGX agent

arXiv:2608.09537v1 Announce Type: new Abstract: Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence. However, optimizing a pretrained world model t

Verication-driven closed-loop multi-agent large language modelframework for code-compliant structural design

Model ReleasesDGX agent

arXiv:2608.07978v1 Announce Type: cross Abstract: Multi-agent large language model(LLM)systems are applied to structural design,yet most use one-shot generation and cannot verify their output,leaving

Verifiably grounded machine interpretation of lunar geology

Local AiDGX agent

arXiv:2608.09276v1 Announce Type: new Abstract: Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations. Here, we present a step toward an a

VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative Scaffolding

ResearchDGX agent

arXiv:2608.09698v1 Announce Type: cross Abstract: Great fiction earns its verisimilitude through precise details, from how a longsword is gripped to pierce armor gaps to why a bleeding corpse cannot y

Vid2WAM: Distilling Video Diffusion Priors into World Action Models

SafetyDGX agent

arXiv:2608.08558v1 Announce Type: new Abstract: World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generali

VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

Model ReleasesDGX agent

arXiv:2608.09573v1 Announce Type: new Abstract: Natural-language-driven 'vibe coding' enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of thei

View-Adaptive Renderer for View-Consistent 2D-to-3D Generation

ApplicationsDGX agent

arXiv:2608.09110v1 Announce Type: new Abstract: Reconstructing 3D shapes from a single image remains a fundamental yet challenging problem in computer vision. Traditional monocular 3D generation pipel

VIGIL: Tackling Hallucination Detection in Image Recontextualization

Model ReleasesDGX agent

arXiv:2602.14633v2 Announce Type: replace Abstract: We introduce VIGIL (Visual Inconsistency & Generative In-context Lucidity), a benchmark dataset and framework that provides a fine-grained categoriz

Vision-Language Grounding as Bidirectional Concept Correspondence

Local AiDGX agent

arXiv:2608.07886v1 Announce Type: cross Abstract: Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization proble

Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties

ResearchDGX agent

arXiv:2608.07726v1 Announce Type: new Abstract: Estimating volumetric mechanical properties, including Young's modulus, Poisson's ratio, and density at each voxel, is intrinsically ambiguous from visi

VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs

ResearchDGX agent

arXiv:2510.16598v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) encounter significant computational and memory bottlenecks from the massive number of visual tokens generat

Visual Distortion Detection in UGC Images Using Large Multimodal Models

ApplicationsDGX agent

arXiv:2608.09122v1 Announce Type: cross Abstract: The localized depiction of perceptual quality has long been a crucial, yet underexplored, challenge in image quality assessment (IQA). Existing approa

Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding

Local AiDGX agent

arXiv:2608.08832v1 Announce Type: new Abstract: Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nod

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling

Model ReleasesDGX agent

arXiv:2608.08630v1 Announce Type: new Abstract: Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-atte

VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression

ResearchDGX agent

arXiv:2608.08366v1 Announce Type: new Abstract: Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few th

VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference

SafetyDGX agent

arXiv:2608.08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks. Despite this progres

VTO: Visual Tool Orchestration for Video Anomaly Detection

Model ReleasesDGX agent

arXiv:2608.08219v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learn

WA-SpecDec: World-Aware Speculative Decoding for Vision-Language-Action Models

ResearchDGX agent

arXiv:2608.08725v1 Announce Type: new Abstract: Vision-language-action (VLA) policies generate robot controls autoregressively, making closed-loop latency dominated by repeated target-model forward pa

Walk-on-Spheres Monte Carlo and deep neural network approximations of elliptic PDEs with drift and killing

ResearchDGX agent

arXiv:2608.09494v1 Announce Type: cross Abstract: In this paper we provide Monte Carlo and deep neural network approximations for stochastic representations of solutions to linear elliptic partial dif

Walking through Discussions: A Mobile Visual Analytics System for In-Situ Group Discussion Analysis

ResearchDGX agent

arXiv:2608.08617v1 Announce Type: new Abstract: Group discussion-based teaching is widely used to foster collaborative learning, yet teachers in physical classrooms often struggle to simultaneously mo

Warp-free Cross-view Geo-localization via Feature-space Consensus Mining

Local AiDGX agent

arXiv:2608.09321v1 Announce Type: new Abstract: Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imager

WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training

SafetyDGX agent

arXiv:2608.09447v1 Announce Type: cross Abstract: On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state mismatch o

Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning Systems

Model ReleasesDGX agent

arXiv:2401.04013v2 Announce Type: replace Abstract: Deep learning models, such as wide neural networks, can be conceptualized as nonlinear dynamical physical systems characterized by a multitude of in

Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning

Local AiDGX agent

arXiv:2608.07593v1 Announce Type: cross Abstract: Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Exi

WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks

Model ReleasesDGX agent

arXiv:2506.01952v2 Announce Type: replace-cross Abstract: Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent

What Irregularity Costs: CUDA C++, Rust, and Triton on a Hash-Blocked GPU Workload

HardwareDGX agent

arXiv:2608.08287v1 Announce Type: new Abstract: GPU language comparisons are almost always run on tiled dense linear algebra, where every toolchain is good and the differences are small. We implement

What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files

SafetyDGX agent

arXiv:2608.08453v1 Announce Type: new Abstract: Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting resources, enabling Large Language Model (LLM) age

What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems

Model ReleasesDGX agent

arXiv:2608.07565v1 Announce Type: cross Abstract: Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactio

What Would Fix This RAG Failure? Auditing Counterfactual Response with Paired Evidence Interventions

Model ReleasesDGX agent

arXiv:2608.08944v1 Announce Type: cross Abstract: A failed retrieval-augmented generation (RAG) answer can be consistent with several unseen responses to evidence repair. We introduce Pair-ID, an offl

When Can Fraud Operations Authorize Automation? A Decision-Support Framework for Fresh Audit Evidence and Review Workload

ResearchDGX agent

arXiv:2608.08577v1 Announce Type: new Abstract: Fraud operations must allocate events among automatic approval, analyst review, and automatic blocking even though the labels needed to evaluate these a

When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information

ResearchDGX agent

arXiv:2608.09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability u

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation

Model ReleasesDGX agent

arXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model 'value dispositions,' and a concentration/ex

When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition

Model ReleasesDGX agent

arXiv:2608.09490v1 Announce Type: new Abstract: Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predi

When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery

Model ReleasesDGX agent

arXiv:2608.08132v1 Announce Type: new Abstract: Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical informa

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes

SafetyDGX agent

arXiv:2608.07911v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard. This makes expert cache

← Previous
1…3839404142…989
Next →