AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
59,851 results
Research

Understanding Counting Mechanisms in Large Language and Vision-Language Models

DGX agent

arXiv:2511.17699v2 Announce Type: replace Abstract: Counting is one of the fundamental abilities of large language models (LLMs) and large vision-language models (LVLMs). This paper examines how these

researcharxiv-cs-cv
21 Apr 2026
Research

Understanding the Prompt Sensitivity

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.18389v1 Announce Type: new Abstract: Prompt sensitivity, which refers to how strongly the output of a large language model (LLM) depends on the exact wording of its input prompt, raises con

researcharxiv-cs-cl
21 Apr 2026
Model Releases

Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis

DGX agent

arXiv:2604.16538v1 Announce Type: cross Abstract: Automatic translation of natural language mathematics into faithful Lean 4 code is hindered by the fundamental dissonance between informal set-theoret

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark

DGX agent

arXiv:2510.13759v3 Announce Type: replace Abstract: Unified multimodal models aim to jointly enable visual understanding and generation, yet current benchmarks rarely examine their true integration. E

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation

DGX agent

arXiv:2602.09130v3 Announce Type: replace Abstract: Model compression is increasingly essential for deploying large language models (LLMs), yet existing comparative studies largely focus on pruning an

safetyarxiv-cs-lg
21 Apr 2026
Safety

UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels

DGX agent

arXiv:2604.16678v1 Announce Type: new Abstract: Contrastive objectives power state-of-the-art multimodal models, but their training remains slow, relying on long stochastic optimization. We propose a

safetyarxiv-cs-lg
21 Apr 2026
Safety

UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement

DGX agent

arXiv:2604.17850v1 Announce Type: new Abstract: Style transfer must match a target style while preserving content semantics. DiT-based diffusion models often suffer from content-style entanglement, le

safetyarxiv-cs-cv
21 Apr 2026
Applications

UniDomain: Pretraining a Unified PDDL Domain from Real-World Demonstrations for Generalizable Robot Task Planning

DGX agent

arXiv:2507.21545v3 Announce Type: replace Abstract: Robotic task planning in real-world environments requires reasoning over implicit constraints from language and vision. While LLMs and VLMs offer st

applicationsarxiv-cs-ro
21 Apr 2026
Model Releases

Unified Multimodal Brain Decoding via Cross-Subject Soft-ROI Fusion

DGX agent

arXiv:2512.20249v3 Announce Type: replace-cross Abstract: Multimodal brain decoding aims to reconstruct semantic information that is consistent with visual stimuli from brain activity signals such as

model-releasesarxiv-cs-cv
21 Apr 2026
Agents

Unified Ultrasound Intelligence Toward an End-to-End Agentic System

DGX agent

arXiv:2604.16914v1 Announce Type: new Abstract: Clinical ultrasound analysis demands models that generalize across heterogeneous organs, views, and devices, while supporting interpretable workflow-lev

agentsarxiv-cs-cv
21 Apr 2026
Research

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

DGX agent

arXiv:2604.17565v1 Announce Type: new Abstract: Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geomet

researcharxiv-cs-cv
21 Apr 2026
Model Releases

UniMamba: A Unified Spatial-Temporal Modeling Framework with State-Space and Attention Integration

DGX agent

arXiv:2604.16325v1 Announce Type: new Abstract: Multivariate time series forecasting is fundamental to numerous domains such as energy, finance, and environmental monitoring, where complex temporal de

model-releasesarxiv-cs-lg
21 Apr 2026
Research

UniMesh: Unifying 3D Mesh Understanding and Generation

DGX agent

arXiv:2604.17472v1 Announce Type: new Abstract: Recent advances in 3D vision have led to specialized models for either 3D understanding (e.g., shape classification, segmentation, reconstruction) or 3D

researcharxiv-cs-cv
21 Apr 2026
Tutorials

UniSim: A Unified Simulator for Time-Coarsened Dynamics of Biomolecules

DGX agent

arXiv:2506.03157v4 Announce Type: replace-cross Abstract: Molecular Dynamics (MD) simulations are essential for understanding the atomic-level behavior of molecular systems, giving insights into their

tutorialsarxiv-cs-lg
21 Apr 2026
Research

Universal Diffusion-Based Probabilistic Downscaling

DGX agent

arXiv:2602.11893v3 Announce Type: replace Abstract: We introduce a universal diffusion-based downscaling framework that lifts deterministic low-resolution weather forecasts into probabilistic high-res

researcharxiv-cs-lg
21 Apr 2026
Model Releases

Universally Empowering Zeroth-Order Optimization via Adaptive Layer-wise Sampling

DGX agent

arXiv:2604.18264v1 Announce Type: new Abstract: Zeroth-Order optimization presents a promising memory-efficient paradigm for fine-tuning Large Language Models by relying solely on forward passes. Howe

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning

DGX agent

arXiv:2603.23404v2 Announce Type: replace-cross Abstract: Existing Multimodal Large Language Models (MLLMs) struggle with 3D spatial reasoning, as they fail to construct structured abstractions of the

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models

DGX agent

arXiv:2604.18000v1 Announce Type: new Abstract: Recent Vision-Language-Action (VLA) models report impressive success rates on standard robotic benchmarks, fueling optimism about general-purpose physic

model-releasesarxiv-cs-ro
21 Apr 2026
Tutorials

Unraveling the Key of Machine Learning-based Android Malware Detection

DGX agent

arXiv:2402.02953v2 Announce Type: replace-cross Abstract: With the rapid advancement of machine learning (ML), ML-based Android malware detection has gained significant popularity due to its ability t

tutorialsarxiv-cs-lg
21 Apr 2026
Model Releases

Unsupervised Discovery of Intermediate Phase Order in the Frustrated J_1-J_2 Heisenberg Model via Prometheus Framework

DGX agent

arXiv:2602.21468v4 Announce Type: replace-cross Abstract: The spin-1/2 J_1-J_2 Heisenberg model on the square lattice exhibits a debated intermediate phase between Neel antiferromagnetic and stripe or

model-releasesarxiv-cs-lg
21 Apr 2026
Safety

Untrained CNNs Match Backpropagation at V1: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI

DGX agent

arXiv:2604.16875v1 Announce Type: new Abstract: A central question in computational neuroscience is whether the learning rule used to train a neural network determines how well its internal representa

safetyarxiv-cs-lg
21 Apr 2026
Model Releases

Unveiling Deepfakes: A Frequency-Aware Triple Branch Network for Deepfake Detection

DGX agent

arXiv:2604.17477v1 Announce Type: new Abstract: Advanced deepfake technologies are blurring the lines between real and fake, presenting both revolutionary opportunities and alarming threats. While it

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Upper Approximation Bounds for Neural Oscillators

DGX agent

arXiv:2512.01015v2 Announce Type: replace Abstract: Neural oscillators, originating from second-order ordinary differential equations (ODEs), have demonstrated strong performance in stably learning ca

researcharxiv-cs-lg
21 Apr 2026
Model Releases

User-Assistant Bias in LLMs

DGX agent

arXiv:2508.15815v3 Announce Type: replace Abstract: Modern large language models (LLMs) are typically trained and deployed using structured role tags (e.g. system, user, assistant, tool) that explicit

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Using large language models for embodied planning introduces systematic safety risks

DGX agent

arXiv:2604.18463v1 Announce Type: cross Abstract: Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe plann

model-releasesarxiv-cs-lg
21 Apr 2026
Research

Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models

DGX agent

arXiv:2506.00065v2 Announce Type: replace Abstract: Multimodal language models (MLMs) increasingly demonstrate human-like communication, yet their use of everyday perspectival words remains poorly und

researcharxiv-cs-cl
21 Apr 2026
Model Releases

VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning

DGX agent

arXiv:2402.13243v2 Announce Type: replace Abstract: Learning a human-like driving policy from large-scale driving demonstrations is promising, but the uncertainty and non-deterministic nature of plann

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Variational Autoencoder Domain Adaptation for Cross-System Generalization in ML-Based SOP Monitoring

DGX agent

arXiv:2604.18035v1 Announce Type: new Abstract: Machine learning (ML) models trained to detect physical-layer threats on one optical fiber system often fail catastrophically when applied to a differen

researcharxiv-cs-lg
21 Apr 2026
Research

VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis

DGX agent

arXiv:2509.16538v3 Announce Type: replace-cross Abstract: We propose VC-Inspector, a lightweight, open-source large multimodal model (LMM) for reference-free evaluation of video captions, with a focus

researcharxiv-cs-cl
21 Apr 2026
Model Releases

VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision

DGX agent

arXiv:2510.27462v2 Announce Type: replace Abstract: Supervised fine-tuning (SFT) on long chain-of-thought (CoT) trajectories has emerged as a crucial technique for enhancing the reasoning abilities of

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

DGX agent

arXiv:2604.17248v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing sp

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Video Panels for Long Video Understanding

DGX agent

arXiv:2509.23724v2 Announce Type: replace Abstract: Recent Video-Language Models (VLMs) achieve promising results on long-video understanding, but their performance still lags behind that achieved on

model-releasesarxiv-cs-cv
21 Apr 2026
Local Ai

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

DGX agent

arXiv:2604.17656v1 Announce Type: cross Abstract: Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typic

local-aiarxiv-cs-cl
21 Apr 2026
Safety

VIDEOP2R: Video Understanding from Perception to Reasoning

DGX agent

arXiv:2511.11113v2 Announce Type: replace Abstract: Reinforcement fine-tuning (RFT), a two-stage framework consisting of supervised fine-tuning (SFT) and reinforcement learning (RL) has shown promisin

safetyarxiv-cs-cv
21 Apr 2026
Local Ai

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning

DGX agent

arXiv:2601.15724v2 Announce Type: replace Abstract: Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning

local-aiarxiv-cs-cv
21 Apr 2026
Model Releases

VIDS: A Verified Imaging Dataset Standard for Medical AI

DGX agent

arXiv:2604.17525v1 Announce Type: cross Abstract: Medical imaging AI development is fundamentally dependent on annotated datasets, yet no existing standard provides machine-enforceable validation acro

model-releasesarxiv-cs-cv
21 Apr 2026
Research

View-Consistent 3D Scene Editing via Dual-Path Structural Correspondense and Semantic Continuity

DGX agent

arXiv:2604.17801v1 Announce Type: new Abstract: Text-driven 3D scene editing has recently attracted increasing attention. Most existing methods follow a render-edit-optimize pipeline, where multi-view

researcharxiv-cs-cv
21 Apr 2026
Research

ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes

DGX agent

arXiv:2604.17623v1 Announce Type: new Abstract: Kinematic rigs provide a structured interface for articulating 3D meshes, but they lack an inherent representation of the plausible manifold of joint co

researcharxiv-cs-cv
21 Apr 2026
Applications

Vision-Braille: A Curriculum Learning Toolkit and Braille-Chinese Corpus for Braille Translation

DGX agent

arXiv:2407.06048v2 Announce Type: replace Abstract: We present Vision-Braille, the first publicly available end-to-end system for translating Chinese Braille extracted from images into written Chinese

applicationsarxiv-cs-cl
21 Apr 2026
Research

Vision Language Models are Biased

DGX agent

arXiv:2505.23941v4 Announce Type: replace-cross Abstract: Large language models (LLMs) memorize a vast amount of prior knowledge from the Internet that helps them on downstream tasks but also may noto

researcharxiv-cs-cv
21 Apr 2026
Applications

Visual-RRT: Finding Paths toward Visual-Goals via Differentiable Rendering

DGX agent

arXiv:2604.16388v1 Announce Type: cross Abstract: Rapidly-exploring random trees (RRTs) have been widely adopted for robot motion planning due to their robustness and theoretical guarantees. However,

applicationsarxiv-cs-cv
21 Apr 2026
Research

ViT^3: Unlocking Test-Time Training in Vision

DGX agent

arXiv:2512.01643v2 Announce Type: replace Abstract: Test-Time Training (TTT) has recently emerged as a promising direction for efficient sequence modeling. TTT reformulates attention operation as an o

researcharxiv-cs-cv
21 Apr 2026
Model Releases

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

DGX agent

arXiv:2505.20279v4 Announce Type: replace-cross Abstract: The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes,

model-releasesarxiv-cs-cl
21 Apr 2026
Research

Vocab Diet: Reshaping the Vocabulary of LLMs via Vector Arithmetic

DGX agent

arXiv:2510.17001v2 Announce Type: replace Abstract: Large language models (LLMs) often encode word-form variation (e.g., walk vs. walked) as linear directions in the embedding space. However, standard

researcharxiv-cs-cl
21 Apr 2026
Local Ai

VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models

DGX agent

arXiv:2508.15229v3 Announce Type: replace Abstract: Small Language Models (SLMs) provide computational advantages in resource-constrained environments, yet memory limitations remain a critical bottlen

local-aiarxiv-cs-cl
21 Apr 2026
Model Releases

Voronoi-guided Bilateral 2D Gaussian Splatting for Arbitrary-Scale Hyperspectral Image Super-Resolution

DGX agent

arXiv:2604.17727v1 Announce Type: new Abstract: Most existing hyperspectral image super-resolution methods require modifications for different scales, limiting their flexibility in arbitrary-scale rec

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception

DGX agent

arXiv:2604.17475v1 Announce Type: cross Abstract: Small Vision-Language Models (SVLMs) are efficient task controllers but often suffer from visual brittleness and poor tool orchestration. They typical

safetyarxiv-cs-cl
21 Apr 2026
Research

Wasserstein Distributionally Robust Risk-Sensitive Estimation via Conditional Value-at-Risk

DGX agent

arXiv:2604.18546v1 Announce Type: new Abstract: We propose a distributionally robust approach to risk-sensitive estimation of an unknown signal x from an observed signal y. The unknown signal and obse

researcharxiv-cs-lg
21 Apr 2026
← Previous
1…11281129113011311132…1247
Next →