AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
Human
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
59,851 results
Research

Understanding Counting Mechanisms in Large Language and Vision-Language Models

DGX agent

arXiv:2511.17699v2 Announce Type: replace Abstract: Counting is one of the fundamental abilities of large language models (LLMs) and large vision-language models (LVLMs). This paper examines how these

researcharxiv-cs-cv
21 Apr 2026
Research
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Understanding the Prompt Sensitivity

DGX agent

arXiv:2604.18389v1 Announce Type: new Abstract: Prompt sensitivity, which refers to how strongly the output of a large language model (LLM) depends on the exact wording of its input prompt, raises con

researcharxiv-cs-cl
21 Apr 2026
Model Releases

Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis

DGX agent

arXiv:2604.16538v1 Announce Type: cross Abstract: Automatic translation of natural language mathematics into faithful Lean 4 code is hindered by the fundamental dissonance between informal set-theoret

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark

DGX agent

arXiv:2510.13759v3 Announce Type: replace Abstract: Unified multimodal models aim to jointly enable visual understanding and generation, yet current benchmarks rarely examine their true integration. E

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation

DGX agent

arXiv:2602.09130v3 Announce Type: replace Abstract: Model compression is increasingly essential for deploying large language models (LLMs), yet existing comparative studies largely focus on pruning an

safetyarxiv-cs-lg
21 Apr 2026
Safety

UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels

DGX agent

arXiv:2604.16678v1 Announce Type: new Abstract: Contrastive objectives power state-of-the-art multimodal models, but their training remains slow, relying on long stochastic optimization. We propose a

safetyarxiv-cs-lg
21 Apr 2026
Safety

UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement

DGX agent

arXiv:2604.17850v1 Announce Type: new Abstract: Style transfer must match a target style while preserving content semantics. DiT-based diffusion models often suffer from content-style entanglement, le

safetyarxiv-cs-cv
21 Apr 2026
Applications

UniDomain: Pretraining a Unified PDDL Domain from Real-World Demonstrations for Generalizable Robot Task Planning

DGX agent

arXiv:2507.21545v3 Announce Type: replace Abstract: Robotic task planning in real-world environments requires reasoning over implicit constraints from language and vision. While LLMs and VLMs offer st

applicationsarxiv-cs-ro
21 Apr 2026
Model Releases

Unified Multimodal Brain Decoding via Cross-Subject Soft-ROI Fusion

DGX agent

arXiv:2512.20249v3 Announce Type: replace-cross Abstract: Multimodal brain decoding aims to reconstruct semantic information that is consistent with visual stimuli from brain activity signals such as

model-releasesarxiv-cs-cv
21 Apr 2026
Agents

Unified Ultrasound Intelligence Toward an End-to-End Agentic System

DGX agent

arXiv:2604.16914v1 Announce Type: new Abstract: Clinical ultrasound analysis demands models that generalize across heterogeneous organs, views, and devices, while supporting interpretable workflow-lev

agentsarxiv-cs-cv
21 Apr 2026
Research

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

DGX agent

arXiv:2604.17565v1 Announce Type: new Abstract: Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geomet

researcharxiv-cs-cv
21 Apr 2026
Model Releases

UniMamba: A Unified Spatial-Temporal Modeling Framework with State-Space and Attention Integration

DGX agent

arXiv:2604.16325v1 Announce Type: new Abstract: Multivariate time series forecasting is fundamental to numerous domains such as energy, finance, and environmental monitoring, where complex temporal de

model-releasesarxiv-cs-lg
21 Apr 2026
Research

UniMesh: Unifying 3D Mesh Understanding and Generation

DGX agent

arXiv:2604.17472v1 Announce Type: new Abstract: Recent advances in 3D vision have led to specialized models for either 3D understanding (e.g., shape classification, segmentation, reconstruction) or 3D

researcharxiv-cs-cv
21 Apr 2026
Tutorials

UniSim: A Unified Simulator for Time-Coarsened Dynamics of Biomolecules

DGX agent

arXiv:2506.03157v4 Announce Type: replace-cross Abstract: Molecular Dynamics (MD) simulations are essential for understanding the atomic-level behavior of molecular systems, giving insights into their

tutorialsarxiv-cs-lg
21 Apr 2026
Research

Universal Diffusion-Based Probabilistic Downscaling

DGX agent

arXiv:2602.11893v3 Announce Type: replace Abstract: We introduce a universal diffusion-based downscaling framework that lifts deterministic low-resolution weather forecasts into probabilistic high-res

researcharxiv-cs-lg
21 Apr 2026
Model Releases

Universally Empowering Zeroth-Order Optimization via Adaptive Layer-wise Sampling

DGX agent

arXiv:2604.18264v1 Announce Type: new Abstract: Zeroth-Order optimization presents a promising memory-efficient paradigm for fine-tuning Large Language Models by relying solely on forward passes. Howe

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning

DGX agent

arXiv:2603.23404v2 Announce Type: replace-cross Abstract: Existing Multimodal Large Language Models (MLLMs) struggle with 3D spatial reasoning, as they fail to construct structured abstractions of the

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models

DGX agent

arXiv:2604.18000v1 Announce Type: new Abstract: Recent Vision-Language-Action (VLA) models report impressive success rates on standard robotic benchmarks, fueling optimism about general-purpose physic

model-releasesarxiv-cs-ro
21 Apr 2026
Tutorials

Unraveling the Key of Machine Learning-based Android Malware Detection

DGX agent

arXiv:2402.02953v2 Announce Type: replace-cross Abstract: With the rapid advancement of machine learning (ML), ML-based Android malware detection has gained significant popularity due to its ability t

tutorialsarxiv-cs-lg
21 Apr 2026
Model Releases

Unsupervised Discovery of Intermediate Phase Order in the Frustrated J_1-J_2 Heisenberg Model via Prometheus Framework

DGX agent

arXiv:2602.21468v4 Announce Type: replace-cross Abstract: The spin-1/2 J_1-J_2 Heisenberg model on the square lattice exhibits a debated intermediate phase between Neel antiferromagnetic and stripe or

model-releasesarxiv-cs-lg
21 Apr 2026
Safety

Untrained CNNs Match Backpropagation at V1: A Systematic RSA Comparison of Four Learning Rules Against Human fMRI

DGX agent

arXiv:2604.16875v1 Announce Type: new Abstract: A central question in computational neuroscience is whether the learning rule used to train a neural network determines how well its internal representa

safetyarxiv-cs-lg
21 Apr 2026
Model Releases

Unveiling Deepfakes: A Frequency-Aware Triple Branch Network for Deepfake Detection

DGX agent

arXiv:2604.17477v1 Announce Type: new Abstract: Advanced deepfake technologies are blurring the lines between real and fake, presenting both revolutionary opportunities and alarming threats. While it

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Upper Approximation Bounds for Neural Oscillators

DGX agent

arXiv:2512.01015v2 Announce Type: replace Abstract: Neural oscillators, originating from second-order ordinary differential equations (ODEs), have demonstrated strong performance in stably learning ca

researcharxiv-cs-lg
21 Apr 2026
Model Releases

User-Assistant Bias in LLMs

DGX agent

arXiv:2508.15815v3 Announce Type: replace Abstract: Modern large language models (LLMs) are typically trained and deployed using structured role tags (e.g. system, user, assistant, tool) that explicit

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Using large language models for embodied planning introduces systematic safety risks

DGX agent

arXiv:2604.18463v1 Announce Type: cross Abstract: Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe plann

model-releasesarxiv-cs-lg
21 Apr 2026
Research

Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models

DGX agent

arXiv:2506.00065v2 Announce Type: replace Abstract: Multimodal language models (MLMs) increasingly demonstrate human-like communication, yet their use of everyday perspectival words remains poorly und

researcharxiv-cs-cl
21 Apr 2026
Model Releases

VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning

DGX agent

arXiv:2402.13243v2 Announce Type: replace Abstract: Learning a human-like driving policy from large-scale driving demonstrations is promising, but the uncertainty and non-deterministic nature of plann

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Variational Autoencoder Domain Adaptation for Cross-System Generalization in ML-Based SOP Monitoring

DGX agent

arXiv:2604.18035v1 Announce Type: new Abstract: Machine learning (ML) models trained to detect physical-layer threats on one optical fiber system often fail catastrophically when applied to a differen

researcharxiv-cs-lg
21 Apr 2026
Research

VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis

DGX agent

arXiv:2509.16538v3 Announce Type: replace-cross Abstract: We propose VC-Inspector, a lightweight, open-source large multimodal model (LMM) for reference-free evaluation of video captions, with a focus

researcharxiv-cs-cl
21 Apr 2026
Model Releases

VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision

DGX agent

arXiv:2510.27462v2 Announce Type: replace Abstract: Supervised fine-tuning (SFT) on long chain-of-thought (CoT) trajectories has emerged as a crucial technique for enhancing the reasoning abilities of

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

DGX agent

arXiv:2604.17248v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing sp

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Video Panels for Long Video Understanding

DGX agent

arXiv:2509.23724v2 Announce Type: replace Abstract: Recent Video-Language Models (VLMs) achieve promising results on long-video understanding, but their performance still lags behind that achieved on

model-releasesarxiv-cs-cv
21 Apr 2026
Local Ai

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

DGX agent

arXiv:2604.17656v1 Announce Type: cross Abstract: Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typic

local-aiarxiv-cs-cl
21 Apr 2026
Safety

VIDEOP2R: Video Understanding from Perception to Reasoning

DGX agent

arXiv:2511.11113v2 Announce Type: replace Abstract: Reinforcement fine-tuning (RFT), a two-stage framework consisting of supervised fine-tuning (SFT) and reinforcement learning (RL) has shown promisin

safetyarxiv-cs-cv
21 Apr 2026
Local Ai

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning

DGX agent

arXiv:2601.15724v2 Announce Type: replace Abstract: Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning

local-aiarxiv-cs-cv
21 Apr 2026
Model Releases

VIDS: A Verified Imaging Dataset Standard for Medical AI

DGX agent

arXiv:2604.17525v1 Announce Type: cross Abstract: Medical imaging AI development is fundamentally dependent on annotated datasets, yet no existing standard provides machine-enforceable validation acro

model-releasesarxiv-cs-cv
21 Apr 2026
Research

View-Consistent 3D Scene Editing via Dual-Path Structural Correspondense and Semantic Continuity

DGX agent

arXiv:2604.17801v1 Announce Type: new Abstract: Text-driven 3D scene editing has recently attracted increasing attention. Most existing methods follow a render-edit-optimize pipeline, where multi-view

researcharxiv-cs-cv
21 Apr 2026
Research

ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes

DGX agent

arXiv:2604.17623v1 Announce Type: new Abstract: Kinematic rigs provide a structured interface for articulating 3D meshes, but they lack an inherent representation of the plausible manifold of joint co

researcharxiv-cs-cv
21 Apr 2026
Applications

Vision-Braille: A Curriculum Learning Toolkit and Braille-Chinese Corpus for Braille Translation

DGX agent

arXiv:2407.06048v2 Announce Type: replace Abstract: We present Vision-Braille, the first publicly available end-to-end system for translating Chinese Braille extracted from images into written Chinese

applicationsarxiv-cs-cl
21 Apr 2026
Research

Vision Language Models are Biased

DGX agent

arXiv:2505.23941v4 Announce Type: replace-cross Abstract: Large language models (LLMs) memorize a vast amount of prior knowledge from the Internet that helps them on downstream tasks but also may noto

researcharxiv-cs-cv
21 Apr 2026
Applications

Visual-RRT: Finding Paths toward Visual-Goals via Differentiable Rendering

DGX agent

arXiv:2604.16388v1 Announce Type: cross Abstract: Rapidly-exploring random trees (RRTs) have been widely adopted for robot motion planning due to their robustness and theoretical guarantees. However,

applicationsarxiv-cs-cv
21 Apr 2026
Research

ViT^3: Unlocking Test-Time Training in Vision

DGX agent

arXiv:2512.01643v2 Announce Type: replace Abstract: Test-Time Training (TTT) has recently emerged as a promising direction for efficient sequence modeling. TTT reformulates attention operation as an o

researcharxiv-cs-cv
21 Apr 2026
Model Releases

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

DGX agent

arXiv:2505.20279v4 Announce Type: replace-cross Abstract: The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes,

model-releasesarxiv-cs-cl
21 Apr 2026
Research

Vocab Diet: Reshaping the Vocabulary of LLMs via Vector Arithmetic

DGX agent

arXiv:2510.17001v2 Announce Type: replace Abstract: Large language models (LLMs) often encode word-form variation (e.g., walk vs. walked) as linear directions in the embedding space. However, standard

researcharxiv-cs-cl
21 Apr 2026
Local Ai

VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models

DGX agent

arXiv:2508.15229v3 Announce Type: replace Abstract: Small Language Models (SLMs) provide computational advantages in resource-constrained environments, yet memory limitations remain a critical bottlen

local-aiarxiv-cs-cl
21 Apr 2026
Model Releases

Voronoi-guided Bilateral 2D Gaussian Splatting for Arbitrary-Scale Hyperspectral Image Super-Resolution

DGX agent

arXiv:2604.17727v1 Announce Type: new Abstract: Most existing hyperspectral image super-resolution methods require modifications for different scales, limiting their flexibility in arbitrary-scale rec

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception

DGX agent

arXiv:2604.17475v1 Announce Type: cross Abstract: Small Vision-Language Models (SVLMs) are efficient task controllers but often suffer from visual brittleness and poor tool orchestration. They typical

safetyarxiv-cs-cl
21 Apr 2026
Research

Wasserstein Distributionally Robust Risk-Sensitive Estimation via Conditional Value-at-Risk

DGX agent

arXiv:2604.18546v1 Announce Type: new Abstract: We propose a distributionally robust approach to risk-sensitive estimation of an unknown signal x from an observed signal y. The unknown signal and obse

researcharxiv-cs-lg
21 Apr 2026
← Previous
1…11281129113011311132…1247
Next →