AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
30 Jun 2026

Stability and Concentration in Nonlinear Inverse Problems with Block-Structured Parameters: Lipschitz Geometry, Identifiability, and an Application to Gaussian Splatting

Model ReleasesDGX agent

arXiv:2602.09415v2 Announce Type: replace Abstract: We develop an operator-theoretic framework for stability and statistical concentration in nonlinear inverse problems with block-structured parameter

Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging

Model ReleasesDGX agent

arXiv:2512.01461v2 Announce Type: replace-cross Abstract: Model merging has emerged as a promising paradigm for enabling multi-task capabilities without additional training. However, traditional basic

StereoGS: Sparse-View 3D Gaussian Splatting via Stereo Priors

Research

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2606.30545v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has achieved remarkable success in real-time novel view synthesis, yet it suffers from severe overfitting under sparse-view

Stimulus Motion Perception Studies Imply Specific Neural Computations in Human Visual Stabilization

ResearchDGX agent

arXiv:2506.13506v3 Announce Type: replace Abstract: Even during fixation the human eye is constantly in low amplitude motion, jittering over small angles in random directions at up to 100Hz. This moti

Stochastic Optimal Control Sampling for Diffusion Inverse Problems

ResearchDGX agent

arXiv:2606.28785v1 Announce Type: new Abstract: Benefiting from the strong ability to capture data distributions, diffusion models have become powerful tools for solving image inverse problems. The ke

StrucTab: A Structured Optimization Framework for Table Parsing

Model ReleasesDGX agent

arXiv:2606.29905v1 Announce Type: new Abstract: Table parsing aims to convert table images into structured, machine-readable representations, a task requiring the joint perception of complex spatial l

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance

Model ReleasesDGX agent

arXiv:2603.12703v3 Announce Type: replace Abstract: Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video u

T2LDM++: A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation

TutorialsDGX agent

arXiv:2606.30147v1 Announce Type: new Abstract: Recent progress in Text-to-Image generation benefits from large-scale Text-Image pairs. However, the scarcity of Text-LiDAR pairs often causes over-smoo

TerraDiT: Point-Conditioned Diffusion Transformer for Satellite Image Synthesis

SafetyDGX agent

arXiv:2603.02172v2 Announce Type: replace Abstract: We introduce TerraDiT, a diffusion transformer designed for text-to-satellite image generation with point-based control. Existing controlled satelli

Text-Conditioned Background Generation for Editable Multi-Layer Documents

ResearchDGX agent

arXiv:2512.17151v2 Announce Type: replace Abstract: We present a framework for document-centric background generation with multi-page editing and thematic continuity. To ensure text regions remain rea

The Calibrated Deepfake Trust Score (CDTS): Competence-Coupled Trust Degradation Across Deepfake Detectors

ResearchDGX agent

arXiv:2606.29484v1 Announce Type: cross Abstract: Modern deepfake detectors are rarely consumed as bare classifiers. In moderation, provenance, and verification pipelines their output probability is r

The Platonic Defense: Backdoor Defense for Self-Supervised Encoders in the Era of Large Scale Pre-training

ResearchDGX agent

arXiv:2606.29451v1 Announce Type: new Abstract: Self-supervised learning (SSL) pretrained models have become a dominant paradigm for visual representation learning, but they are vulnerable to backdoor

The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction

ResearchDGX agent

arXiv:2606.30308v1 Announce Type: new Abstract: 4D hand motion reconstruction from egocentric video is bottlenecked by clear limitations of existing methods: image-based pipelines depend on a detector

Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation

Model ReleasesDGX agent

arXiv:2606.30598v1 Announce Type: new Abstract: Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing lea

Towards Long-Form Spatio-Temporal Video Grounding

Local AiDGX agent

arXiv:2602.23294v2 Announce Type: replace Abstract: In real scenarios, videos can span several minutes or even hours. However, existing research on spatio-temporal video grounding (STVG), given a text

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

Model ReleasesDGX agent

arXiv:2512.13660v3 Announce Type: replace-cross Abstract: Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded

TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment

ResearchDGX agent

arXiv:2606.30313v1 Announce Type: new Abstract: Longitudinal glioblastoma response assessment requires comparing subtle tumor changes across MRI time points using structured clinical criteria such as

Traffic-CBM: A Structurally Interpretable Multimodal Framework for Encrypted Traffic Classification

TutorialsDGX agent

arXiv:2606.29909v1 Announce Type: new Abstract: Encrypted traffic classification has achieved strong performance, but its decision process remains difficult to interpret. Existing methods usually comb

TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation

SafetyDGX agent

arXiv:2606.29097v1 Announce Type: new Abstract: Recent research has investigated the use of large language models (LLMs) to generate traffic scenarios for autonomous driving. However, pretrained LLMs

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision

Model ReleasesDGX agent

arXiv:2606.30552v1 Announce Type: cross Abstract: Cross-embodiment transfer in vision-language-action (VLA) models remains challenging because low-level state and action spaces differ fundamentally ac

Transition-Aware best-of-N sampling for Longitudinal Chest X-ray Reports

ResearchDGX agent

arXiv:2606.28393v1 Announce Type: new Abstract: In longitudinal clinical practice, every chest X-ray is read in the context of the patients prior exam, and much of what the radiologist communicates is

TUGS: Physics-based Compact Representation of Underwater Scenes by Tensorized Gaussian

ApplicationsDGX agent

arXiv:2505.08811v3 Announce Type: replace Abstract: Underwater 3D scene reconstruction is crucial for multimedia applications in adverse environments, such as underwater robotic perception and navigat

Tumor-aware augmentation with task-guided attention analysis improves rectal cancer segmentation from magnetic resonance images

Model ReleasesDGX agent

arXiv:2605.05522v2 Announce Type: replace-cross Abstract: Although self-supervised pretraining is expected to learn broadly transferable representations, its effectiveness across imaging modalities su

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models

ApplicationsDGX agent

arXiv:2602.22960v2 Announce Type: replace Abstract: World models based on video generation demonstrate remarkable potential for simulating interactive environments yet suffer from persistent difficult

UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention

Local AiDGX agent

arXiv:2510.16325v3 Announce Type: replace Abstract: Ultra-high-resolution text-to-image generation is increasingly vital for applications requiring fine-grained textures and global structural fidelity

Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning

ResearchDGX agent

arXiv:2606.30020v1 Announce Type: new Abstract: Pathology foundation models (PFMs) offer generalizable representations for whole-slide image (WSI) analysis, yet their clinical adoption remains limited

UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image

AgentsDGX agent

arXiv:2606.30608v1 Announce Type: new Abstract: Articulated 3D objects are essential for interactive environments in embodied AI, robotics, and virtual reality, but reconstructing their structure and

UniCA: Bi-directional Cross-Attention with Positive Similarity Loss for Robust Multi-Modal Retrieval

Model ReleasesDGX agent

arXiv:2606.28350v1 Announce Type: cross Abstract: Multi-modal retrieval has become increasingly critical for handling the growing volume of integrated visual-textual data in real-world applications, b

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

SafetyDGX agent

arXiv:2606.30332v1 Announce Type: new Abstract: Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing app

UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer

ResearchDGX agent

arXiv:2512.21078v3 Announce Type: replace Abstract: Visual Place Recognition (VPR) has been traditionally formulated as a single-image retrieval task. Using multiple views offers clear advantages, yet

UniTriSplat: A Unified 3D Gaussian Splatting Framework with Uniform Spherical Rasterization for Universal Cameras

ResearchDGX agent

arXiv:2606.29794v1 Announce Type: new Abstract: Existing 3D Gaussian Splatting (3DGS) frameworks rely on camera-specific rasterization, suffering from inconsistent solid-angle sampling and degraded pe

UniVAD v2: Unified Visual Anomaly Detection via Support-Conditioned Boundary Construction

ResearchDGX agent

arXiv:2606.29714v1 Announce Type: new Abstract: Unified visual anomaly detection seeks to train a single detector that can be deployed across categories, domains, and application scenarios. In the few

UrbanCDNet: Appearance-Robust and Boundary-Aware Bitemporal Change Detection for Korean Urban Building Monitoring

Model ReleasesDGX agent

arXiv:2606.29781v1 Announce Type: new Abstract: Urban building change detection from bi-temporal aerial imagery is important for redevelopment monitoring, infrastructure management, and unauthorized-c

Variance Reduction on the Camera Axis: Multi-View Score Distillation for 3D

Model ReleasesDGX agent

arXiv:2606.29964v1 Announce Type: new Abstract: Score distillation turns a pretrained 2D diffusion model into a 3D generator, but the per-step gradient is estimated from a single randomly chosen view:

VCS-SLAM: Geometry-Validated Semantic Evidence Fusion for 3D Gaussian SLAM

ApplicationsDGX agent

arXiv:2606.29494v1 Announce Type: new Abstract: Visual SLAM performance often deteriorates in complex real-world applications. Semantic 3D Gaussian SLAM commonly fuses 2D semantic priors into a persis

VIB-AVSR: Variational Information Bottleneck for Noise-Robust LLM-Based Audio-Visual Speech Recognition

ResearchDGX agent

arXiv:2606.29632v1 Announce Type: cross Abstract: Audio-Visual Speech Recognition takes two input modalities, acoustic and visual streams, where visual information from lip movements aids recognition

VibES: Induced Vibration for Persistent Event-Based Sensing

ApplicationsDGX agent

arXiv:2508.19094v3 Announce Type: replace Abstract: Event cameras are a bio-inspired class of sensors that asynchronously measure per-pixel intensity changes. Under fixed illumination conditions in st

ViewSplat: View-Adaptive 3D Gaussian Splatting for Feed-Forward Synthesis

ResearchDGX agent

arXiv:2603.25265v2 Announce Type: replace Abstract: We present ViewSplat, a view-adaptive 3D Gaussian splatting network for novel view synthesis from unposed images. While recent feed-forward 3D Gauss

VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake Detection

Model ReleasesDGX agent

arXiv:2603.21526v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) offer a promising path toward interpretable deepfake detection by generating textual explanations. However,

ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models

Model ReleasesDGX agent

arXiv:2606.28804v1 Announce Type: new Abstract: Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for advancing embodied intelligence, enabling the safety-critical evaluat

Virtual Ring Try-On

ResearchDGX agent

arXiv:2606.28792v1 Announce Type: new Abstract: This paper presents an innovative approach that enables the users to capture their hand and try the jewel ring on their hand. The user captures the imag

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs

SafetyDGX agent

arXiv:2606.28401v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown strong performance in visual understanding, yet they still suffer from hallucinations, generating content that

Vision-Language-Action Models: Experimental Insights from a Real-World UR5 Platform

SafetyDGX agent

arXiv:2606.30456v1 Announce Type: cross Abstract: This project investigates whether recent Vision-Language-Action (VLA) models can be transferred from controlled research benchmarks to a real-world ro

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context

Local AiDGX agent

arXiv:2606.30288v1 Announce Type: new Abstract: Large Vision Language Models (LVLMs) have achieved remarkable success on vision-language tasks, yet fine-grained perception over high-resolution images

Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration

SafetyDGX agent

arXiv:2508.14483v4 Announce Type: replace Abstract: We present Vivid-VR, a DiT-based generative video restoration method built upon an advanced T2V foundation model, where ControlNet is leveraged to c

VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors

SafetyDGX agent

arXiv:2510.00458v3 Announce Type: replace Abstract: Vision-language object detectors (VLODs) such as YOLO-World and Grounding DINO exhibit strong zero-shot generalization, but their performance degrad

VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On

Model ReleasesDGX agent

arXiv:2603.11734v2 Announce Type: replace Abstract: As virtual try-on (VTON) continues to advance, a growing number of real-world scenarios have emerged, pushing beyond the ability of the existing spe

W4A4 Quantization for Inference on Wan2.2-I2V-A14B

ResearchDGX agent

arXiv:2606.29337v1 Announce Type: new Abstract: We summarize our submission to Sub-Challenge 1: W4A4 Quantization for Inference (HiF4 / MXFP4) of the ICME 2026 Low-Bit-width Large-Model Quantization C

Walking in the Implicit: Interactive World Exploration via Neural Scene Representation

Local AiDGX agent

arXiv:2606.30045v1 Announce Type: new Abstract: Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent video frames, entangling state transit

What Color is the Sky (for a non-human) ?

ResearchDGX agent

arXiv:2606.28912v1 Announce Type: new Abstract: The light of the daytime sky contains a mixture of many colors yet is perceived as blue by human observers. This is largely due to the particular respon

When Does Synthetic CT Transfer? A Label-Free Donor/Host Diagnostic for Medical Vision-Language Model Routing on Real Lung CT

ResearchDGX agent

arXiv:2606.29232v1 Announce Type: new Abstract: A synthetic measurement of model competence is useful only if it survives the move to real data, yet the real labels that would verify it are exactly wh

XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity

Model ReleasesDGX agent

arXiv:2506.00599v3 Announce Type: replace Abstract: While current 6D pose estimation benchmarks have reached near-saturation on household objects, they often fail to capture the stochastic and optical

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation

Local AiDGX agent

arXiv:2606.30248v1 Announce Type: new Abstract: Recent text-to-video (T2V) diffusion models rely heavily on auxiliary reward signals (e.g., via reward models or DPO) to align generated content with hu

Zero-Gated Language-conditioned Human Motion Prediction

Model ReleasesDGX agent

arXiv:2606.29208v1 Announce Type: new Abstract: Pose histories provide the core kinematic evidence for 3D human motion prediction, but they lack explicit high-level semantic guidance. This paper intro

Zero-Label Driving Scenario Complexity Detection via Joint Embedding Predictive Architecture

SafetyDGX agent

arXiv:2606.28383v1 Announce Type: new Abstract: Identifying complex and safety-critical driving scenarios in large unlabelled datasets is an important but expensive problem. Existing approaches rely o

Zero-Shot Depth from Defocus

Model ReleasesDGX agent

arXiv:2603.26658v2 Announce Type: replace Abstract: Depth from Defocus (DfD) is the task of estimating a dense metric depth map from a focus stack. Unlike previous works overfitting to a certain datas

29 Jun 2026

A Comprehensive Survey on World Models for Embodied AI

AgentsDGX agent

arXiv:2510.16732v3 Announce Type: replace Abstract: Embodied AI requires agents that perceive, act, and anticipate how actions reshape future world states. World models serve as internal simulators th

A Multi-Attribute Latent Space for Visual Analysis of Watches

Model ReleasesDGX agent

arXiv:2606.27897v1 Announce Type: new Abstract: We present a design rationale, embedding model, and interactive visual-analysis system for exploring large wristwatch collections through heterogeneous

A Unified Framework for Vision Transformers Equivariant to Discrete Subgroups of O(2)

Model ReleasesDGX agent

arXiv:2606.27864v1 Announce Type: new Abstract: Vision transformers have become a dominant architecture for visual recognition. However, standard models do not explicitly encode the planar symmetries

AI-Generated Image Recognition via Fusion of CNNs and Vision Transformers

ResearchDGX agent

arXiv:2606.27637v1 Announce Type: new Abstract: Recent advancements in synthetic data technology have opened a new era where images of remarkable quality are generated, blurring the lines between real

← Previous
1…6465666768…209
Next →