AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer

DGX agent

arXiv:2512.21078v3 Announce Type: replace Abstract: Visual Place Recognition (VPR) has been traditionally formulated as a single-image retrieval task. Using multiple views offers clear advantages, yet

researcharxiv-cs-cv
30 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

UniTriSplat: A Unified 3D Gaussian Splatting Framework with Uniform Spherical Rasterization for Universal Cameras

DGX agent

arXiv:2606.29794v1 Announce Type: new Abstract: Existing 3D Gaussian Splatting (3DGS) frameworks rely on camera-specific rasterization, suffering from inconsistent solid-angle sampling and degraded pe

researcharxiv-cs-cv
30 Jun 2026
Research

UniVAD v2: Unified Visual Anomaly Detection via Support-Conditioned Boundary Construction

DGX agent

arXiv:2606.29714v1 Announce Type: new Abstract: Unified visual anomaly detection seeks to train a single detector that can be deployed across categories, domains, and application scenarios. In the few

researcharxiv-cs-cv
30 Jun 2026
Model Releases

UrbanCDNet: Appearance-Robust and Boundary-Aware Bitemporal Change Detection for Korean Urban Building Monitoring

DGX agent

arXiv:2606.29781v1 Announce Type: new Abstract: Urban building change detection from bi-temporal aerial imagery is important for redevelopment monitoring, infrastructure management, and unauthorized-c

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Variance Reduction on the Camera Axis: Multi-View Score Distillation for 3D

DGX agent

arXiv:2606.29964v1 Announce Type: new Abstract: Score distillation turns a pretrained 2D diffusion model into a 3D generator, but the per-step gradient is estimated from a single randomly chosen view:

model-releasesarxiv-cs-cv
30 Jun 2026
Applications

VCS-SLAM: Geometry-Validated Semantic Evidence Fusion for 3D Gaussian SLAM

DGX agent

arXiv:2606.29494v1 Announce Type: new Abstract: Visual SLAM performance often deteriorates in complex real-world applications. Semantic 3D Gaussian SLAM commonly fuses 2D semantic priors into a persis

applicationsarxiv-cs-cv
30 Jun 2026
Research

VIB-AVSR: Variational Information Bottleneck for Noise-Robust LLM-Based Audio-Visual Speech Recognition

DGX agent

arXiv:2606.29632v1 Announce Type: cross Abstract: Audio-Visual Speech Recognition takes two input modalities, acoustic and visual streams, where visual information from lip movements aids recognition

researcharxiv-cs-cv
30 Jun 2026
Applications

VibES: Induced Vibration for Persistent Event-Based Sensing

DGX agent

arXiv:2508.19094v3 Announce Type: replace Abstract: Event cameras are a bio-inspired class of sensors that asynchronously measure per-pixel intensity changes. Under fixed illumination conditions in st

applicationsarxiv-cs-cv
30 Jun 2026
Research

ViewSplat: View-Adaptive 3D Gaussian Splatting for Feed-Forward Synthesis

DGX agent

arXiv:2603.25265v2 Announce Type: replace Abstract: We present ViewSplat, a view-adaptive 3D Gaussian splatting network for novel view synthesis from unposed images. While recent feed-forward 3D Gauss

researcharxiv-cs-cv
30 Jun 2026
Model Releases

VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake Detection

DGX agent

arXiv:2603.21526v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) offer a promising path toward interpretable deepfake detection by generating textual explanations. However,

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models

DGX agent

arXiv:2606.28804v1 Announce Type: new Abstract: Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for advancing embodied intelligence, enabling the safety-critical evaluat

model-releasesarxiv-cs-cv
30 Jun 2026
Research

Virtual Ring Try-On

DGX agent

arXiv:2606.28792v1 Announce Type: new Abstract: This paper presents an innovative approach that enables the users to capture their hand and try the jewel ring on their hand. The user captures the imag

researcharxiv-cs-cv
30 Jun 2026
Safety

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs

DGX agent

arXiv:2606.28401v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown strong performance in visual understanding, yet they still suffer from hallucinations, generating content that

safetyarxiv-cs-cv
30 Jun 2026
Safety

Vision-Language-Action Models: Experimental Insights from a Real-World UR5 Platform

DGX agent

arXiv:2606.30456v1 Announce Type: cross Abstract: This project investigates whether recent Vision-Language-Action (VLA) models can be transferred from controlled research benchmarks to a real-world ro

safetyarxiv-cs-cv
30 Jun 2026
Local Ai

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context

DGX agent

arXiv:2606.30288v1 Announce Type: new Abstract: Large Vision Language Models (LVLMs) have achieved remarkable success on vision-language tasks, yet fine-grained perception over high-resolution images

local-aiarxiv-cs-cv
30 Jun 2026
Safety

Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration

DGX agent

arXiv:2508.14483v4 Announce Type: replace Abstract: We present Vivid-VR, a DiT-based generative video restoration method built upon an advanced T2V foundation model, where ControlNet is leveraged to c

safetyarxiv-cs-cv
30 Jun 2026
Safety

VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors

DGX agent

arXiv:2510.00458v3 Announce Type: replace Abstract: Vision-language object detectors (VLODs) such as YOLO-World and Grounding DINO exhibit strong zero-shot generalization, but their performance degrad

safetyarxiv-cs-cv
30 Jun 2026
Model Releases

VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On

DGX agent

arXiv:2603.11734v2 Announce Type: replace Abstract: As virtual try-on (VTON) continues to advance, a growing number of real-world scenarios have emerged, pushing beyond the ability of the existing spe

model-releasesarxiv-cs-cv
30 Jun 2026
Research

W4A4 Quantization for Inference on Wan2.2-I2V-A14B

DGX agent

arXiv:2606.29337v1 Announce Type: new Abstract: We summarize our submission to Sub-Challenge 1: W4A4 Quantization for Inference (HiF4 / MXFP4) of the ICME 2026 Low-Bit-width Large-Model Quantization C

researcharxiv-cs-cv
30 Jun 2026
Local Ai

Walking in the Implicit: Interactive World Exploration via Neural Scene Representation

DGX agent

arXiv:2606.30045v1 Announce Type: new Abstract: Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent video frames, entangling state transit

local-aiarxiv-cs-cv
30 Jun 2026
Research

What Color is the Sky (for a non-human) ?

DGX agent

arXiv:2606.28912v1 Announce Type: new Abstract: The light of the daytime sky contains a mixture of many colors yet is perceived as blue by human observers. This is largely due to the particular respon

researcharxiv-cs-cv
30 Jun 2026
Research

When Does Synthetic CT Transfer? A Label-Free Donor/Host Diagnostic for Medical Vision-Language Model Routing on Real Lung CT

DGX agent

arXiv:2606.29232v1 Announce Type: new Abstract: A synthetic measurement of model competence is useful only if it survives the move to real data, yet the real labels that would verify it are exactly wh

researcharxiv-cs-cv
30 Jun 2026
Model Releases

XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity

DGX agent

arXiv:2506.00599v3 Announce Type: replace Abstract: While current 6D pose estimation benchmarks have reached near-saturation on household objects, they often fail to capture the stochastic and optical

model-releasesarxiv-cs-cv
30 Jun 2026
Local Ai

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation

DGX agent

arXiv:2606.30248v1 Announce Type: new Abstract: Recent text-to-video (T2V) diffusion models rely heavily on auxiliary reward signals (e.g., via reward models or DPO) to align generated content with hu

local-aiarxiv-cs-cv
30 Jun 2026
Model Releases

Zero-Gated Language-conditioned Human Motion Prediction

DGX agent

arXiv:2606.29208v1 Announce Type: new Abstract: Pose histories provide the core kinematic evidence for 3D human motion prediction, but they lack explicit high-level semantic guidance. This paper intro

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

Zero-Label Driving Scenario Complexity Detection via Joint Embedding Predictive Architecture

DGX agent

arXiv:2606.28383v1 Announce Type: new Abstract: Identifying complex and safety-critical driving scenarios in large unlabelled datasets is an important but expensive problem. Existing approaches rely o

safetyarxiv-cs-cv
30 Jun 2026
Model Releases

Zero-Shot Depth from Defocus

DGX agent

arXiv:2603.26658v2 Announce Type: replace Abstract: Depth from Defocus (DfD) is the task of estimating a dense metric depth map from a focus stack. Unlike previous works overfitting to a certain datas

model-releasesarxiv-cs-cv
30 Jun 2026
Agents

A Comprehensive Survey on World Models for Embodied AI

DGX agent

arXiv:2510.16732v3 Announce Type: replace Abstract: Embodied AI requires agents that perceive, act, and anticipate how actions reshape future world states. World models serve as internal simulators th

agentsarxiv-cs-cv
29 Jun 2026
Model Releases

A Multi-Attribute Latent Space for Visual Analysis of Watches

DGX agent

arXiv:2606.27897v1 Announce Type: new Abstract: We present a design rationale, embedding model, and interactive visual-analysis system for exploring large wristwatch collections through heterogeneous

model-releasesarxiv-cs-cv
29 Jun 2026
Model Releases

A Unified Framework for Vision Transformers Equivariant to Discrete Subgroups of O(2)

DGX agent

arXiv:2606.27864v1 Announce Type: new Abstract: Vision transformers have become a dominant architecture for visual recognition. However, standard models do not explicitly encode the planar symmetries

model-releasesarxiv-cs-cv
29 Jun 2026
Research

AI-Generated Image Recognition via Fusion of CNNs and Vision Transformers

DGX agent

arXiv:2606.27637v1 Announce Type: new Abstract: Recent advancements in synthetic data technology have opened a new era where images of remarkable quality are generated, blurring the lines between real

researcharxiv-cs-cv
29 Jun 2026
Model Releases

AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration

DGX agent

arXiv:2606.28049v1 Announce Type: new Abstract: In recent years, multimodal large language models (MLLMs) have shown strong potential for embodied intelligence, yet their ability to maintain geometric

model-releasesarxiv-cs-cv
29 Jun 2026
Research

An Approach to Enriching Surgical Video Datasets for Fine-Grained Spatial-Temporal Understanding of Vision-Language Models

DGX agent

arXiv:2604.00784v2 Announce Type: replace Abstract: Surgical video understanding is a crucial prerequisite for advancing Computer-Assisted Surgery. While vision-language models (VLMs) have recently be

researcharxiv-cs-cv
29 Jun 2026
Research

An Embedded Real-Time License Plate Recognition System for Complex Traffic Scenes

DGX agent

arXiv:2606.27772v1 Announce Type: new Abstract: Vehicle license plate recognition is an integral component of intelligent transportation systems. In this work, we present an embedded real-time license

researcharxiv-cs-cv
29 Jun 2026
Research

An Expectation-Maximization Algorithm for Training Clean Diffusion Models from Corrupted Observations

DGX agent

arXiv:2407.01014v2 Announce Type: replace Abstract: Diffusion models excel in solving imaging inverse problems due to their ability to model complex image priors. However, their reliance on large, cle

researcharxiv-cs-cv
29 Jun 2026
Research

Beyond MoCap: Scaling Motion Tokenizers with Synthetic Human Motion for Generative Modeling

DGX agent

arXiv:2606.27547v1 Announce Type: new Abstract: Human motion generation models are fundamentally constrained by the limited diversity of motion capture datasets, which predominantly contain common, re

researcharxiv-cs-cv
29 Jun 2026
Research

Beyond Points: Spherical Distributional Part Prototypes for Interpretable Classification

DGX agent

arXiv:2606.27582v1 Announce Type: new Abstract: Prototype-based neural networks aim to provide intrinsic interpretability by grounding predictions in a small set of part prototypes. However, modern vi

researcharxiv-cs-cv
29 Jun 2026
Local Ai

Beyond Sequential Distance: Inter-Modal Distance Invariant Position Encoding

DGX agent

arXiv:2603.10863v2 Announce Type: replace Abstract: Despite the remarkable capabilities of Multimodal Large Language Models (MLLMs), they still suffer from visual fading in long-context scenarios. Spe

local-aiarxiv-cs-cv
29 Jun 2026
Agents

CascadeOcc: Rethinking 3D Occupancy World Models with Cascaded VQ Representations

DGX agent

arXiv:2606.27644v1 Announce Type: new Abstract: This letter proposes CascadeOcc, a novel occupancy world model that prioritizes intrinsic structural hierarchy over extrinsic auxiliary modalities for a

agentsarxiv-cs-cv
29 Jun 2026
Model Releases

Complex-Valued 2D Gaussian Representation for Computer-Generated Holography

DGX agent

arXiv:2511.15022v2 Announce Type: replace Abstract: Complex-valued Gaussian primitives have recently been explored for representing holographic radiance fields in 3D novel view synthesis. In this work

model-releasesarxiv-cs-cv
29 Jun 2026
Safety

Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding

DGX agent

arXiv:2604.02546v2 Announce Type: replace Abstract: Pretraining 3D encoders by aligning with Contrastive Language Image Pretraining (CLIP) has emerged as a promising direction to learn generalizable r

safetyarxiv-cs-cv
29 Jun 2026
Research

Controllable Histopathology Image Synthesis with Training-free Structural Initialization and Textural Modulation

DGX agent

arXiv:2606.27935v1 Announce Type: new Abstract: Deep learning has demonstrated remarkable success in high-throughput histopathology image analysis. However, the performance of learning-based models cr

researcharxiv-cs-cv
29 Jun 2026
Research

Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training

DGX agent

arXiv:2606.28104v1 Announce Type: new Abstract: Vision-based assessment can provide convenient and cost-effective evaluation in Traditional Chinese Medicine (TCM) rehabilitation training, where action

researcharxiv-cs-cv
29 Jun 2026
Safety

CSD: Content-aware Speculative Decoding for Efficient Image Generation

DGX agent

arXiv:2606.27829v1 Announce Type: new Abstract: Speculative decoding (SD) has emerged as a key solution to accelerate the inference of autoregressive models. However, in the field of image generation,

safetyarxiv-cs-cv
29 Jun 2026
Model Releases

Curriculum-guided Change Detection Training: Toward Accurate Serac Fall Monitoring

DGX agent

arXiv:2606.28012v1 Announce Type: new Abstract: Change Detection (CD) aims to identify semantic or structural changes from nearly registered multi-temporal images. While recent advances in training me

model-releasesarxiv-cs-cv
29 Jun 2026
Tutorials

DeLux: Cross-Modal Local Artifact Restoration in Video Using Neuromorphic Data

DGX agent

arXiv:2606.27576v1 Announce Type: new Abstract: Conventional RGB cameras suffer from lighting artifacts such as flare, glare, flicker, and overexposure, leading to irrecoverable information loss that

tutorialsarxiv-cs-cv
29 Jun 2026
Research

Denoising ICF Images with Multiplicative Uniform Noise: A Self-Supervised Study Based on the Log-Domain Noisier2Inverse Framework

DGX agent

arXiv:2606.27635v1 Announce Type: new Abstract: This paper documents the implementation and evaluation of a self-supervised denoising framework on Inertial Confinement Fusion (ICF) images corrupted by

researcharxiv-cs-cv
29 Jun 2026
Research

Differentiable design of the PIAA-ZWFS: a flexible wavefront sensor that approaches the fundamental limit

DGX agent

arXiv:2606.28136v1 Announce Type: cross Abstract: Extreme adaptive optics (AO) is necessary for high contrast astronomy at scales of the habitable zone of nearby systems. We seek to evaluate wavefront

researcharxiv-cs-cv
29 Jun 2026
← Previous
1…8384858687…263
Next →