AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Detail Consistent Stage-Wise Distillation for Efficient 3D MRI Segmentation

DGX agent

arXiv:2605.26382v1 Announce Type: new Abstract: Deploying high-performing 3D medical image segmenters (e.g., nnU-Net) is often limited by memory footprint and inference latency. Compression is therefo

researcharxiv-cs-cv
27 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Dimensional Distribution Emotion State: Leveraging Valence and Arousal as a Common Embedding Space for Visual Emotion Analysis

DGX agent

arXiv:2605.26262v1 Announce Type: new Abstract: Museums are important sites for the dissemination of culture and art. They are institutions rooted in history and tradition; their exhibitions are often

safetyarxiv-cs-cv
27 May 2026
Applications

DinoComplete: 3D Shape Completion with Distilled Semantic Priors and State Space Models

DGX agent

arXiv:2605.26949v1 Announce Type: new Abstract: 3D shape completion from partial scans remains challenging for unseen categories and noisy real-world observations, where geometry alone is often insuff

applicationsarxiv-cs-cv
27 May 2026
Research

DirectFisheye-GS: Enabling Native Fisheye Input in Gaussian Splatting with Cross-View Joint Optimization

DGX agent

arXiv:2604.00648v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has enabled efficient 3D scene reconstruction from everyday images with real-time, high-fidelity rendering, greatly adv

researcharxiv-cs-cv
27 May 2026
Applications

Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?

DGX agent

arXiv:2605.27135v1 Announce Type: cross Abstract: With the rapid proliferation of generative models, such as diffusion models, digital watermarking has emerged as a crucial solution for identifying AI

applicationsarxiv-cs-cv
27 May 2026
Research

Dual-Thresholded Heatmap-Guided Proposal Clustering and Negative Certainty Supervision with Enhanced Base Network for Weakly Supervised Object Detection

DGX agent

arXiv:2509.08289v3 Announce Type: replace Abstract: Weakly supervised object detection (WSOD) has attracted significant attention in recent years, as it does not require box-level annotations. State-o

researcharxiv-cs-cv
27 May 2026
Safety

DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation

DGX agent

arXiv:2605.26236v1 Announce Type: new Abstract: Co-speech gesture generation requires both semantic expressivity and biomechanically plausible rhythmic motion. Existing holistic gesture models mix lex

safetyarxiv-cs-cv
27 May 2026
Safety

DV-SFT: Direct Vision Supervision for Fine-Grained Visual Understanding

DGX agent

arXiv:2605.26656v1 Announce Type: new Abstract: Multimodal large language models are typically trained end-to-end to predict ground-truth answers, yet supervision signals are applied exclusively to te

safetyarxiv-cs-cv
27 May 2026
Local Ai

Efficient All-Pairs Correlation Volume Sampling for Optical Flow Estimation

DGX agent

arXiv:2505.16942v2 Announce Type: replace Abstract: Recent optical flow estimation methods often employ local cost sampling from a dense all-pairs correlation volume. This results in quadratic computa

local-aiarxiv-cs-cv
27 May 2026
Model Releases

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy

DGX agent

arXiv:2605.24456v2 Announce Type: replace Abstract: Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life.

model-releasesarxiv-cs-cv
27 May 2026
Local Ai

Feedforward 3D Editing Learns from Semantic-Part Transformation

DGX agent

arXiv:2605.27351v1 Announce Type: new Abstract: 3D editing is a fundamental capability for scalable 3D content creation. While image editing has rapidly evolved toward large-scale feedforward generati

local-aiarxiv-cs-cv
27 May 2026
Research

Frequency-Guided Fusion For RGB-Thermal Semantic Segmentation

DGX agent

arXiv:2605.26273v1 Announce Type: new Abstract: Semantic segmentation in complex environments such as urban driving scenes remains challenging under adverse lighting conditions, where RGB images alone

researcharxiv-cs-cv
27 May 2026
Research

From Contrast to Consistency: Rethinking Event-based Continuous-Time Optical Flow Estimation

DGX agent

arXiv:2605.25570v2 Announce Type: replace Abstract: Estimating continuous optical flow is a fundamental yet challenging problem in dynamic visual perception. Event-based cameras, with microsecond late

researcharxiv-cs-cv
27 May 2026
Safety

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling

DGX agent

arXiv:2605.26601v1 Announce Type: new Abstract: Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible trainin

safetyarxiv-cs-cv
27 May 2026
Applications

G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing

DGX agent

arXiv:2605.27372v1 Announce Type: new Abstract: Modern feed-forward 3D reconstruction methods like VGGT predict pixel-aligned pointmaps in camera-centric coordinate frames. However, this choice of coo

applicationsarxiv-cs-cv
27 May 2026
Research

Garment Particles: A 2D--3D Symmetric Garment Representation for Generation and Editing

DGX agent

arXiv:2605.26391v1 Announce Type: cross Abstract: Practical garment design spans two modes: intuitive creation from high-level intent, such as a reference image or text description, and complex low-le

researcharxiv-cs-cv
27 May 2026
Applications

Gaussian-Voxel Duet: A Dual-Scaffolding Hybrid Representation for Fast and Accurate Monocular Surface Reconstruction

DGX agent

arXiv:2605.26616v1 Announce Type: new Abstract: While 3D Gaussian Splatting has achieved remarkable success in photorealistic novel view synthesis, its pursuit of fast and high-fidelity 3D reconstruct

applicationsarxiv-cs-cv
27 May 2026
Model Releases

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

DGX agent

arXiv:2605.27295v1 Announce Type: new Abstract: We introduce Gemini Embedding 2, a native multimodal embedding model that allows embedding video, audio, image, and text modalities in a unified represe

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction

DGX agent

arXiv:2605.26230v1 Announce Type: new Abstract: Multi-view 3D reconstruction has achieved remarkable progress with the advent of feed-forward 3D reconstruction models. However, these models are typica

model-releasesarxiv-cs-cv
27 May 2026
Tutorials

GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision

DGX agent

arXiv:2603.09551v2 Announce Type: replace Abstract: While Vision-Language Models (VLMs) have significantly advanced remote sensing interpretation, enabling them to perform complex, step-by-step reason

tutorialsarxiv-cs-cv
27 May 2026
Research

Global Structure-from-Motion Meets Feedforward Reconstruction

DGX agent

arXiv:2605.26103v2 Announce Type: replace Abstract: Structure-from-Motion -- the process of simultaneously estimating camera poses and 3D scene structure from a collection of images -- remains a centr

researcharxiv-cs-cv
27 May 2026
Research

GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning

DGX agent

arXiv:2602.19206v3 Announce Type: replace Abstract: Zero-shot 3D Anomaly Detection is an emerging task that aims to detect anomalies in a target dataset without any target training data, which is part

researcharxiv-cs-cv
27 May 2026
Model Releases

Guiding Token-Sparse Diffusion Models

DGX agent

arXiv:2601.01608v2 Announce Type: replace Abstract: Diffusion models deliver high quality in image synthesis but remain expensive during training and inference. Recent works have leveraged the inheren

model-releasesarxiv-cs-cv
27 May 2026
Tutorials

How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning

DGX agent

arXiv:2605.27310v1 Announce Type: new Abstract: Cross-view spatial reasoning remains a weak spot for vision-language models (VLMs): they often reason in language and lose the fine-grained geometry nee

tutorialsarxiv-cs-cv
27 May 2026
Research

HydraPrompt: An Adaptive and Asymmetric Framework of Vision-Language Models for Synthetic Image Detection

DGX agent

arXiv:2605.26421v1 Announce Type: new Abstract: The rapid evolution of generative models has precipitated a proliferation of fabricated content, posing significant challenges to existing Synthetic Ima

researcharxiv-cs-cv
27 May 2026
Tutorials

I2PRef: Image-Driven Point Completion with Iterative Refinement

DGX agent

arXiv:2605.26914v1 Announce Type: new Abstract: We present an image-conditioned point cloud completion approach that treats images as the primary geometric source rather than a secondary guide. To thi

tutorialsarxiv-cs-cv
27 May 2026
Safety

Image Thresholding: Understanding Bias of Evaluation Metrics towards Specific Evaluation Functions

DGX agent

arXiv:2605.27132v1 Announce Type: new Abstract: Multilevel image thresholding is widely used for segmentation in applications ranging from medical imaging to remote sensing. Classical objective functi

safetyarxiv-cs-cv
27 May 2026
Model Releases

ImViD: Immersive Volumetric Videos for Enhanced VR Engagement

DGX agent

arXiv:2503.14359v2 Announce Type: replace Abstract: User engagement is greatly enhanced by fully immersive multi-modal experiences that combine visual and auditory stimuli. Consequently, the next fron

model-releasesarxiv-cs-cv
27 May 2026
Research

Innovative Silicosis and Pneumonia Classification: Leveraging Graph Transformer Post-hoc Modeling and Ensemble Techniques

DGX agent

arXiv:2501.00520v2 Announce Type: replace Abstract: This paper presents a comprehensive study on the classification and detection of Silicosis-related lung inflammation. Our main contributions include

researcharxiv-cs-cv
27 May 2026
Safety

Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification

DGX agent

arXiv:2510.00902v2 Announce Type: replace Abstract: Transfer learning is crucial for medical imaging, yet the selection of source datasets often relies on researchers' intuition rather than systematic

safetyarxiv-cs-cv
27 May 2026
Model Releases

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams

DGX agent

arXiv:2605.27074v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) achieve strong performance on reactive question answering, but real-world streaming assistants require p

model-releasesarxiv-cs-cv
27 May 2026
Research

Is an Image Also Worth 16x16=256 Superpixels? A Framework for Attentional Image Classification

DGX agent

arXiv:2605.27144v1 Announce Type: new Abstract: Superpixel-based image classification has traditionally leveraged graph neural networks (GNNs) for processing irregular image representations. Recent ad

researcharxiv-cs-cv
27 May 2026
Research

ISTASTrack: Bridging ANN and SNN via ISTA Adapter for RGB-Event Tracking

DGX agent

arXiv:2509.09977v2 Announce Type: replace Abstract: RGB-Event tracking has become a promising trend in visual object tracking to leverage the complementary strengths of both RGB images and dynamic spi

researcharxiv-cs-cv
27 May 2026
Research

JLT: Clean-Latent Prediction in Latent Diffusion Transformers

DGX agent

arXiv:2605.27102v1 Announce Type: new Abstract: Flow matching with clean-data prediction has shown that regressing the clean point can exploit low-dimensional structure more effectively than predictin

researcharxiv-cs-cv
27 May 2026
Tutorials

Joint 2D-3D Segmentation and Association in Street-level Imaging

DGX agent

arXiv:2605.26725v1 Announce Type: new Abstract: Accurate interpretation of street-level imagery is essential for large-scale urban mapping and the creation of Spatial Digital Twin (SDT) environments.

tutorialsarxiv-cs-cv
27 May 2026
Research

Joint Instance Segmentation and Geometric Attribute Regression for Roof Structures in Aerial Imagery

DGX agent

arXiv:2605.26370v1 Announce Type: new Abstract: We present a method for jointly predicting instance-level roof segment masks together with three continuous geometric attributes -- building height, roo

researcharxiv-cs-cv
27 May 2026
Local Ai

LDP-Slicing: Local Differential Privacy for Images via Randomized Bit-Plane Slicing

DGX agent

arXiv:2603.03711v3 Announce Type: replace Abstract: Local Differential Privacy (LDP) is the gold standard trust model for privacy-preserving machine learning by guaranteeing privacy at the data source

local-aiarxiv-cs-cv
27 May 2026
Research

Learning Reference-Guided Exposure Correction with Hybrid Illumination Characteristics

DGX agent

arXiv:2605.26729v1 Announce Type: new Abstract: We present HICNet, a reference-guided exposure correction framework. A lightweight, content-agnostic encoder distills each image into a compact illumina

researcharxiv-cs-cv
27 May 2026
Tutorials

Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking

DGX agent

arXiv:2605.26933v1 Announce Type: new Abstract: Unsupervised visual object tracking is a challenging task that requires following arbitrary targets in videos without training on ground-truth annotatio

tutorialsarxiv-cs-cv
27 May 2026
Applications

Leveraging Visual Signals for Robust Token-Level Uncertainty in Vision-Language Generation

DGX agent

arXiv:2605.27136v1 Announce Type: new Abstract: Uncertainty quantification (UQ) remains a critical challenge in Large Vision Language Models (LVLMs) for reliable predictions and real-world deployment.

applicationsarxiv-cs-cv
27 May 2026
Tutorials

Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models

DGX agent

arXiv:2506.11253v2 Announce Type: replace Abstract: Machine unlearning removes certain training data points and their influence from AI models (e.g., when a data owner revokes their consent to allow m

tutorialsarxiv-cs-cv
27 May 2026
Research

LiM-YOLO: Less is More with Pyramid Level Shift for Ship Detection in Optical Remote Sensing

DGX agent

arXiv:2512.09700v3 Announce Type: replace Abstract: General-purpose object detectors face fundamental structural limitations when applied to ship detection in satellite imagery, where the ship scale d

researcharxiv-cs-cv
27 May 2026
Model Releases

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV

DGX agent

arXiv:2605.26244v1 Announce Type: new Abstract: Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to sho

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

LongCat-Video-Avatar 1.5 Technical Report

DGX agent

arXiv:2605.26486v1 Announce Type: new Abstract: Despite advances in audio-driven video generation, achieving commercial-grade stability remains challenging. We present LongCat-Video-Avatar 1.5, an upg

model-releasesarxiv-cs-cv
27 May 2026
Applications

LuxRemix: Lighting Decomposition and Remixing for Indoor Scenes

DGX agent

arXiv:2601.15283v2 Announce Type: replace Abstract: We present a novel approach for interactive light editing in indoor scenes from a single multi-view scene capture. Our method leverages a generative

applicationsarxiv-cs-cv
27 May 2026
Research

Memory-Distilled Selection for Noise-Robust Anomaly Detection

DGX agent

arXiv:2605.26676v1 Announce Type: new Abstract: Anomaly detection (AD) under data contamination is critical for deploying unsupervised defect detection in industrial environments, where curating perfe

researcharxiv-cs-cv
27 May 2026
Model Releases

METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition

DGX agent

arXiv:2605.26712v1 Announce Type: new Abstract: Benchmarks that reflect the diversity and complexity of real-world documents are essential for accurately evaluating Automatic Text Recognition (ATR) sy

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Model discovery for dynamical systems with complex-valued product units

DGX agent

arXiv:2605.27158v1 Announce Type: new Abstract: Discovering the governing equations of a dynamical system from observed trajectories provides deeper insight into its structure than mere prediction of

model-releasesarxiv-cs-cv
27 May 2026
← Previous
1…141142143144145…263
Next →