AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Research

HD-VGGT: High-Resolution Visual Geometry Transformer

DGX agent

arXiv:2603.27222v2 Announce Type: replace Abstract: High-resolution imagery is essential for accurate 3D reconstruction, as many geometric details only emerge at fine spatial scales. Recent feed-forwa

researcharxiv-cs-cv
13 Apr 2026
Safety

Hitem3D 2.0: Multi-View Guided Native 3D Texture Generation

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2604.09231v1 Announce Type: new Abstract: Although recent advances have improved the quality of 3D texture generation, existing methods still struggle with incomplete texture coverage, cross-vie

safetyarxiv-cs-cv
13 Apr 2026
Research

How Noise Benefits AI-generated Image Detection

DGX agent

arXiv:2511.16136v2 Announce Type: replace Abstract: The rapid advancement of generative models has made real and synthetic images increasingly indistinguishable. Although extensive efforts have been d

researcharxiv-cs-cv
13 Apr 2026
Model Releases

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms

DGX agent

arXiv:2604.08966v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have advanced Video Temporal Grounding (VTG), existing methods often couple output paradigms with differe

model-releasesarxiv-cs-cv
13 Apr 2026
Research

Incremental Semantics-Aided Meshing from LiDAR-Inertial Odometry and RGB Direct Label Transfer

DGX agent

arXiv:2604.09478v1 Announce Type: new Abstract: Geometric high-fidelity mesh reconstruction from LiDAR-inertial scans remains challenging in large, complex indoor environments -- such as cultural buil

researcharxiv-cs-cv
13 Apr 2026
Research

InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation

DGX agent

arXiv:2604.08646v1 Announce Type: new Abstract: Instruction-based video editing is a natural way to control video content with text, but adapting a video generation model into an editor usually appear

researcharxiv-cs-cv
13 Apr 2026
Research

Intrinsic Concept Extraction Based on Compositional Interpretability

DGX agent

arXiv:2603.11795v2 Announce Type: replace Abstract: Unsupervised Concept Extraction aims to extract concepts from a single image; however, existing methods suffer from the inability to extract composa

researcharxiv-cs-cv
13 Apr 2026
Research

LoBE-GS: Load-Balanced and Efficient 3D Gaussian Splatting for Large-Scale Scene Reconstruction

DGX agent

arXiv:2510.01767v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has established itself as an efficient representation for real-time, high-fidelity 3D scene reconstruction. However, sc

researcharxiv-cs-cv
13 Apr 2026
Safety

Long-SCOPE: Fully Sparse Long-Range Cooperative 3D Perception

DGX agent

arXiv:2604.09206v1 Announce Type: new Abstract: Cooperative 3D perception via Vehicle-to-Everything communication is a promising paradigm for enhancing autonomous driving, offering extended sensing ho

safetyarxiv-cs-cv
13 Apr 2026
Model Releases

Low-Data Supervised Adaptation Outperforms Prompting for Cloud Segmentation Under Domain Shift

DGX agent

arXiv:2604.08956v1 Announce Type: new Abstract: Adapting vision-language models to remote sensing imagery presents a fundamental challenge: both the visual and linguistic distributions of satellite da

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

LPLCv2: An Expanded Dataset for Fine-Grained License Plate Legibility Classification

DGX agent

arXiv:2604.08741v1 Announce Type: new Abstract: Modern Automatic License Plate Recognition (ALPR) systems achieve outstanding performance in controlled, well-defined scenarios. However, large-scale re

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

LuMon: A Comprehensive Benchmark and Development Suite with Novel Datasets for Lunar Monocular Depth Estimation

DGX agent

arXiv:2604.09352v1 Announce Type: new Abstract: Monocular Depth Estimation (MDE) is crucial for autonomous lunar rover navigation using electro-optical cameras. However, deploying terrestrial MDE netw

model-releasesarxiv-cs-cv
13 Apr 2026
Tutorials

M-IDoL: Information Decomposition for Modality-Specific and Diverse Representation Learning in Medical Foundation Model

DGX agent

arXiv:2604.08936v1 Announce Type: new Abstract: Medical foundation models (MFMs) aim to learn universal representations from multimodal medical images that can generalize effectively to diverse downst

tutorialsarxiv-cs-cv
13 Apr 2026
Agents

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding

DGX agent

arXiv:2604.09167v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance in multimodal understanding and reasoning, yet grounded reasoning in 3D scenes remains un

agentsarxiv-cs-cv
13 Apr 2026
Research

MASS: Mesh-inellipse Aligned Deformable Surfel Splatting for Hand Reconstruction and Rendering from Egocentric Monocular Video

DGX agent

arXiv:2604.08943v1 Announce Type: new Abstract: Reconstructing high-fidelity 3D hands from egocentric monocular videos remains a challenge due to the limitations in capturing high-resolution geometry,

researcharxiv-cs-cv
13 Apr 2026
Applications

Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory

DGX agent

arXiv:2604.08995v1 Announce Type: new Abstract: With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing

applicationsarxiv-cs-cv
13 Apr 2026
Research

Measurement-Consistent Langevin Corrector for Stabilizing Latent Diffusion Inverse Problem Solvers

DGX agent

arXiv:2601.04791v3 Announce Type: replace Abstract: While latent diffusion models (LDMs) have emerged as powerful priors for inverse problems, existing LDM-based solvers frequently suffer from instabi

researcharxiv-cs-cv
13 Apr 2026
Model Releases

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation

DGX agent

arXiv:2604.09088v1 Announce Type: new Abstract: Memory-efficient transfer learning (METL) approaches have recently achieved promising performance in adapting pre-trained models to downstream tasks. Th

model-releasesarxiv-cs-cv
13 Apr 2026
Safety

MeshOn: Intersection-Free Mesh-to-Mesh Composition

DGX agent

arXiv:2604.08799v1 Announce Type: cross Abstract: We propose MeshOn, a method that finds physically and semantically realistic compositions of two input meshes. Given an accessory, a base mesh with a

safetyarxiv-cs-cv
13 Apr 2026
Safety

MixFlow: Mixed Source Distributions Improve Rectified Flows

DGX agent

arXiv:2604.09181v1 Announce Type: new Abstract: Diffusion models and their variations, such as rectified flows, generate diverse and high-quality images, but they are still hindered by slow iterative

safetyarxiv-cs-cv
13 Apr 2026
Research

Multi-task Just Recognizable Difference for Video Coding for Machines: Database, Model, and Coding Application

DGX agent

arXiv:2604.09421v1 Announce Type: cross Abstract: Just Recognizable Difference (JRD) boosts coding efficiency for machine vision through visibility threshold modeling, but is currently limited to a si

researcharxiv-cs-cv
13 Apr 2026
Safety

Multimodal Anomaly Detection for Human-Robot Interaction

DGX agent

arXiv:2604.09326v1 Announce Type: cross Abstract: Ensuring safety and reliability in human-robot interaction (HRI) requires the timely detection of unexpected events that could lead to system failures

safetyarxiv-cs-cv
13 Apr 2026
Research

Music Audio-Visual Question Answering Requires Specialized Multimodal Designs

DGX agent

arXiv:2505.20638v2 Announce Type: replace-cross Abstract: While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music nec

researcharxiv-cs-cv
13 Apr 2026
Tutorials

MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation

DGX agent

arXiv:2604.08916v1 Announce Type: new Abstract: Conventional 3D instance segmentation methods rely on labor-intensive 3D annotations for supervised training, which limits their scalability and general

tutorialsarxiv-cs-cv
13 Apr 2026
Research

Nested Radially Monotone Polar Occupancy Estimation: Clinically-Grounded Optic Disc and Cup Segmentation for Glaucoma Screening

DGX agent

arXiv:2604.09062v1 Announce Type: new Abstract: Valid segmentation of the optic disc (OD) and optic cup (OC) from fundus photographs is essential for glaucoma screening. Unfortunately, existing deep l

researcharxiv-cs-cv
13 Apr 2026
Model Releases

NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: Multi-Exposure Image Fusion in Dynamic Scenes (Track 2)

DGX agent

arXiv:2604.09030v1 Announce Type: new Abstract: This paper presents NTIRE 2026, the 3rd Restore Any Image Model (RAIM) challenge on multi-exposure image fusion in dynamic scenes. We introduce a benchm

model-releasesarxiv-cs-cv
13 Apr 2026
Local Ai

Off-the-shelf Vision Models Benefit Image Manipulation Localization

DGX agent

arXiv:2604.09096v1 Announce Type: new Abstract: Image manipulation localization (IML) and general vision tasks are typically treated as two separate research directions due to the fundamental differen

local-aiarxiv-cs-cv
13 Apr 2026
Local Ai

Online3R: Online Learning for Consistent Sequential Reconstruction Based on Geometry Foundation Model

DGX agent

arXiv:2604.09480v1 Announce Type: new Abstract: We present Online3R, a new sequential reconstruction framework that is capable of adapting to new scenes through online learning, effectively resolving

local-aiarxiv-cs-cv
13 Apr 2026
Research

P3P Made Easy

DGX agent

arXiv:2508.01312v4 Announce Type: replace Abstract: We revisit the classical Perspective-Three-Point (P3P) problem, which aims to recover the absolute pose of a calibrated camera from three 2D-3D corr

researcharxiv-cs-cv
13 Apr 2026
Tutorials

Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch

DGX agent

arXiv:2604.09100v1 Announce Type: new Abstract: We propose a multimodal, physically grounded approach for metric-scale amodal object reconstruction and pose estimation under severe hand occlusion. Unl

tutorialsarxiv-cs-cv
13 Apr 2026
Research

PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation

DGX agent

arXiv:2508.05091v2 Announce Type: replace Abstract: Generating temporally coherent, long-duration videos with precise control over subject identity and movement remains a fundamental challenge for con

researcharxiv-cs-cv
13 Apr 2026
Safety

Post-Hoc Guidance for Consistency Models by Joint Flow Distribution Learning

DGX agent

arXiv:2604.08828v1 Announce Type: cross Abstract: Classifier-free Guidance (CFG) lets practitioners trade-off fidelity against diversity in Diffusion Models (DMs). The practicality of CFG is however h

safetyarxiv-cs-cv
13 Apr 2026
Research

PRADA: Probability-Ratio-Based Attribution and Detection of Autoregressive-Generated Images

DGX agent

arXiv:2511.20068v2 Announce Type: replace Abstract: Autoregressive (AR) image generation has recently emerged as a powerful paradigm for image synthesis. Leveraging the generation principle of large l

researcharxiv-cs-cv
13 Apr 2026
Model Releases

Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance

DGX agent

arXiv:2604.08881v1 Announce Type: new Abstract: In real-world deployments, Vision-Language Large Models (VLLMs) face critical challenges from multilingual and multimodal composite attacks: harmful ima

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search

DGX agent

arXiv:2604.08598v1 Announce Type: cross Abstract: Text-based person search faces inherent limitations due to data scarcity, driven by stringent privacy constraints and the high cost of manual annotati

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

R2G: A Multi-View Circuit Graph Benchmark Suite from RTL to GDSII

DGX agent

arXiv:2604.08810v1 Announce Type: new Abstract: Graph neural networks (GNNs) are increasingly applied to physical design tasks such as congestion prediction and wirelength estimation, yet progress is

model-releasesarxiv-cs-cv
13 Apr 2026
Applications

R3PM-Net: Real-time, Robust, Real-world Point Matching Network

DGX agent

arXiv:2604.05060v2 Announce Type: replace Abstract: Accurate Point Cloud Registration (PCR) is an important task in 3D data processing, involving the estimation of a rigid transformation between two p

applicationsarxiv-cs-cv
13 Apr 2026
Model Releases

RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models

DGX agent

arXiv:2511.19704v2 Announce Type: replace Abstract: Open-vocabulary semantic segmentation (OVSS) underpins many vision and robotics tasks that require generalizable semantic understanding. Existing ap

model-releasesarxiv-cs-cv
13 Apr 2026
Research

Ranked Activation Shift for Post-Hoc Out-of-Distribution Detection

DGX agent

arXiv:2604.08572v1 Announce Type: cross Abstract: State-of-the-art post-hoc out-of-distribution detection methods rely on intermediate layer activation editing. However, they exhibit inconsistent perf

researcharxiv-cs-cv
13 Apr 2026
Research

REACT3D: Recovering Articulations for Interactive Physical 3D Scenes

DGX agent

arXiv:2510.11340v4 Announce Type: replace Abstract: Interactive 3D scenes are increasingly vital for embodied intelligence, yet existing datasets remain limited due to the labor-intensive process of a

researcharxiv-cs-cv
13 Apr 2026
Applications

Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement

DGX agent

arXiv:2604.09473v1 Announce Type: new Abstract: Fully immersive experiences that tightly integrate 6-DoF visual and auditory interaction are essential for virtual and augmented reality. While such exp

applicationsarxiv-cs-cv
13 Apr 2026
Safety

Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing

DGX agent

arXiv:2604.09386v1 Announce Type: new Abstract: Instruction-guided image editing requires balancing target modification with non-target preservation. Recently, flow-based models have emerged as a stro

safetyarxiv-cs-cv
13 Apr 2026
Tutorials

RetinexDualV2: Physically-Grounded Dual Retinex for Generalized UHD Image Restoration

DGX agent

arXiv:2603.27979v2 Announce Type: replace Abstract: We propose RetinexDualV2, a unified, physically grounded dual-branch framework for diverse Ultra-High-Definition (UHD) image restoration. Unlike gen

tutorialsarxiv-cs-cv
13 Apr 2026
Model Releases

Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios

DGX agent

arXiv:2509.20006v3 Announce Type: replace Abstract: With the large models easing the labor-intensive manipulation process, image manipulations in today's real scenarios often entail a complex manipula

model-releasesarxiv-cs-cv
13 Apr 2026
Agents

RIRF: Reasoning Image Restoration Framework

DGX agent

arXiv:2604.09511v1 Announce Type: new Abstract: Universal image restoration (UIR) aims to recover clean images from diverse and unknown degradations using a unified model. Existing UIR methods primari

agentsarxiv-cs-cv
13 Apr 2026
Local Ai

Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors

DGX agent

arXiv:2604.09366v1 Announce Type: new Abstract: Reconstructing dynamic 4D scenes is an important yet challenging task. While 3D foundation models like VGGT excel in static settings, they often struggl

local-aiarxiv-cs-cv
13 Apr 2026
Agents

Robust by Design: A Continuous Monitoring and Data Integration Framework for Medical AI

DGX agent

arXiv:2604.09009v1 Announce Type: new Abstract: Adaptive medical AI models often face performance drops in dynamic clinical environments due to data drift. We propose an autonomous continuous monitori

agentsarxiv-cs-cv
13 Apr 2026
Applications

RS-OVC: Open-Vocabulary Counting for Remote-Sensing Data

DGX agent

arXiv:2604.08704v1 Announce Type: new Abstract: Object-Counting for remote-sensing (RS) imagery is attracting increasing research interest due to its crucial role in a wide and diverse set of applicat

applicationsarxiv-cs-cv
13 Apr 2026
← Previous
1…251252253254255…259
Next →