AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlog
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification

DGX agent

arXiv:2606.30376v1 Announce Type: cross Abstract: Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing

safetyarxiv-cs-cv
30 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

FR-DETR: Frequency and Recurrent Feature Refinement for Robust Object Detection under Adverse Weather

DGX agent

arXiv:2606.30471v1 Announce Type: new Abstract: Object detection under adverse weather remains challenging due to severe visual degradations and domain shifts. Existing enhancer-based approaches attem

researcharxiv-cs-cv
30 Jun 2026
Research

Frames2Residual: Spatiotemporal Decoupling for Self-Supervised Video Denoising

DGX agent

arXiv:2603.10417v2 Announce Type: replace Abstract: Self-supervised video denoising methods typically extend image-based frameworks into the temporal dimension, yet they often struggle to integrate in

researcharxiv-cs-cv
30 Jun 2026
Applications

FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution

DGX agent

arXiv:2606.28745v1 Announce Type: new Abstract: Diffusion prior-based methods have shown impressive results in real-world image super-resolution (ISR), yet two key challenges persist: balancing pixel-

applicationsarxiv-cs-cv
30 Jun 2026
Model Releases

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA

DGX agent

arXiv:2606.30220v1 Announce Type: new Abstract: High benchmark accuracy does not guarantee genuine use of visual evidence. We study this problem in traffic accident Video Question Answering (VideoQA),

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

From Fog Chamber to Aircraft Window: Pixel-Registered Imaging and Synthetic Fine-Tuning Enable Cross-Domain Defogging

DGX agent

arXiv:2606.29093v1 Announce Type: new Abstract: A deep defogging pipeline pretrained on controlled laboratory fog and fine-tuned with domain-randomized synthetic fog applied to clear outdoor scenes ge

model-releasesarxiv-cs-cv
30 Jun 2026
Research

From Local Windows to Adaptive Candidates via Individualized Exploratory: Rethinking Attention for Image Super-Resolution

DGX agent

arXiv:2601.08341v2 Announce Type: replace Abstract: Single Image Super-Resolution (SISR) is a fundamental computer vision task that aims to reconstruct a high-resolution (HR) image from a low-resoluti

researcharxiv-cs-cv
30 Jun 2026
Research

From Phase to Phenomenon: Self-Supervised Learning of Subsurface Scattering with Minimal Phase-shift Inputs

DGX agent

arXiv:2606.29461v1 Announce Type: new Abstract: We propose a self-supervised pretraining framework for learning sub-surface scattering (SSS) light transport representations from minimal input. Our met

researcharxiv-cs-cv
30 Jun 2026
Safety

GarmentZoom: Generating Zoomable Images from Garment Listings

DGX agent

arXiv:2606.29535v1 Announce Type: new Abstract: Online product listings for garments often include an overview photo and a close-up to show garment details. However, each photo focuses on either field

safetyarxiv-cs-cv
30 Jun 2026
Research

GCN-DevLSTM: Path Development for Skeleton-Based Action Recognition

DGX agent

arXiv:2403.15212v3 Announce Type: replace Abstract: Skeleton-based action recognition (SAR) in videos is an important but challenging task in computer vision. The recent state-of-the-art (SOTA) models

researcharxiv-cs-cv
30 Jun 2026
Tutorials

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

DGX agent

arXiv:2603.19235v2 Announce Type: replace Abstract: While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-

tutorialsarxiv-cs-cv
30 Jun 2026
Model Releases

GeoEdit: Geometry-Aware Object Editing via Dual-Branch Denoising

DGX agent

arXiv:2606.30003v1 Announce Type: new Abstract: Precisely manipulating objects in a single photograph (translation, rotation, scaling) while obeying 3D physical constraints remains unsolved for diffus

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

GeoISF: Instance Semantic Forest Inspired Large-Scale Cross-View Geo-Localization via Ground LiDAR-to-Satellite Image

DGX agent

arXiv:2606.28371v1 Announce Type: new Abstract: The problem of localization on a large-scale satellite image given a frame of query ground view point clouds remains challenging. Existing LiDAR-to-imag

safetyarxiv-cs-cv
30 Jun 2026
Model Releases

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing

DGX agent

arXiv:2606.30599v1 Announce Type: new Abstract: Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real

model-releasesarxiv-cs-cv
30 Jun 2026
Research

Good Enough? An Investigation on the Impact of Label Quality in Large-Scale Medical Datasets

DGX agent

arXiv:2505.20928v2 Announce Type: replace Abstract: Manually refining radiological segmentation masks is highly resource-intensive. To determine when this expert commitment is truly justified for the

researcharxiv-cs-cv
30 Jun 2026
Hardware

GPU-Accelerated Inverse Structural Anastylosis from Block Collapse Dynamics

DGX agent

arXiv:2606.28394v1 Announce Type: new Abstract: The physical anastylosis of collapsed architectural monuments -- the meticulous reassembly of fallen stone elements into their original structural confi

hardwarearxiv-cs-cv
30 Jun 2026
Research

Graph-GSReg: Leveraging 3D Scene Graphs for Gaussian Splatting Registration

DGX agent

arXiv:2606.29782v1 Announce Type: new Abstract: Merging multiple 3D Gaussian Splatting (3DGS) scenes into a single unified Gaussian representation is essential for large-scale 3D mapping and long-term

researcharxiv-cs-cv
30 Jun 2026
Research

Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video

DGX agent

arXiv:2606.28828v1 Announce Type: new Abstract: Learning a 4D scene representation from a single monocular video that supports dynamic novel-view synthesis while maintaining faithful geometry over tim

researcharxiv-cs-cv
30 Jun 2026
Research

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning

DGX agent

arXiv:2606.29915v1 Announce Type: new Abstract: Vision-Language Models (VLMs) often achieve high performance on benchmarks while remaining 'black boxes', yet they remain prone to hallucination or rely

researcharxiv-cs-cv
30 Jun 2026
Research

HASTE: A Framework for Training-Free, Dynamic, and Steerable Compression of Pre-Trained Convolutional Neural Networks

DGX agent

arXiv:2606.30516v1 Announce Type: new Abstract: Deploying large convolutional neural networks (CNNs) on resource-constrained devices is challenging due to their high computational cost. While dynamic

researcharxiv-cs-cv
30 Jun 2026
Research

HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers

DGX agent

arXiv:2603.12222v2 Announce Type: replace Abstract: Vision Transformers require significant computational resources and memory bandwidth, severely limiting their deployment on resource-constraint hard

researcharxiv-cs-cv
30 Jun 2026
Research

High-Resolution Flood Mapping With Sentinel-1 and Sentinel-2 via Misalignment-Robust Cross-Sensor Learning and Generative Despeckling

DGX agent

arXiv:2606.30511v1 Announce Type: new Abstract: Reliable high-resolution flood extent mapping from satellite imagery remains constrained by limited data fidelity and sensor-specific artifacts. Multisp

researcharxiv-cs-cv
30 Jun 2026
Research

HiReFF: High-Resolution Feedforward Human Reconstruction from Uncalibrated Sparse-View Video

DGX agent

arXiv:2606.29333v1 Announce Type: new Abstract: Uncalibrated volumetric video streaming for human reconstruction is essential for holographic communication and AR/VR, yet remains challenging due to th

researcharxiv-cs-cv
30 Jun 2026
Applications

HiRes: A Hierarchical Cascaded Method for Resistor Value Identification

DGX agent

arXiv:2606.30179v1 Announce Type: new Abstract: Accurate identification of resistor values from unconstrained images remains a challenging computer vision task due to variations in lighting, orientati

applicationsarxiv-cs-cv
30 Jun 2026
Local Ai

HKVLM: Faithful Reasoning Grounding by Binding Language Queries to a Frozen Detector

DGX agent

arXiv:2606.28862v1 Announce Type: new Abstract: Many visual requests -- ``the object to open this bottle'', ``the person not wearing a helmet'' -- require reasoning, not just category matching. Pure o

local-aiarxiv-cs-cv
30 Jun 2026
Research

HomeDiffusion: Zero-Shot Object Customization with Multi-View Representation Learning for Indoor Scenes

DGX agent

arXiv:2606.29828v1 Announce Type: new Abstract: Recently, zero-shot object customization generation methods have rapidly developed and shown tremendous potential for applications. For instance, in the

researcharxiv-cs-cv
30 Jun 2026
Research

HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers

DGX agent

arXiv:2606.29095v1 Announce Type: new Abstract: Diffusion-based video relighting enables controllable relighting from a single input video, but modern video diffusion backbones are trained on short cl

researcharxiv-cs-cv
30 Jun 2026
Research

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding

DGX agent

arXiv:2602.12957v3 Announce Type: replace Abstract: Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extractio

researcharxiv-cs-cv
30 Jun 2026
Research

HTC-SGA Former: A Hybrid Transformer-CNN Network with Self-Guided Attention and a New Boundary-Weighted Adaptive Loss for Coronary DSA Vessel Segmentation

DGX agent

arXiv:2606.29744v1 Announce Type: new Abstract: Accurate coronary Digital Subtraction Angiography (DSA) vessel segmentation is essential for computer-aided diagnosis and treatment planning of coronary

researcharxiv-cs-cv
30 Jun 2026
Applications

HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement in Game Engines

DGX agent

arXiv:2603.10604v3 Announce Type: replace Abstract: Generative models are increasingly used in video game engines to enhance the photorealism of rendered images for visual synthetic data generation an

applicationsarxiv-cs-cv
30 Jun 2026
Research

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation

DGX agent

arXiv:2606.30054v1 Announce Type: new Abstract: The advancement of generative AI models capable of producing text and image marks a critical step forward in the realm of multimodal intelligence, parti

researcharxiv-cs-cv
30 Jun 2026
Research

IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion

DGX agent

arXiv:2606.28604v1 Announce Type: new Abstract: Capturing full-body human motion with object interactions is crucial for AR/VR and robotics applications, yet it remains challenging for conventional vi

researcharxiv-cs-cv
30 Jun 2026
Tutorials

Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction

DGX agent

arXiv:2511.14540v2 Announce Type: replace Abstract: This paper focuses on a challenging setting of simultaneously modeling geometry and appearance of hand-object interaction scenes without any object

tutorialsarxiv-cs-cv
30 Jun 2026
Model Releases

InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing

DGX agent

arXiv:2603.13082v2 Announce Type: replace Abstract: Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limite

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment

DGX agent

arXiv:2606.30262v1 Announce Type: new Abstract: Text-to-image (T2I) diffusion models often fail to faithfully render explicit textual descriptions, instead defaulting to strongly learned visual priors

model-releasesarxiv-cs-cv
30 Jun 2026
Research

Invisible Yet Detected: PelFANet with Attention-Guided Anatomical Fusion for Pelvic Fracture Diagnosis

DGX agent

arXiv:2509.13873v3 Announce Type: replace Abstract: Pelvic fractures pose significant diagnostic challenges, particularly in cases where fracture signs are subtle or invisible on standard radiographs.

researcharxiv-cs-cv
30 Jun 2026
Research

IREU: Identity-Related Encoder-Only Unlearning for Customized Portrait Generation

DGX agent

arXiv:2606.29880v1 Announce Type: new Abstract: Customized Portrait Generation (CPG) technologies have been widely used to generate high-fidelity person images given an input image indicating the iden

researcharxiv-cs-cv
30 Jun 2026
Research

ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation

DGX agent

arXiv:2505.20935v3 Announce Type: replace Abstract: Recent open-weight text-to-image (T2I) diffusion models still struggle with multi-instance prompts, often omitting or merging instances and mixing s

researcharxiv-cs-cv
30 Jun 2026
Tutorials

JASPR: Joint Spatial Representation learning of histology and spatial genomics for improved virtual genomic screening and clinical prognostication

DGX agent

arXiv:2606.28395v1 Announce Type: new Abstract: Recent studies have shown that spatial properties of tumors are critical for understanding disease biology and predicting patient outcomes. These spatia

tutorialsarxiv-cs-cv
30 Jun 2026
Research

JOPP-3D: Joint Open Vocabulary Semantic Segmentation on Point Clouds and Panoramas

DGX agent

arXiv:2603.06168v3 Announce Type: replace Abstract: Semantic segmentation across visual modalities such as 3D point clouds and panoramic images remains a challenging task, primarily due to the scarcit

researcharxiv-cs-cv
30 Jun 2026
Tutorials

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization

DGX agent

arXiv:2606.28568v1 Announce Type: new Abstract: Speech-driven 3D facial animation methods face significant challenges in simultaneously achieving high-fidelity motion and precise artistic control at p

tutorialsarxiv-cs-cv
30 Jun 2026
Safety

L2D2-GS: Learning to Densify for Feedforward Dynamic Gaussian Scene Reconstruction

DGX agent

arXiv:2606.29374v1 Announce Type: new Abstract: High-fidelity reconstruction of dynamic urban environments is a cornerstone of autonomous driving simulation and large-scale world modeling. While 3D Ga

safetyarxiv-cs-cv
30 Jun 2026
Safety

LaGen: Towards Autoregressive LiDAR Scene Generation

DGX agent

arXiv:2511.21256v2 Announce Type: replace Abstract: Generative world models for autonomous driving (AD) are of great value in applications such as data augmentation, closed-loop simulation, and safety

safetyarxiv-cs-cv
30 Jun 2026
Research

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models

DGX agent

arXiv:2606.30168v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundan

researcharxiv-cs-cv
30 Jun 2026
Model Releases

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion

DGX agent

arXiv:2603.14526v2 Announce Type: replace Abstract: The recent success of inference-time scaling in large language models has inspired similar explorations in video diffusion. In particular, motivated

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

LaVPR: Benchmarking Language and Vision for Place Recognition

DGX agent

arXiv:2602.03253v2 Announce Type: replace Abstract: Visual Place Recognition (VPR) often fails under extreme environmental changes and perceptual aliasing. Beyond these limitations, standard systems c

model-releasesarxiv-cs-cv
30 Jun 2026
Research

LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection

DGX agent

arXiv:2512.05663v3 Announce Type: replace Abstract: Real-time monocular 3D object detection remains challenging due to severe depth ambiguity, viewpoint shifts, and the high computational cost of 3D r

researcharxiv-cs-cv
30 Jun 2026
Model Releases

Learning Cross-view Correspondences for Geo-localization on Planetary Surfaces

DGX agent

arXiv:2606.29821v1 Announce Type: new Abstract: Maintaining global position awareness is a fundamental challenge for planetary surface exploration, since satellite-based positioning systems are unavai

model-releasesarxiv-cs-cv
30 Jun 2026
← Previous
1…7980818283…263
Next →