AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
30 Jun 2026

Envisage: Diffusion-Based Rhinoplasty Goal Visualization with Mask-Decomposed Evaluation

Model ReleasesDGX agent

arXiv:2606.28628v1 Announce Type: cross Abstract: Localized generative editing needs localized evaluation: full-image identity metrics are structurally confounded under hard-composited edits. We prese

EpiSAM: Character Segmentation in Challenging Stone Inscriptions

ResearchDGX agent

arXiv:2606.28859v1 Announce Type: new Abstract: Stone inscriptions are invaluable sources of historical and linguistic knowledge, yet their automated analysis remains a major challenge due to surface

EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2512.21545v2 Announce Type: replace Abstract: Object removal must prevent the masked target from reappearing and reconstruct the occluded background with structural and contextual fidelity, rath

Establishing the Minimal Clinically Important Difference (MCID) for Smartphone-Derived Gait Measures in Multiple Sclerosis

ApplicationsDGX agent

arXiv:2606.28449v1 Announce Type: cross Abstract: Background: Digital health technologies allow for frequent, remote gait monitoring in people with multiple sclerosis (MS). However, to differentiate d

Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems

AgentsDGX agent

arXiv:2504.02918v3 Announce Type: replace Abstract: Recent advances in image and video generation raise hopes that these models possess world modeling capabilities-the ability to generate realistic, p

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model

ApplicationsDGX agent

arXiv:2606.29384v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable

EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

Model ReleasesDGX agent

arXiv:2606.20092v2 Announce Type: replace Abstract: Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-r

EvLIR: Learning Illumination Residuals from Ordered Events for Low-Light Image Enhancement

ResearchDGX agent

arXiv:2606.29430v1 Announce Type: new Abstract: Low-light image enhancement is severely ill-posed when the input frame contains missing structure, saturated noise, and weak local contrast. Event camer

Evolutionary Hyperparameter Optimization to Find Lightweight CNN Models for Autonomous Steering

AgentsDGX agent

arXiv:2606.29684v1 Announce Type: cross Abstract: This research investigates the optimization of Convolutional and Dense Neural Networks (CNNs and DNNs) for autonomous steering using the (N+M) Evoluti

ExACT: Exemplar-Driven Calibrated Refinement for Training-Free Visual Grounding in Remote Sensing Images

Local AiDGX agent

arXiv:2606.28920v1 Announce Type: new Abstract: Remote sensing visual grounding (RSVG) aims to locate specific objects in high-resolution RS imagery using free-form natural language descriptions. Whil

Explainability-Aware Frustum Attack: Exposing Structural Vulnerabilities in LiDAR-Based 3D Object Detectors

Model ReleasesDGX agent

arXiv:2606.29963v1 Announce Type: new Abstract: The structural vulnerabilities of point cloud-based 3D object detectors remain poorly understood. Prior work has studied adversarial robustness primaril

ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2604.02714v2 Announce Type: replace Abstract: End-to-end autonomous driving models based on Vision-Language-Action (VLA) architectures have shown promising results by learning driving policies t

FAIL: Flow Matching Adversarial Imitation Learning for Image Generation

SafetyDGX agent

arXiv:2602.12155v2 Announce Type: replace Abstract: Post-training of flow matching models-aligning the output distribution with a high-quality target-is mathematically equivalent to imitation learning

Fast Equivariant Imaging: Accelerating Unsupervised Learning and Model Adaptation via Inexact Splitting

ResearchDGX agent

arXiv:2507.06764v5 Announce Type: replace-cross Abstract: In this work, we propose Fast Equivariant Imaging (FEI), a novel unsupervised learning framework to rapidly and efficiently train deep imaging

FastPano3D: Feed-Forward Indoor Panoramic 3D Reconstruction from a Single Image

Model ReleasesDGX agent

arXiv:2606.30352v1 Announce Type: new Abstract: Recent advances in 3D scene reconstruction have highlighted the intricate trade-offs among rendering quality, inference efficiency, and data dependency.

FDM-MFVT: Few-step Sampling Diffusion Model for Mask-Free Virtual Try-On

SafetyDGX agent

arXiv:2606.29319v1 Announce Type: new Abstract: Image-based Virtual Try-On (IVTON) has greatly advanced through diffusion models, yet existing methods require many sampling steps and depend on masks w

FiRe: Frequency Reparameterization as a Preconditioner for Periodic Implicit Neural Representations

Model ReleasesDGX agent

arXiv:2606.29414v1 Announce Type: new Abstract: Periodic Implicit Neural Representations (INRs) such as SIREN and FINER assign every neuron, the same global frequency, spending the representational bu

FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification

SafetyDGX agent

arXiv:2606.30376v1 Announce Type: cross Abstract: Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing

FR-DETR: Frequency and Recurrent Feature Refinement for Robust Object Detection under Adverse Weather

ResearchDGX agent

arXiv:2606.30471v1 Announce Type: new Abstract: Object detection under adverse weather remains challenging due to severe visual degradations and domain shifts. Existing enhancer-based approaches attem

Frames2Residual: Spatiotemporal Decoupling for Self-Supervised Video Denoising

ResearchDGX agent

arXiv:2603.10417v2 Announce Type: replace Abstract: Self-supervised video denoising methods typically extend image-based frameworks into the temporal dimension, yet they often struggle to integrate in

FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution

ApplicationsDGX agent

arXiv:2606.28745v1 Announce Type: new Abstract: Diffusion prior-based methods have shown impressive results in real-world image super-resolution (ISR), yet two key challenges persist: balancing pixel-

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA

Model ReleasesDGX agent

arXiv:2606.30220v1 Announce Type: new Abstract: High benchmark accuracy does not guarantee genuine use of visual evidence. We study this problem in traffic accident Video Question Answering (VideoQA),

From Fog Chamber to Aircraft Window: Pixel-Registered Imaging and Synthetic Fine-Tuning Enable Cross-Domain Defogging

Model ReleasesDGX agent

arXiv:2606.29093v1 Announce Type: new Abstract: A deep defogging pipeline pretrained on controlled laboratory fog and fine-tuned with domain-randomized synthetic fog applied to clear outdoor scenes ge

From Local Windows to Adaptive Candidates via Individualized Exploratory: Rethinking Attention for Image Super-Resolution

ResearchDGX agent

arXiv:2601.08341v2 Announce Type: replace Abstract: Single Image Super-Resolution (SISR) is a fundamental computer vision task that aims to reconstruct a high-resolution (HR) image from a low-resoluti

From Phase to Phenomenon: Self-Supervised Learning of Subsurface Scattering with Minimal Phase-shift Inputs

ResearchDGX agent

arXiv:2606.29461v1 Announce Type: new Abstract: We propose a self-supervised pretraining framework for learning sub-surface scattering (SSS) light transport representations from minimal input. Our met

GarmentZoom: Generating Zoomable Images from Garment Listings

SafetyDGX agent

arXiv:2606.29535v1 Announce Type: new Abstract: Online product listings for garments often include an overview photo and a close-up to show garment details. However, each photo focuses on either field

GCN-DevLSTM: Path Development for Skeleton-Based Action Recognition

ResearchDGX agent

arXiv:2403.15212v3 Announce Type: replace Abstract: Skeleton-based action recognition (SAR) in videos is an important but challenging task in computer vision. The recent state-of-the-art (SOTA) models

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

TutorialsDGX agent

arXiv:2603.19235v2 Announce Type: replace Abstract: While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-

GeoEdit: Geometry-Aware Object Editing via Dual-Branch Denoising

Model ReleasesDGX agent

arXiv:2606.30003v1 Announce Type: new Abstract: Precisely manipulating objects in a single photograph (translation, rotation, scaling) while obeying 3D physical constraints remains unsolved for diffus

GeoISF: Instance Semantic Forest Inspired Large-Scale Cross-View Geo-Localization via Ground LiDAR-to-Satellite Image

SafetyDGX agent

arXiv:2606.28371v1 Announce Type: new Abstract: The problem of localization on a large-scale satellite image given a frame of query ground view point clouds remains challenging. Existing LiDAR-to-imag

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing

Model ReleasesDGX agent

arXiv:2606.30599v1 Announce Type: new Abstract: Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real

Good Enough? An Investigation on the Impact of Label Quality in Large-Scale Medical Datasets

ResearchDGX agent

arXiv:2505.20928v2 Announce Type: replace Abstract: Manually refining radiological segmentation masks is highly resource-intensive. To determine when this expert commitment is truly justified for the

GPU-Accelerated Inverse Structural Anastylosis from Block Collapse Dynamics

HardwareDGX agent

arXiv:2606.28394v1 Announce Type: new Abstract: The physical anastylosis of collapsed architectural monuments -- the meticulous reassembly of fallen stone elements into their original structural confi

Graph-GSReg: Leveraging 3D Scene Graphs for Gaussian Splatting Registration

ResearchDGX agent

arXiv:2606.29782v1 Announce Type: new Abstract: Merging multiple 3D Gaussian Splatting (3DGS) scenes into a single unified Gaussian representation is essential for large-scale 3D mapping and long-term

Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video

ResearchDGX agent

arXiv:2606.28828v1 Announce Type: new Abstract: Learning a 4D scene representation from a single monocular video that supports dynamic novel-view synthesis while maintaining faithful geometry over tim

H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning

ResearchDGX agent

arXiv:2606.29915v1 Announce Type: new Abstract: Vision-Language Models (VLMs) often achieve high performance on benchmarks while remaining 'black boxes', yet they remain prone to hallucination or rely

HASTE: A Framework for Training-Free, Dynamic, and Steerable Compression of Pre-Trained Convolutional Neural Networks

ResearchDGX agent

arXiv:2606.30516v1 Announce Type: new Abstract: Deploying large convolutional neural networks (CNNs) on resource-constrained devices is challenging due to their high computational cost. While dynamic

HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers

ResearchDGX agent

arXiv:2603.12222v2 Announce Type: replace Abstract: Vision Transformers require significant computational resources and memory bandwidth, severely limiting their deployment on resource-constraint hard

High-Resolution Flood Mapping With Sentinel-1 and Sentinel-2 via Misalignment-Robust Cross-Sensor Learning and Generative Despeckling

ResearchDGX agent

arXiv:2606.30511v1 Announce Type: new Abstract: Reliable high-resolution flood extent mapping from satellite imagery remains constrained by limited data fidelity and sensor-specific artifacts. Multisp

HiReFF: High-Resolution Feedforward Human Reconstruction from Uncalibrated Sparse-View Video

ResearchDGX agent

arXiv:2606.29333v1 Announce Type: new Abstract: Uncalibrated volumetric video streaming for human reconstruction is essential for holographic communication and AR/VR, yet remains challenging due to th

HiRes: A Hierarchical Cascaded Method for Resistor Value Identification

ApplicationsDGX agent

arXiv:2606.30179v1 Announce Type: new Abstract: Accurate identification of resistor values from unconstrained images remains a challenging computer vision task due to variations in lighting, orientati

HKVLM: Faithful Reasoning Grounding by Binding Language Queries to a Frozen Detector

Local AiDGX agent

arXiv:2606.28862v1 Announce Type: new Abstract: Many visual requests -- ``the object to open this bottle'', ``the person not wearing a helmet'' -- require reasoning, not just category matching. Pure o

HomeDiffusion: Zero-Shot Object Customization with Multi-View Representation Learning for Indoor Scenes

ResearchDGX agent

arXiv:2606.29828v1 Announce Type: new Abstract: Recently, zero-shot object customization generation methods have rapidly developed and shown tremendous potential for applications. For instance, in the

HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers

ResearchDGX agent

arXiv:2606.29095v1 Announce Type: new Abstract: Diffusion-based video relighting enables controllable relighting from a single input video, but modern video diffusion backbones are trained on short cl

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding

ResearchDGX agent

arXiv:2602.12957v3 Announce Type: replace Abstract: Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extractio

HTC-SGA Former: A Hybrid Transformer-CNN Network with Self-Guided Attention and a New Boundary-Weighted Adaptive Loss for Coronary DSA Vessel Segmentation

ResearchDGX agent

arXiv:2606.29744v1 Announce Type: new Abstract: Accurate coronary Digital Subtraction Angiography (DSA) vessel segmentation is essential for computer-aided diagnosis and treatment planning of coronary

HyPER-GAN: Hybrid Patch-Based Image-to-Image Translation for Real-Time Photorealism Enhancement in Game Engines

ApplicationsDGX agent

arXiv:2603.10604v3 Announce Type: replace Abstract: Generative models are increasingly used in video game engines to enhance the photorealism of rendered images for visual synthetic data generation an

Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation

ResearchDGX agent

arXiv:2606.30054v1 Announce Type: new Abstract: The advancement of generative AI models capable of producing text and image marks a critical step forward in the realm of multimodal intelligence, parti

IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion

ResearchDGX agent

arXiv:2606.28604v1 Announce Type: new Abstract: Capturing full-body human motion with object interactions is crucial for AR/VR and robotics applications, yet it remains challenging for conventional vi

Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction

TutorialsDGX agent

arXiv:2511.14540v2 Announce Type: replace Abstract: This paper focuses on a challenging setting of simultaneously modeling geometry and appearance of hand-object interaction scenes without any object

InterEdit: Navigating Text-Guided 3D Dyadic Human Motion Editing

Model ReleasesDGX agent

arXiv:2603.13082v2 Announce Type: replace Abstract: Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limite

Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment

Model ReleasesDGX agent

arXiv:2606.30262v1 Announce Type: new Abstract: Text-to-image (T2I) diffusion models often fail to faithfully render explicit textual descriptions, instead defaulting to strongly learned visual priors

Invisible Yet Detected: PelFANet with Attention-Guided Anatomical Fusion for Pelvic Fracture Diagnosis

ResearchDGX agent

arXiv:2509.13873v3 Announce Type: replace Abstract: Pelvic fractures pose significant diagnostic challenges, particularly in cases where fracture signs are subtle or invisible on standard radiographs.

IREU: Identity-Related Encoder-Only Unlearning for Customized Portrait Generation

ResearchDGX agent

arXiv:2606.29880v1 Announce Type: new Abstract: Customized Portrait Generation (CPG) technologies have been widely used to generate high-fidelity person images given an input image indicating the iden

ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation

ResearchDGX agent

arXiv:2505.20935v3 Announce Type: replace Abstract: Recent open-weight text-to-image (T2I) diffusion models still struggle with multi-instance prompts, often omitting or merging instances and mixing s

JASPR: Joint Spatial Representation learning of histology and spatial genomics for improved virtual genomic screening and clinical prognostication

TutorialsDGX agent

arXiv:2606.28395v1 Announce Type: new Abstract: Recent studies have shown that spatial properties of tumors are critical for understanding disease biology and predicting patient outcomes. These spatia

JOPP-3D: Joint Open Vocabulary Semantic Segmentation on Point Clouds and Panoramas

ResearchDGX agent

arXiv:2603.06168v3 Announce Type: replace Abstract: Semantic segmentation across visual modalities such as 3D point clouds and panoramic images remains a challenging task, primarily due to the scarcit

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization

TutorialsDGX agent

arXiv:2606.28568v1 Announce Type: new Abstract: Speech-driven 3D facial animation methods face significant challenges in simultaneously achieving high-fidelity motion and precise artistic control at p

L2D2-GS: Learning to Densify for Feedforward Dynamic Gaussian Scene Reconstruction

SafetyDGX agent

arXiv:2606.29374v1 Announce Type: new Abstract: High-fidelity reconstruction of dynamic urban environments is a cornerstone of autonomous driving simulation and large-scale world modeling. While 3D Ga

LaGen: Towards Autoregressive LiDAR Scene Generation

SafetyDGX agent

arXiv:2511.21256v2 Announce Type: replace Abstract: Generative world models for autonomous driving (AD) are of great value in applications such as data augmentation, closed-loop simulation, and safety

← Previous
1…6162636465…209
Next →