AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
2 Jun 2026

COLLAR: Cascaded Object-Level Latent Refinement for High-Fidelity Conditional Generation

SafetyDGX agent

arXiv:2606.00954v1 Announce Type: new Abstract: Achieving high-fidelity object-level control in Diffusion Transformers remains a significant challenge despite the introduction of structural priors lik

Conditional Collapse in Sign Language Production: A Diagnostic and a Scaling Argument

ApplicationsDGX agent

arXiv:2606.01643v1 Announce Type: new Abstract: Sign Language Production (SLP) is the task of generating avatar sign language motion from natural language text. The quality of the generated motion is

Contrastive Augmented Transformer with Domain-specific Enhancement for Robust Multi-scenario Metal Surface Defect Detection

ApplicationsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.01962v1 Announce Type: new Abstract: Metal surface defect detection is critical for maintaining product quality in industrial manufacturing. However, it faces significant challenges, includ

Contrastive meta-domain adaptation for robust skin lesion classification across clinical and acquisition conditions

ResearchDGX agent

arXiv:2602.19857v2 Announce Type: replace Abstract: Deep learning models for dermatological image analysis remain sensitive to acquisition variability and domain-specific visual characteristics, leadi

CORE-MTL: Rethinking Gradient Balancing via Causal Orthogonal Representations

ResearchDGX agent

arXiv:2606.02221v1 Announce Type: new Abstract: Multi-task learning (MTL) aims to construct a joint model for multiple tasks by sharing a common representation across domains. To achieve this goal, ex

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection

Local AiDGX agent

arXiv:2606.01149v1 Announce Type: new Abstract: Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wis

Counterfactual Intervention Feature Transfer for Visible-Infrared Person Re-identification

ResearchDGX agent

arXiv:2208.00967v4 Announce Type: replace Abstract: Graph-based models have achieved great success in person re-identification tasks recently, which compute the graph topology structure (affinities) a

CountGD++: Generalized Prompting for Open-World Counting

AgentsDGX agent

arXiv:2512.23351v2 Announce Type: replace Abstract: The flexibility and accuracy of methods for automatically counting objects in images and videos are limited by the way the object can be specified.

CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval

SafetyDGX agent

arXiv:2606.00706v1 Announce Type: new Abstract: Cross-modal remote sensing image retrieval aims to retrieve semantically related scenes across heterogeneous sensing modalities. This remains challengin

Cross-Domain Dead Tree Detection via Knowledge Distillation in Aerial Imagery

SafetyDGX agent

arXiv:2606.02303v1 Announce Type: new Abstract: Detecting dead trees in aerial imagery is vital for assessing forest health, especially as tree mortality increases globally due to climate change, but

Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation

ResearchDGX agent

arXiv:2602.05217v2 Announce Type: replace Abstract: Cross-Domain Few-Shot Segmentation aims to segment categories in data-scarce domains conditioned on a few exemplars. Typical methods first establish

DeblurNVS: Geometric Latent Diffusion for Novel View Synthesis from Sparse Motion-Blurred Images

ResearchDGX agent

arXiv:2606.01315v1 Announce Type: new Abstract: Novel view synthesis (NVS) is a fundamental problem in computer vision and graphics. Recent advances in neural radiance fields (NeRF), 3D Gaussian Splat

Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation

ResearchDGX agent

arXiv:2606.01048v1 Announce Type: new Abstract: We propose Decoupled Residual Denoising Diffusion models (DRDD) for unified and data-efficient image-to-image (I2I) translation. While diffusion models

Deep Learning for Generating Computational PIN-4 Immunohistochemistry Staining from Prostate Biopsy H&E Images

ResearchDGX agent

arXiv:2606.01871v1 Announce Type: new Abstract: Immunohistochemistry (IHC)is frequently used to resolve diagnostically ambiguous prostate cancer biopsy findings on hematoxylin and eosin (H&E)-stained

Deep Learning for Remote Sensing to Improve Flood Inundation Mapping

TutorialsDGX agent

arXiv:2606.02310v1 Announce Type: new Abstract: Flooding is the most pervasive natural disaster worldwide. Timely and accurate flood inundation mapping are essential for informing disaster risk manage

DeepLatent: Think with Images via Parallel Latent Visual Reasoning

ResearchDGX agent

arXiv:2606.00562v1 Announce Type: new Abstract: The emerging paradigm of 'thinking with images' embeds visual states into intermediate reasoning steps, defining a new frontier for Vision-Language Mode

DefocusTrackerAI -- A Generalized Framework for the Automatic Detection of Defocused Particle Images

ResearchDGX agent

arXiv:2606.00076v1 Announce Type: new Abstract: The present work introduces DefocusTrackerAI, a generalized deep-learning framework for the automatic detection and position estimation of defocused par

Deformable Wiener Filter for Future Video Coding

Model ReleasesDGX agent

arXiv:2606.01576v1 Announce Type: new Abstract: In-loop filters have attracted increasing attention due to the remarkable noise-reduction capability in the hybrid video coding framework. However, the

Degradation-Aware Metric Prompting for Hyperspectral Image Restoration

ResearchDGX agent

arXiv:2512.20251v3 Announce Type: replace Abstract: Unified hyperspectral image (HSI) restoration aims to recover diverse degradations within a single model. However, current methods often rely on imp

DENSER: Depth-Guided Ensemble with Staged EFA-GS Reconstruction for Soccer Novel View Synthesis

ResearchDGX agent

arXiv:2606.01419v1 Announce Type: new Abstract: We propose DENSER, a Depth-guided ENSemble with Staged EFA-GS Reconstruction for soccer novel view synthesis. DENSER extends EFA-GS with three key contr

Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs

Model ReleasesDGX agent

arXiv:2606.01710v1 Announce Type: new Abstract: Vision-Language models (VLMs), such as CLIP, achieve powerful zero-shot classification. However, their predictions remain sensitive to spurious correlat

DerMAE: Improving skin lesion classification through conditioned latent diffusion and MAE distillation

Local AiDGX agent

arXiv:2602.19848v2 Announce Type: replace Abstract: Skin lesion classification datasets often suffer from severe class imbalance, with malignant cases significantly underrepresented, leading to biased

Detecting Pen-In-Air States from Video: A Proof-of-Concept Toward Complementary Handwriting Analysis

ResearchDGX agent

arXiv:2606.02342v1 Announce Type: new Abstract: Dynamic aspects of handwriting are critical for assessing developmental disorders such as dysgraphia and are typically captured using digitizing tablets

Diamonds in the Sky: Pareidolic Animals in Clouds

ResearchDGX agent

arXiv:2606.01361v1 Announce Type: new Abstract: People often see animal shapes in clouds, a phenomenon known as pareidolia. We propose an AI-based method that aims to predict which animals people are

Differing Roles of Leisure and Productivity in GDP - A Machine Learning based comparative analysis of Germany and USA

ResearchDGX agent

arXiv:2606.01234v1 Announce Type: cross Abstract: The GDP of a country is modelled as the relative interaction between two agents - working hours, reflecting the social choice of a population, and Tot

Diffusion Models for Hyperspectral Image Analysis: A Comprehensive Review

ResearchDGX agent

arXiv:2505.11158v4 Announce Type: replace-cross Abstract: Hyperspectral image (HSI) analysis plays a critical role in remote sensing, agriculture, and environmental monitoring. However, traditional me

DINO-GFSA: Geo-Localization via Semantic Gated Fusion and Mamba-based Sequential Aggregation

Model ReleasesDGX agent

arXiv:2606.00784v1 Announce Type: new Abstract: Cross-view geo-localization (CVGL) is critical for Unmanned Aerial Vehicle (UAV) self-positioning and target localization in GNSS-denied environments. H

Directed Distance Fields for Constant-Time Ray Queries on Gaussian Splatting

ResearchDGX agent

arXiv:2606.00817v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) renders new views of a scene in real time. Like every rasterizer, it answers only primary rays, the rays from the camera

Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

ResearchDGX agent

arXiv:2508.20072v4 Announce Type: replace Abstract: Vision-Language-Action (VLA) models adapt large vision-language backbones to map images and instructions into robot actions. However, prevailing VLA

Disentanglement-Based Equivariant Learning for Compositional VQA

Model ReleasesDGX agent

arXiv:2606.02168v1 Announce Type: new Abstract: Compositional visual question answering (VQA) represents a challenging yet fundamental task that requires models to comprehend novel combinations of pre

Distortion-Aware Fusion of Statistical and Vision-Language Features for Blind Image Quality Assessment

TutorialsDGX agent

arXiv:2606.02002v1 Announce Type: new Abstract: Blind image quality assessment (BIQA) aims to predict perceived image quality without access to a reference image. Classical natural scene statistics (N

Divide and Conquer: Reliable Multi-View Evidential Learning for Deepfake Detection

ResearchDGX agent

arXiv:2606.01885v1 Announce Type: new Abstract: With the evolution of generative models, deepfakes have achieved near-perfect semantic realism, leaving forensic traces only in subtle structural anomal

Domain Adaptation with a Single Vision-Language Embedding

AgentsDGX agent

arXiv:2410.21361v2 Announce Type: replace Abstract: Domain adaptation has been extensively investigated in computer vision but still requires access to target data at the training time, which might be

DPsurv: Dual-Prototype Evidential Fusion for Uncertainty-Aware and Interpretable Whole-Slide Image Survival Prediction

ResearchDGX agent

arXiv:2510.00053v2 Announce Type: replace-cross Abstract: Pathology whole-slide images (WSIs) are widely used for cancer survival analysis because of their comprehensive histopathological information

Drifting Preference Optimization for One-Step Generative Models

SafetyDGX agent

arXiv:2606.02521v1 Announce Type: cross Abstract: One-step text-to-image generators are attractive for deployment because they generate an image with a single forward pass, but preference finetuning t

Dual-Route Top-K Retrieval with 1v1 VLM Reranking for the CoVR-R

ResearchDGX agent

arXiv:2606.01097v1 Announce Type: new Abstract: We describe Dual-Route Top-K Retrieval with 1v1 VLM Reranking for the CoVR-R challenge. The method treats composed video retrieval as two coupled proble

Edge-directed geometric partitioning for versatile video coding

ResearchDGX agent

arXiv:2606.01641v1 Announce Type: new Abstract: To improve the coding performance, geometric partition (GEO) was proposed for the upcoming VVC standard. GEO provides 140 partition candidates. The inde

Edge Prediction for Roof Wireframe Reconstruction with Transformers

ResearchDGX agent

arXiv:2606.02406v1 Announce Type: new Abstract: This paper presents a competitive solution to the S23DR Challenge 2026, which aims to reconstruct 3D house roof wireframe models from sparse SfM point c

Effective Multi-sensor Conditioning for Street-view Novel-view Synthesis

ResearchDGX agent

arXiv:2606.01590v1 Announce Type: new Abstract: Modern vehicle platforms are equipped with a rich sensor suite, including LiDAR, calibrated multi-camera rigs, and accurate ego-motion, that in principl

Ego-METAS: Egocentric online Multimodal Energy-efficient Temporal Action Segmentation benchmark

Model ReleasesDGX agent

arXiv:2606.02246v1 Announce Type: new Abstract: To operate in the physical world, embodied agents must perceive their environment in an 'always-on' fashion, selectively accessing the most informative

EIVE: End-to-End Instance-Specific Visual Explanations for Detection Transformers

ResearchDGX agent

arXiv:2606.01601v1 Announce Type: new Abstract: Visual explainability for object detection remains challenging due to the multi-instance nature of detection. Existing approaches predominantly adopt po

Enhancing Blind Source Separation with Dissociative Principal Component Analysis

Model ReleasesDGX agent

arXiv:2411.12321v2 Announce Type: replace Abstract: Principal component analysis (PCA) and its sparse variants (sPCA) are widely used as a precursor to independent component analysis (ICA) for blind s

Entropy Minimization without Model Collapse: Mitigating Prediction Bias in Medical Imaging

SafetyDGX agent

arXiv:2606.02339v1 Announce Type: cross Abstract: Entropy minimization (EM) is the dominant objective for test-time adaptation, yet its failure mode, model collapse, remains poorly understood. In this

Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image Customization

SafetyDGX agent

arXiv:2606.02129v1 Announce Type: new Abstract: Image customization learns target subjects from reference concept images and generates conditioned images per text prompts, mainly modifying styles or b

ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs

ResearchDGX agent

arXiv:2606.00543v1 Announce Type: new Abstract: In Vision-Language Models (VLMs), high-resolution images produce a large number of visual tokens, resulting in high computational costs and KV-cache ove

Event-Based Vision in Space: Applications, Trends, and Future Directions

ResearchDGX agent

arXiv:2606.01280v1 Announce Type: new Abstract: Earth Observation (EO) is undergoing a significant transformation driven by the deployment of novel sensing technologies. Traditional frame-based optica

EvoCut: Multi-Layer Evolution-Aware Visual Token Compression for Efficient Large Vision-Language Models

ResearchDGX agent

arXiv:2606.01756v1 Announce Type: new Abstract: Large vision-language models (LVLMs) achieve strong performance on image and video understanding tasks, but their inference efficiency is constrained by

Evolving to the Aesthetics of a Vision-Language Model

ApplicationsDGX agent

arXiv:2606.00112v1 Announce Type: cross Abstract: Evolutionary systems have demonstrated remarkable results in creative domains, with recent applications in generative typography, design, and music. H

Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene Graphs

SafetyDGX agent

arXiv:2606.01072v1 Announce Type: cross Abstract: Imitation learning enables robots to learn how to execute tasks via observation. However, real-world environments like homes and offices are often sev

Explainable Forensics of Manipulated Segments in Untrimmed Long Videos

Model ReleasesDGX agent

arXiv:2606.02402v1 Announce Type: new Abstract: The rapid advancement of AI-driven video generation has transformed content creation, while simultaneously increasing the risk of misinformation through

Exploiting In-Sensor Computing for Energy-Efficient Earth Observation

ResearchDGX agent

arXiv:2606.01271v1 Announce Type: new Abstract: The rapid growth of the satellite industry has driven a significant increase in geospatial data acquisition, highlighting a critical bottleneck: the sev

Exploiting Semantic and Pixel Representations for Ultra-Low Bitrate Image Compression

Model ReleasesDGX agent

arXiv:2606.01608v1 Announce Type: new Abstract: Most existing extreme compression methods fail to achieve an optimal rate-distortion-perception trade-off, as they typically prioritize perceptual fidel

Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays

Model ReleasesDGX agent

arXiv:2509.15234v2 Announce Type: replace Abstract: Multimodal learning from paired medical images and clinical text is a central challenge in medical data-driven informatics, where effective cross-mo

ext{VG}^2GT: Voxel-Gaussian Splatting Visual Geometry Grounded Transformer

ResearchDGX agent

arXiv:2606.01573v1 Announce Type: new Abstract: Gaussian splatting has shown strong potential for 3D reconstruction and novel view synthesis. However, most existing methods require accurate camera par

FACT: A Simple and Efficient Framework for Active Finetuning

Model ReleasesDGX agent

arXiv:2606.02079v1 Announce Type: new Abstract: The main goal of active finetuning is to improve a pretrained model's performance on a specific task or domain by finetuning it with carefully selected

Family Matters: A Systematic Study of Spatial vs. Frequency Masking for Continual Test-Time Adaptation

SafetyDGX agent

arXiv:2512.08048v3 Announce Type: replace Abstract: Recent continual test-time adaptation (CTTA) methods adopt masked image modeling to stabilize learning under distribution shift, yet each treats its

Fast-SAM3D: 3Dfy Anything in Images but Faster

Model ReleasesDGX agent

arXiv:2602.05293v2 Announce Type: replace Abstract: SAM3D enables scalable, open-world 3D reconstruction from complex scenes, yet its deployment is hindered by prohibitive inference latency. In this w

FDIO: Frequency Decomposed Inertial Odometry

Local AiDGX agent

arXiv:2511.15645v3 Announce Type: replace Abstract: Pedestrian inertial odometry (PIO) estimates autonomous pedestrian motion using only acceleration and angular velocity measurements collected by an

Feature Alignment Determines Fusion Strategy: A Comparative Study of Cross-Attention and Concatenation in Multimodal Learning

SafetyDGX agent

arXiv:2606.01207v1 Announce Type: new Abstract: The choice between cross-attention and concatenation for multimodal fusion remains governed by practitioner intuition rather than principled understandi

FiSeR: Fine-Grained Source Representations for Cross-Domain AI Image Detection

TutorialsDGX agent

arXiv:2606.00606v1 Announce Type: new Abstract: Real-world synthetic image detectors often generalize poorly under domain shift despite strong in-domain performance. Using unsupervised UMAP projection

← Previous
1…100101102103104…211
Next →