AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
13 Apr 2026

Cross-Modal Knowledge Distillation from Spatial Transcriptomics to Histology

ResearchDGX agent

arXiv:2604.09076v1 Announce Type: new Abstract: Spatial transcriptomics provides a molecularly rich description of tissue organization, enabling unsupervised discovery of tissue niches -- spatially co

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation

TutorialsDGX agent

arXiv:2604.09201v1 Announce Type: new Abstract: Camera-controllable video generation aims to synthesize videos with flexible and physically plausible camera movements. However, existing methods either

Customized Fusion: A Closed-Loop Dynamic Network for Adaptive Multi-Task-Aware Infrared-Visible Image Fusion

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.08924v1 Announce Type: new Abstract: Infrared-visible image fusion aims to integrate complementary information for robust visual understanding, but existing fusion methods struggle with sim

Deep Light Pollution Removal in Night Cityscape Photographs

ResearchDGX agent

arXiv:2604.09145v1 Announce Type: new Abstract: Nighttime photography is severely degraded by light pollution induced by pervasive artificial lighting in urban environments. After long-range scatterin

DeFakeQ: Enabling Real-Time Deepfake Detection on Edge Devices via Adaptive Bidirectional Quantization

Model ReleasesDGX agent

arXiv:2604.08847v1 Announce Type: new Abstract: Deepfake detection has become a fundamental component of modern media forensics. Despite significant progress in detection accuracy, most existing metho

Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenarios

TutorialsDGX agent

arXiv:2604.08922v1 Announce Type: new Abstract: Complex degradations like noise, blur, and low resolution are typical challenges in real world image fusion tasks, limiting the performance and practica

Detecting Diffusion-generated Images via Dynamic Assembly ForestsDetecting Diffusion-generated Images via Dynamic Assembly Forests

ResearchDGX agent

arXiv:2604.09106v1 Announce Type: new Abstract: Diffusion models are known for generating high-quality images, causing serious security concerns. To combat this, most efforts rely on deep neural netwo

Do Vision Language Models Need to Process Image Tokens?

ResearchDGX agent

arXiv:2604.09425v1 Announce Type: new Abstract: Vision Language Models (VLMs) have achieved remarkable success by integrating visual encoders with large language models (LLMs). While VLMs process dens

Domain-generalizable Face Anti-Spoofing with Patch-based Multi-tasking and Artifact Pattern Conversion

ResearchDGX agent

arXiv:2604.09018v1 Announce Type: new Abstract: Face Anti-Spoofing (FAS) algorithms, designed to secure face recognition systems against spoofing, struggle with limited dataset diversity, impairing th

DSVTLA: Deep Swin Vision Transformer-Based Transfer Learning Architecture for Multi-Type Cancer Histopathological Cancer Image Classification

Model ReleasesDGX agent

arXiv:2604.09468v1 Announce Type: cross Abstract: In this study, we proposed a deep Swin-Vision Transformer-based transfer learning architecture for robust multi-cancer histopathological image classif

Dynamic Class-Aware Active Learning for Unbiased Satellite Image Segmentation

SafetyDGX agent

arXiv:2604.08965v1 Announce Type: new Abstract: Semantic segmentation of satellite imagery plays a vital role in land cover mapping and environmental monitoring. However, annotating large-scale, high-

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection

ApplicationsDGX agent

arXiv:2604.09164v1 Announce Type: new Abstract: Temporal human action detection aims to identify and localize action segments within untrimmed videos, serving as a pivotal task in video understanding.

Efficient Unlearning through Maximizing Relearning Convergence Delay

ResearchDGX agent

arXiv:2604.09391v1 Announce Type: cross Abstract: Machine unlearning poses challenges in removing mislabeled, contaminated, or problematic data from a pretrained model. Current unlearning approaches a

EfficientSign: An Attention-Enhanced Lightweight Architecture for Indian Sign Language Recognition

ResearchDGX agent

arXiv:2604.08694v1 Announce Type: new Abstract: How do you build a sign language recognizer that works on a phone? That question drove this work. We built EfficientSign, a lightweight model which take

EGLOCE: Training-Free Energy-Guided Latent Optimization for Concept Erasure

SafetyDGX agent

arXiv:2604.09405v1 Announce Type: new Abstract: As text-to-image diffusion models grow increasingly prevalent, the ability to remove specific concepts-mostly explicit content and many copyrighted char

EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks

Model ReleasesDGX agent

arXiv:2604.09535v1 Announce Type: new Abstract: Large foundation models have made significant advances in embodied intelligence, enabling synthesis and reasoning over egocentric input for household ta

ELT: Elastic Looped Transformers for Visual Generation

Model ReleasesDGX agent

arXiv:2604.09168v1 Announce Type: new Abstract: We introduce Elastic Looped Transformers (ELT), a highly parameter-efficient class of visual generative models based on a recurrent transformer architec

EmoCtrl: Controllable Emotional Image Content Generation

SafetyDGX agent

arXiv:2512.22437v2 Announce Type: replace Abstract: An image conveys meaning through both its visual content and emotional tone, jointly shaping human perception. We introduce Controllable Emotional I

Enhanced Self-Supervised Multi-Image Super-Resolution for Camera Array Images

ApplicationsDGX agent

arXiv:2604.06816v2 Announce Type: replace-cross Abstract: Conventional multi-image super-resolution (MISR) methods, such as burst and video SR, rely on sequential frames from a single camera. Conseque

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration

AgentsDGX agent

arXiv:2604.09367v1 Announce Type: new Abstract: Ancient inscriptions, as repositories of cultural memory, have suffered from centuries of environmental and human-induced degradation. Restoring their i

FaceLiVTv2: An Improved Hybrid Architecture for Efficient Mobile Face Recognition

Local AiDGX agent

arXiv:2604.09127v1 Announce Type: new Abstract: Lightweight face recognition is increasingly important for deployment on edge and mobile devices, where strict constraints on latency, memory, and energ

FashionStylist: An Expert Knowledge-enhanced Multimodal Dataset for Fashion Understanding

Model ReleasesDGX agent

arXiv:2604.09249v1 Announce Type: new Abstract: Fashion understanding requires both visual perception and expert-level reasoning about style, occasion, compatibility, and outfit rationale. However, ex

Fast Model-guided Instance-wise Adaptation Framework for Real-world Pansharpening with Fidelity Constraints

HardwareDGX agent

arXiv:2604.08903v1 Announce Type: new Abstract: Pansharpening aims to generate high-resolution multispectral (HRMS) images by fusing low-resolution multispectral (LRMS) and high-resolution panchromati

FDIF: Formula-Driven supervised Learning with Implicit Functions for 3D Medical Image Segmentation

ResearchDGX agent

arXiv:2603.23199v2 Announce Type: replace Abstract: Deep learning-based 3D medical image segmentation methods relies on large-scale labeled datasets, yet acquiring such data is difficult due to privac

Few-Shot Personalized Age Estimation

Model ReleasesDGX agent

arXiv:2604.09125v1 Announce Type: new Abstract: Existing age estimation methods treat each face as an independent sample, learning a global mapping from appearance to age. This ignores a well-document

Fine-Grained Action Segmentation for Renorrhaphy in Robot-Assisted Partial Nephrectomy

Model ReleasesDGX agent

arXiv:2604.09051v1 Announce Type: new Abstract: Fine-grained action segmentation during renorrhaphy in robot-assisted partial nephrectomy requires frame-level recognition of visually similar suturing

FIRE-CIR: Fine-grained Reasoning for Composed Fashion Image Retrieval

Model ReleasesDGX agent

arXiv:2604.09114v1 Announce Type: new Abstract: Composed image retrieval (CIR) aims to retrieve a target image that depicts a reference image modified by a textual description. While recent vision-lan

FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs

HardwareDGX agent

arXiv:2512.20033v2 Announce Type: replace Abstract: We present FlashLips, a two-stage, mask-free lip-sync system that decouples lips control from rendering and achieves real-time performance, with our

From Frames to Events: Rethinking Evaluation in Human-Centric Video Anomaly Detection

Local AiDGX agent

arXiv:2604.09327v1 Announce Type: new Abstract: Pose-based Video Anomaly Detection (VAD) has gained significant attention for its privacy-preserving nature and robustness to environmental variations.

Generative View Stitching

TutorialsDGX agent

arXiv:2510.24718v3 Announce Type: replace Abstract: Autoregressive video diffusion models are capable of long rollouts that are stable and consistent with history, but they are unable to guide the cur

Geometry Reinforced Efficient Attention Tuning Equipped with Normals for Robust Stereo Matching

ResearchDGX agent

arXiv:2604.09142v1 Announce Type: new Abstract: Despite remarkable advances in image-driven stereo matching over the past decade, Synthetic-to-Realistic Zero-Shot (Syn-to-Real) generalization remains

GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing

Model ReleasesDGX agent

arXiv:2604.08896v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) have accelerated progress in domain-oriented AI, yet their development in geoscience and rem

GeRM: A Generative Rendering Model From Physically Realistic to Photorealistic

AgentsDGX agent

arXiv:2604.09304v1 Announce Type: new Abstract: For decades, Physically-Based Rendering (PBR) is the fundation of synthesizing photorealisitic images, and therefore sometimes roughly referred as Photo

Globally Optimal Pose from Orthographic Silhouettes

ResearchDGX agent

arXiv:2604.09199v1 Announce Type: new Abstract: We solve the problem of determining the pose of known shapes in R^3 from their unoccluded silhouettes. The pose is determined up to global optimality us

HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models

ResearchDGX agent

arXiv:2604.06165v2 Announce Type: replace Abstract: Large vision-language models can produce object hallucinations in image descriptions, highlighting the need for effective detection and mitigation s

Harnessing Weak Pair Uncertainty for Text-based Person Search

ResearchDGX agent

arXiv:2604.08877v1 Announce Type: new Abstract: In this paper, we study the text-based person search, which is to retrieve the person of interest via natural language description. Prevailing methods u

HD-VGGT: High-Resolution Visual Geometry Transformer

ResearchDGX agent

arXiv:2603.27222v2 Announce Type: replace Abstract: High-resolution imagery is essential for accurate 3D reconstruction, as many geometric details only emerge at fine spatial scales. Recent feed-forwa

Hitem3D 2.0: Multi-View Guided Native 3D Texture Generation

SafetyDGX agent

arXiv:2604.09231v1 Announce Type: new Abstract: Although recent advances have improved the quality of 3D texture generation, existing methods still struggle with incomplete texture coverage, cross-vie

How Noise Benefits AI-generated Image Detection

ResearchDGX agent

arXiv:2511.16136v2 Announce Type: replace Abstract: The rapid advancement of generative models has made real and synthetic images increasingly indistinguishable. Although extensive efforts have been d

How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms

Model ReleasesDGX agent

arXiv:2604.08966v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have advanced Video Temporal Grounding (VTG), existing methods often couple output paradigms with differe

Incremental Semantics-Aided Meshing from LiDAR-Inertial Odometry and RGB Direct Label Transfer

ResearchDGX agent

arXiv:2604.09478v1 Announce Type: new Abstract: Geometric high-fidelity mesh reconstruction from LiDAR-inertial scans remains challenging in large, complex indoor environments -- such as cultural buil

InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation

ResearchDGX agent

arXiv:2604.08646v1 Announce Type: new Abstract: Instruction-based video editing is a natural way to control video content with text, but adapting a video generation model into an editor usually appear

Intrinsic Concept Extraction Based on Compositional Interpretability

ResearchDGX agent

arXiv:2603.11795v2 Announce Type: replace Abstract: Unsupervised Concept Extraction aims to extract concepts from a single image; however, existing methods suffer from the inability to extract composa

LoBE-GS: Load-Balanced and Efficient 3D Gaussian Splatting for Large-Scale Scene Reconstruction

ResearchDGX agent

arXiv:2510.01767v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has established itself as an efficient representation for real-time, high-fidelity 3D scene reconstruction. However, sc

Long-SCOPE: Fully Sparse Long-Range Cooperative 3D Perception

SafetyDGX agent

arXiv:2604.09206v1 Announce Type: new Abstract: Cooperative 3D perception via Vehicle-to-Everything communication is a promising paradigm for enhancing autonomous driving, offering extended sensing ho

Low-Data Supervised Adaptation Outperforms Prompting for Cloud Segmentation Under Domain Shift

Model ReleasesDGX agent

arXiv:2604.08956v1 Announce Type: new Abstract: Adapting vision-language models to remote sensing imagery presents a fundamental challenge: both the visual and linguistic distributions of satellite da

LPLCv2: An Expanded Dataset for Fine-Grained License Plate Legibility Classification

Model ReleasesDGX agent

arXiv:2604.08741v1 Announce Type: new Abstract: Modern Automatic License Plate Recognition (ALPR) systems achieve outstanding performance in controlled, well-defined scenarios. However, large-scale re

LuMon: A Comprehensive Benchmark and Development Suite with Novel Datasets for Lunar Monocular Depth Estimation

Model ReleasesDGX agent

arXiv:2604.09352v1 Announce Type: new Abstract: Monocular Depth Estimation (MDE) is crucial for autonomous lunar rover navigation using electro-optical cameras. However, deploying terrestrial MDE netw

M-IDoL: Information Decomposition for Modality-Specific and Diverse Representation Learning in Medical Foundation Model

TutorialsDGX agent

arXiv:2604.08936v1 Announce Type: new Abstract: Medical foundation models (MFMs) aim to learn universal representations from multimodal medical images that can generalize effectively to diverse downst

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding

AgentsDGX agent

arXiv:2604.09167v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance in multimodal understanding and reasoning, yet grounded reasoning in 3D scenes remains un

MASS: Mesh-inellipse Aligned Deformable Surfel Splatting for Hand Reconstruction and Rendering from Egocentric Monocular Video

ResearchDGX agent

arXiv:2604.08943v1 Announce Type: new Abstract: Reconstructing high-fidelity 3D hands from egocentric monocular videos remains a challenge due to the limitations in capturing high-resolution geometry,

Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory

ApplicationsDGX agent

arXiv:2604.08995v1 Announce Type: new Abstract: With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing

Measurement-Consistent Langevin Corrector for Stabilizing Latent Diffusion Inverse Problem Solvers

ResearchDGX agent

arXiv:2601.04791v3 Announce Type: replace Abstract: While latent diffusion models (LDMs) have emerged as powerful priors for inverse problems, existing LDM-based solvers frequently suffer from instabi

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation

Model ReleasesDGX agent

arXiv:2604.09088v1 Announce Type: new Abstract: Memory-efficient transfer learning (METL) approaches have recently achieved promising performance in adapting pre-trained models to downstream tasks. Th

MeshOn: Intersection-Free Mesh-to-Mesh Composition

SafetyDGX agent

arXiv:2604.08799v1 Announce Type: cross Abstract: We propose MeshOn, a method that finds physically and semantically realistic compositions of two input meshes. Given an accessory, a base mesh with a

MixFlow: Mixed Source Distributions Improve Rectified Flows

SafetyDGX agent

arXiv:2604.09181v1 Announce Type: new Abstract: Diffusion models and their variations, such as rectified flows, generate diverse and high-quality images, but they are still hindered by slow iterative

Multi-task Just Recognizable Difference for Video Coding for Machines: Database, Model, and Coding Application

ResearchDGX agent

arXiv:2604.09421v1 Announce Type: cross Abstract: Just Recognizable Difference (JRD) boosts coding efficiency for machine vision through visibility threshold modeling, but is currently limited to a si

Multimodal Anomaly Detection for Human-Robot Interaction

SafetyDGX agent

arXiv:2604.09326v1 Announce Type: cross Abstract: Ensuring safety and reliability in human-robot interaction (HRI) requires the timely detection of unexpected events that could lead to system failures

Music Audio-Visual Question Answering Requires Specialized Multimodal Designs

ResearchDGX agent

arXiv:2505.20638v2 Announce Type: replace-cross Abstract: While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music nec

MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation

TutorialsDGX agent

arXiv:2604.08916v1 Announce Type: new Abstract: Conventional 3D instance segmentation methods rely on labor-intensive 3D annotations for supervised training, which limits their scalability and general

← Previous
1…200201202203204…207
Next →