AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Research

Beyond Flicker: Detecting Kinematic Inconsistencies for Generalizable Deepfake Video Detection

DGX agent

arXiv:2512.04175v2 Announce Type: replace Abstract: Generalizing deepfake detection to unseen manipulations remains a key challenge. A recent approach to tackle this issue is to train a network with p

researcharxiv-cs-cv
13 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Beyond Segmentation: Structurally Informed Facade Parsing from Imperfect Images

DGX agent

arXiv:2604.09260v1 Announce Type: new Abstract: Standard object detectors typically treat architectural elements independently, often resulting in facade parsings that lack the structural coherence re

safetyarxiv-cs-cv
13 Apr 2026
Safety

BIAS: A Biologically Inspired Algorithm for Video Saliency Detection

DGX agent

arXiv:2604.08858v1 Announce Type: new Abstract: We present BIAS, a fast, biologically inspired model for dynamic visual saliency detection in continuous video streams. Building on the Itti--Koch frame

safetyarxiv-cs-cv
13 Apr 2026
Research

BlendFusion -- Scalable Synthetic Data Generation for Diffusion Model Training

DGX agent

arXiv:2604.09022v1 Announce Type: new Abstract: With the rapid adoption of diffusion models, synthetic data generation has emerged as a promising approach for addressing the growing demand for large-s

researcharxiv-cs-cv
13 Apr 2026
Model Releases

CAD 100K: A Comprehensive Multi-Task Dataset for Car Related Visual Anomaly Detection

DGX agent

arXiv:2604.09023v1 Announce Type: new Abstract: Multi-task visual anomaly detection is critical for car-related manufacturing quality assessment. However, existing methods remain task-specific, hinder

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

CatalogStitch: Dimension-Aware and Occlusion-Preserving Object Compositing for Catalog Image Generation

DGX agent

arXiv:2604.08836v1 Announce Type: new Abstract: Generative object compositing methods have shown remarkable ability to seamlessly insert objects into scenes. However, when applied to real-world catalo

model-releasesarxiv-cs-cv
13 Apr 2026
Safety

CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Intervention

DGX agent

arXiv:2603.18561v2 Announce Type: replace Abstract: Planning-oriented end-to-end driving models show great promise, yet they fundamentally learn statistical correlations instead of true causal relatio

safetyarxiv-cs-cv
13 Apr 2026
Research

Characterizing Lidar Range-Measurement Ambiguity due to Multiple Returns

DGX agent

arXiv:2604.09282v1 Announce Type: cross Abstract: Reliable position and attitude sensing is critical for highly automated vehicles that operate on conventional roadways. Lidar sensors are increasingly

researcharxiv-cs-cv
13 Apr 2026
Model Releases

Cluster-First Labelling: An Automated Pipeline for Segmentation and Morphological Clustering in Histology Whole Slide Images

DGX agent

arXiv:2604.09370v1 Announce Type: cross Abstract: Labelling tissue components in histology whole slide images (WSIs) is prohibitively labour-intensive: a single slide may contain tens of thousands of

model-releasesarxiv-cs-cv
13 Apr 2026
Safety

ClusterMark: Towards Robust Watermarking for Autoregressive Image Generators with Visual Token Clustering

DGX agent

arXiv:2508.06656v2 Announce Type: replace Abstract: In-generation watermarking for latent diffusion models has recently shown high robustness in marking generated images for easier detection and attri

safetyarxiv-cs-cv
13 Apr 2026
Applications

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos

DGX agent

arXiv:2601.10632v2 Announce Type: replace Abstract: In this paper, we find that the generation of 3D human motions and 2D human videos is intrinsically coupled. 3D motions provide the structural prior

applicationsarxiv-cs-cv
13 Apr 2026
Model Releases

Compositional-Degradation UAV Image Restoration: Conditional Decoupled MoE Network and A Benchmark

DGX agent

arXiv:2604.09313v1 Announce Type: cross Abstract: UAV images are critical for applications such as large-area mapping, infrastructure inspection, and emergency response. However, in real-world flight

model-releasesarxiv-cs-cv
13 Apr 2026
Research

Cross-Modal Knowledge Distillation from Spatial Transcriptomics to Histology

DGX agent

arXiv:2604.09076v1 Announce Type: new Abstract: Spatial transcriptomics provides a molecularly rich description of tissue organization, enabling unsupervised discovery of tissue niches -- spatially co

researcharxiv-cs-cv
13 Apr 2026
Tutorials

CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation

DGX agent

arXiv:2604.09201v1 Announce Type: new Abstract: Camera-controllable video generation aims to synthesize videos with flexible and physically plausible camera movements. However, existing methods either

tutorialsarxiv-cs-cv
13 Apr 2026
Research

Customized Fusion: A Closed-Loop Dynamic Network for Adaptive Multi-Task-Aware Infrared-Visible Image Fusion

DGX agent

arXiv:2604.08924v1 Announce Type: new Abstract: Infrared-visible image fusion aims to integrate complementary information for robust visual understanding, but existing fusion methods struggle with sim

researcharxiv-cs-cv
13 Apr 2026
Research

Deep Light Pollution Removal in Night Cityscape Photographs

DGX agent

arXiv:2604.09145v1 Announce Type: new Abstract: Nighttime photography is severely degraded by light pollution induced by pervasive artificial lighting in urban environments. After long-range scatterin

researcharxiv-cs-cv
13 Apr 2026
Model Releases

DeFakeQ: Enabling Real-Time Deepfake Detection on Edge Devices via Adaptive Bidirectional Quantization

DGX agent

arXiv:2604.08847v1 Announce Type: new Abstract: Deepfake detection has become a fundamental component of modern media forensics. Despite significant progress in detection accuracy, most existing metho

model-releasesarxiv-cs-cv
13 Apr 2026
Tutorials

Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenarios

DGX agent

arXiv:2604.08922v1 Announce Type: new Abstract: Complex degradations like noise, blur, and low resolution are typical challenges in real world image fusion tasks, limiting the performance and practica

tutorialsarxiv-cs-cv
13 Apr 2026
Research

Detecting Diffusion-generated Images via Dynamic Assembly ForestsDetecting Diffusion-generated Images via Dynamic Assembly Forests

DGX agent

arXiv:2604.09106v1 Announce Type: new Abstract: Diffusion models are known for generating high-quality images, causing serious security concerns. To combat this, most efforts rely on deep neural netwo

researcharxiv-cs-cv
13 Apr 2026
Research

Do Vision Language Models Need to Process Image Tokens?

DGX agent

arXiv:2604.09425v1 Announce Type: new Abstract: Vision Language Models (VLMs) have achieved remarkable success by integrating visual encoders with large language models (LLMs). While VLMs process dens

researcharxiv-cs-cv
13 Apr 2026
Research

Domain-generalizable Face Anti-Spoofing with Patch-based Multi-tasking and Artifact Pattern Conversion

DGX agent

arXiv:2604.09018v1 Announce Type: new Abstract: Face Anti-Spoofing (FAS) algorithms, designed to secure face recognition systems against spoofing, struggle with limited dataset diversity, impairing th

researcharxiv-cs-cv
13 Apr 2026
Model Releases

DSVTLA: Deep Swin Vision Transformer-Based Transfer Learning Architecture for Multi-Type Cancer Histopathological Cancer Image Classification

DGX agent

arXiv:2604.09468v1 Announce Type: cross Abstract: In this study, we proposed a deep Swin-Vision Transformer-based transfer learning architecture for robust multi-cancer histopathological image classif

model-releasesarxiv-cs-cv
13 Apr 2026
Safety

Dynamic Class-Aware Active Learning for Unbiased Satellite Image Segmentation

DGX agent

arXiv:2604.08965v1 Announce Type: new Abstract: Semantic segmentation of satellite imagery plays a vital role in land cover mapping and environmental monitoring. However, annotating large-scale, high-

safetyarxiv-cs-cv
13 Apr 2026
Applications

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection

DGX agent

arXiv:2604.09164v1 Announce Type: new Abstract: Temporal human action detection aims to identify and localize action segments within untrimmed videos, serving as a pivotal task in video understanding.

applicationsarxiv-cs-cv
13 Apr 2026
Research

Efficient Unlearning through Maximizing Relearning Convergence Delay

DGX agent

arXiv:2604.09391v1 Announce Type: cross Abstract: Machine unlearning poses challenges in removing mislabeled, contaminated, or problematic data from a pretrained model. Current unlearning approaches a

researcharxiv-cs-cv
13 Apr 2026
Research

EfficientSign: An Attention-Enhanced Lightweight Architecture for Indian Sign Language Recognition

DGX agent

arXiv:2604.08694v1 Announce Type: new Abstract: How do you build a sign language recognizer that works on a phone? That question drove this work. We built EfficientSign, a lightweight model which take

researcharxiv-cs-cv
13 Apr 2026
Safety

EGLOCE: Training-Free Energy-Guided Latent Optimization for Concept Erasure

DGX agent

arXiv:2604.09405v1 Announce Type: new Abstract: As text-to-image diffusion models grow increasingly prevalent, the ability to remove specific concepts-mostly explicit content and many copyrighted char

safetyarxiv-cs-cv
13 Apr 2026
Model Releases

EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks

DGX agent

arXiv:2604.09535v1 Announce Type: new Abstract: Large foundation models have made significant advances in embodied intelligence, enabling synthesis and reasoning over egocentric input for household ta

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

ELT: Elastic Looped Transformers for Visual Generation

DGX agent

arXiv:2604.09168v1 Announce Type: new Abstract: We introduce Elastic Looped Transformers (ELT), a highly parameter-efficient class of visual generative models based on a recurrent transformer architec

model-releasesarxiv-cs-cv
13 Apr 2026
Safety

EmoCtrl: Controllable Emotional Image Content Generation

DGX agent

arXiv:2512.22437v2 Announce Type: replace Abstract: An image conveys meaning through both its visual content and emotional tone, jointly shaping human perception. We introduce Controllable Emotional I

safetyarxiv-cs-cv
13 Apr 2026
Applications

Enhanced Self-Supervised Multi-Image Super-Resolution for Camera Array Images

DGX agent

arXiv:2604.06816v2 Announce Type: replace-cross Abstract: Conventional multi-image super-resolution (MISR) methods, such as burst and video SR, rely on sequential frames from a single camera. Conseque

applicationsarxiv-cs-cv
13 Apr 2026
Agents

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration

DGX agent

arXiv:2604.09367v1 Announce Type: new Abstract: Ancient inscriptions, as repositories of cultural memory, have suffered from centuries of environmental and human-induced degradation. Restoring their i

agentsarxiv-cs-cv
13 Apr 2026
Local Ai

FaceLiVTv2: An Improved Hybrid Architecture for Efficient Mobile Face Recognition

DGX agent

arXiv:2604.09127v1 Announce Type: new Abstract: Lightweight face recognition is increasingly important for deployment on edge and mobile devices, where strict constraints on latency, memory, and energ

local-aiarxiv-cs-cv
13 Apr 2026
Model Releases

FashionStylist: An Expert Knowledge-enhanced Multimodal Dataset for Fashion Understanding

DGX agent

arXiv:2604.09249v1 Announce Type: new Abstract: Fashion understanding requires both visual perception and expert-level reasoning about style, occasion, compatibility, and outfit rationale. However, ex

model-releasesarxiv-cs-cv
13 Apr 2026
Hardware

Fast Model-guided Instance-wise Adaptation Framework for Real-world Pansharpening with Fidelity Constraints

DGX agent

arXiv:2604.08903v1 Announce Type: new Abstract: Pansharpening aims to generate high-resolution multispectral (HRMS) images by fusing low-resolution multispectral (LRMS) and high-resolution panchromati

hardwarearxiv-cs-cv
13 Apr 2026
Research

FDIF: Formula-Driven supervised Learning with Implicit Functions for 3D Medical Image Segmentation

DGX agent

arXiv:2603.23199v2 Announce Type: replace Abstract: Deep learning-based 3D medical image segmentation methods relies on large-scale labeled datasets, yet acquiring such data is difficult due to privac

researcharxiv-cs-cv
13 Apr 2026
Model Releases

Few-Shot Personalized Age Estimation

DGX agent

arXiv:2604.09125v1 Announce Type: new Abstract: Existing age estimation methods treat each face as an independent sample, learning a global mapping from appearance to age. This ignores a well-document

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

Fine-Grained Action Segmentation for Renorrhaphy in Robot-Assisted Partial Nephrectomy

DGX agent

arXiv:2604.09051v1 Announce Type: new Abstract: Fine-grained action segmentation during renorrhaphy in robot-assisted partial nephrectomy requires frame-level recognition of visually similar suturing

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

FIRE-CIR: Fine-grained Reasoning for Composed Fashion Image Retrieval

DGX agent

arXiv:2604.09114v1 Announce Type: new Abstract: Composed image retrieval (CIR) aims to retrieve a target image that depicts a reference image modified by a textual description. While recent vision-lan

model-releasesarxiv-cs-cv
13 Apr 2026
Hardware

FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs

DGX agent

arXiv:2512.20033v2 Announce Type: replace Abstract: We present FlashLips, a two-stage, mask-free lip-sync system that decouples lips control from rendering and achieves real-time performance, with our

hardwarearxiv-cs-cv
13 Apr 2026
Local Ai

From Frames to Events: Rethinking Evaluation in Human-Centric Video Anomaly Detection

DGX agent

arXiv:2604.09327v1 Announce Type: new Abstract: Pose-based Video Anomaly Detection (VAD) has gained significant attention for its privacy-preserving nature and robustness to environmental variations.

local-aiarxiv-cs-cv
13 Apr 2026
Tutorials

Generative View Stitching

DGX agent

arXiv:2510.24718v3 Announce Type: replace Abstract: Autoregressive video diffusion models are capable of long rollouts that are stable and consistent with history, but they are unable to guide the cur

tutorialsarxiv-cs-cv
13 Apr 2026
Research

Geometry Reinforced Efficient Attention Tuning Equipped with Normals for Robust Stereo Matching

DGX agent

arXiv:2604.09142v1 Announce Type: new Abstract: Despite remarkable advances in image-driven stereo matching over the past decade, Synthetic-to-Realistic Zero-Shot (Syn-to-Real) generalization remains

researcharxiv-cs-cv
13 Apr 2026
Model Releases

GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing

DGX agent

arXiv:2604.08896v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) have accelerated progress in domain-oriented AI, yet their development in geoscience and rem

model-releasesarxiv-cs-cv
13 Apr 2026
Agents

GeRM: A Generative Rendering Model From Physically Realistic to Photorealistic

DGX agent

arXiv:2604.09304v1 Announce Type: new Abstract: For decades, Physically-Based Rendering (PBR) is the fundation of synthesizing photorealisitic images, and therefore sometimes roughly referred as Photo

agentsarxiv-cs-cv
13 Apr 2026
Research

Globally Optimal Pose from Orthographic Silhouettes

DGX agent

arXiv:2604.09199v1 Announce Type: new Abstract: We solve the problem of determining the pose of known shapes in R^3 from their unoccluded silhouettes. The pose is determined up to global optimality us

researcharxiv-cs-cv
13 Apr 2026
Research

HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models

DGX agent

arXiv:2604.06165v2 Announce Type: replace Abstract: Large vision-language models can produce object hallucinations in image descriptions, highlighting the need for effective detection and mitigation s

researcharxiv-cs-cv
13 Apr 2026
Research

Harnessing Weak Pair Uncertainty for Text-based Person Search

DGX agent

arXiv:2604.08877v1 Announce Type: new Abstract: In this paper, we study the text-based person search, which is to retrieve the person of interest via natural language description. Prevailing methods u

researcharxiv-cs-cv
13 Apr 2026
← Previous
1…250251252253254…259
Next →