AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
30 Apr 2026

3D Generation for Embodied AI and Robotic Simulation: A Survey

AgentsDGX agent

arXiv:2604.26509v1 Announce Type: cross Abstract: Embodied AI and robotic systems increasingly depend on scalable, diverse, and physically grounded 3D content for simulation-based training and real-wo

3D-LENS: A 3D Lifting-based Elevated Novel-view Synthesis method for Single-View Aerial-Ground Re-Identification

Model ReleasesDGX agent

arXiv:2604.26520v1 Announce Type: new Abstract: Aerial-Ground Re-Identification (AG-ReID) is constrained by the viewpoint-domain gap, as drastic viewpoint disparities occlude or distort discriminative

A Diffeomorphism Groupoid and Algebroid Framework for Discontinuous Image Registration

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2603.11806v2 Announce Type: replace-cross Abstract: In this paper, we propose a novel mathematical framework for piecewise diffeomorphic image registration that involves discontinuous sliding mo

A Multimodal Depth-Aware Method For Embodied Reference Understanding

ResearchDGX agent

arXiv:2510.08278v3 Announce Type: replace Abstract: Embodied Reference Understanding requires identifying a target object in a visual scene based on both language instructions and pointing cues. While

A Multimodal Pre-trained Network for Integrated EEG-Video Seizure Detection

SafetyDGX agent

arXiv:2604.26379v1 Announce Type: new Abstract: Reliable seizure detection in mouse models is essential for preclinical epilepsy research, yet manual review of synchronized video-EEG recordings is lab

A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows

Model ReleasesDGX agent

arXiv:2604.26462v1 Announce Type: new Abstract: Structured information extraction from long, multilingual scanned financial documents is a core requirement in industrial KYC and compliance workflows.

Action Hints: Semantic Typicality and Context Uniqueness for Generalizable Skeleton-based Video Anomaly Detection

Local AiDGX agent

arXiv:2509.11058v2 Announce Type: replace Abstract: Zero-Shot Video Anomaly Detection (ZS-VAD) requires temporally localizing anomalies without target domain training data, which is a crucial task due

Adaptive Transform Coding for Semantic Compression

ResearchDGX agent

arXiv:2604.26492v1 Announce Type: cross Abstract: Visual data compression is shifting from human-centered reconstruction to machine-oriented representation coding. In this setting, an image is often m

AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision

Model ReleasesDGX agent

arXiv:2604.26567v1 Announce Type: new Abstract: Despite the rapid progress in data-driven 3D vision, aerial geometric 3D vision remains a formidable challenge due to the severe scarcity of large-scale

AnimateAnyMesh++: A Flexible 4D Foundation Model for High-Fidelity Text-Driven Mesh Animation

ResearchDGX agent

arXiv:2604.26917v1 Announce Type: new Abstract: Recent advances in 4D content generation have attracted increasing attention, yet creating high-quality animated 3D models remains challenging due to th

Are Data Augmentation and Segmentation Always Necessary? Insights from COVID-19 X-Rays and a Methodology Thereof

ResearchDGX agent

arXiv:2604.26437v1 Announce Type: new Abstract: Purpose: Rapid and reliable diagnostic tools are crucial for managing respiratory diseases like COVID-19, where chest X-ray analysis coupled with artifi

Attribution-Guided Multimodal Deepfake Detection via Cross-Modal Forensic Fingerprints

SafetyDGX agent

arXiv:2604.26453v1 Announce Type: new Abstract: Audio-visual deepfakes have reached a level of realism that makes perceptual detection unreliable, threatening media integrity and biometric security. W

Benchmarking Deep Learning and Vision Foundation Models for Atypical vs. Normal Mitosis Classification with Cross-Dataset Evaluation

Model ReleasesDGX agent

arXiv:2506.21444v4 Announce Type: replace Abstract: Atypical mitosis marks a deviation in the cell division process that has been shown be an independent prognostic marker for tumor malignancy. Howeve

Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion Models

Model ReleasesDGX agent

arXiv:2604.26365v1 Announce Type: new Abstract: To address the high sampling cost of Diffusion Transformers (DiTs), feature caching offers a training-free acceleration method. However, existing method

Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning

SafetyDGX agent

arXiv:2604.26250v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have achieved state-of-the-art performance in general visual tasks, their perceptual robustness remains remarkably b

Breaking the Rigid Prior: Towards Articulated 3D Anomaly Detection

Model ReleasesDGX agent

arXiv:2604.26868v1 Announce Type: new Abstract: Existing 3D anomaly detection methods are built on a rigid prior: normal geometry is pose-invariant and can be canonicalized through registration or ali

Bridge: Basis-Driven Causal Inference Marries VFMs for Domain Generalization

Model ReleasesDGX agent

arXiv:2604.26820v1 Announce Type: new Abstract: Detectors often suffer from degraded performance, primarily due to the distributional gap between the source and target domains. This issue is especiall

Camera-RFID Fusion for Robust Asset Tracking in Forested Environments

ResearchDGX agent

arXiv:2604.26241v1 Announce Type: new Abstract: Passive RFID tags offer a cost-effective and scalable solution for tracking numerous deployed assets. However, in forested environments, signal attenuat

ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch

ResearchDGX agent

arXiv:2601.13606v2 Announce Type: replace Abstract: Chart reasoning is a critical capability for Vision Language Models (VLMs). However, the development of open-source models is severely hindered by t

Circular Phase Representation and Geometry-Aware Optimization for Ptychographic Image Reconstruction

ResearchDGX agent

arXiv:2604.26664v1 Announce Type: cross Abstract: Traditional iterative reconstruction methods are accurate but computationally expensive, limiting their use in high-throughput and real-time ptychogra

CO-EVO: Co-evolving Semantic Anchoring and Style Diversification for Federated DG-ReID

TutorialsDGX agent

arXiv:2604.26363v1 Announce Type: new Abstract: Federated domain generalization for person re-identification (FedDG-ReID) aims to collaboratively train a pedestrian retrieval model across multiple dec

Color-Encoded Illumination for High-Speed Volumetric Scene Reconstruction

ApplicationsDGX agent

arXiv:2604.26920v1 Announce Type: new Abstract: The task of capturing and rendering 3D dynamic scenes from 2D images has become increasingly popular in recent years. However, most conventional cameras

COMMA: Coordinate-aware Modulated Mamba Network for 3D Dispersed Vessel Segmentation

ResearchDGX agent

arXiv:2503.02332v3 Announce Type: replace-cross Abstract: Accurate segmentation of 3D vascular structures is essential for various medical imaging applications. The dispersed nature of vascular struct

Contrastive Heliophysical Image Pretraining for Solar Dynamics Observatory Records

ResearchDGX agent

arXiv:2511.22958v2 Announce Type: replace Abstract: Deep learning has revolutionized solar image analysis, yet most approaches train task-specific encoders from scratch or rely on natural-image pretra

COP-GEN: Latent Diffusion Transformer for Copernicus Earth Observation Data

Model ReleasesDGX agent

arXiv:2603.03239v2 Announce Type: replace Abstract: Earth observation applications increasingly rely on data from multiple sensors, including optical, radar, elevation, and land-cover. Relationships b

Cross-Domain Transfer of Hyperspectral Foundation Models

Model ReleasesDGX agent

arXiv:2604.26478v1 Announce Type: new Abstract: Hyperspectral imaging (HSI) semantic segmentation typically relies on in-domain training, but limited data availability often restricts model performanc

CurEvo: Curriculum-Guided Self-Evolution for Video Understanding

Model ReleasesDGX agent

arXiv:2604.26707v1 Announce Type: new Abstract: Recent advances in self-evolution video understanding frameworks have demonstrated the potential of autonomous learning without human annotations. Howev

Decoupled Prototype Matching with Vision Foundation Models for Few-Shot Industrial Object Detection

Model ReleasesDGX agent

arXiv:2604.26404v1 Announce Type: new Abstract: Industrial object detection systems typically rely on large annotated datasets, which are expensive to collect and challenging to maintain in industrial

Delta Score Matters! Spatial Adaptive Multi Guidance in Diffusion Models

SafetyDGX agent

arXiv:2604.26503v1 Announce Type: new Abstract: Diffusion models have achieved remarkable success in synthesizing complex static and temporal visuals, a breakthrough largely driven by Classifier-Free

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation

Model ReleasesDGX agent

arXiv:2604.26565v1 Announce Type: new Abstract: Long-term video understanding requires interpreting complex temporal events and reasoning over procedural activities. While instructional video corpora,

Edge AI for Automotive Vulnerable Road User Safety: Deployable Detection via Knowledge Distillation

Local AiDGX agent

arXiv:2604.26857v1 Announce Type: new Abstract: Deploying accurate object detection for Vulnerable Road User (VRU) safety on edge hardware requires balancing model capacity against computational const

Efficient Zero-Shot Inpainting with Decoupled Diffusion Guidance

Local AiDGX agent

arXiv:2512.18365v2 Announce Type: replace Abstract: Diffusion models have emerged as powerful priors for image editing tasks such as inpainting and local modification, where the objective is to genera

EnerGS: Energy-Based Gaussian Splatting with Partial Geometric Priors

ResearchDGX agent

arXiv:2604.26238v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has been widely adopted for scene reconstruction, where training inherently constitutes a highly coupled and non-convex opt

Event-based Liveness Detection using Temporal Ocular Dynamics: An Exploratory Approach

ResearchDGX agent

arXiv:2604.26285v1 Announce Type: new Abstract: Face liveness detection has been extensively studied using RGB cameras, achieving strong performance under controlled conditions but often failing to ge

ext{PKS}^4:Parallel Kinematic Selective State Space Scanners for Efficient Video Understanding

Model ReleasesDGX agent

arXiv:2604.26461v1 Announce Type: new Abstract: Temporal modeling remains a fundamental challenge in video understanding, particularly as sequence lengths scale. Traditional video models relying on de

FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing

ResearchDGX agent

arXiv:2604.26186v1 Announce Type: new Abstract: Fashion AI systems routinely encode the aesthetic logic of specific houses, editors, and historical moments without disclosing it. We present FASH-iCNN,

FASTER: Rethinking Real-Time Flow VLAs

ApplicationsDGX agent

arXiv:2603.19199v2 Announce Type: replace-cross Abstract: Real-time execution is crucial for deploying Vision-Language-Action (VLA) models in the physical world. Existing asynchronous inference method

Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners

ResearchDGX agent

arXiv:2604.26488v1 Announce Type: new Abstract: One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack r

Federated Medical Image Classification under Class and Domain Imbalance exploiting Synthetic Sample Generation

ResearchDGX agent

arXiv:2604.26324v1 Announce Type: new Abstract: Exploiting deep learning in medical imaging faces critical challenges, including strict privacy constraints, heterogeneous imaging devices with varying

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding

SafetyDGX agent

arXiv:2504.09925v3 Announce Type: replace Abstract: We introduce FLARE, a family of vision language models (VLMs) with a fully vision-language alignment and integration paradigm. Unlike existing appro

Foundation Model-Driven Semantic Change Detection in Remote Sensing Imagery

ResearchDGX agent

arXiv:2602.13780v2 Announce Type: replace Abstract: Remote sensing (RS) change detection is essential for interpreting surface dynamics. Semantic change detection (SCD) further enables pixel-level und

FunFace: Feature Utility and Norm Estimation for Face Recognition

ResearchDGX agent

arXiv:2604.26598v1 Announce Type: new Abstract: Face Recognition (FR) is used in a variety of application domains, from entertainment and banking to security and surveillance. Such applications rely o

GaitKD: A Universal Decoupled Distillation Framework for Efficient Gait Recognition

ApplicationsDGX agent

arXiv:2604.26255v1 Announce Type: new Abstract: Gait recognition is an attractive biometric modality for long-range and contact-free identification, but high-performing gait models often rely on deep

GateMOT: Q-Gated Attention for Dense Object Tracking

ResearchDGX agent

arXiv:2604.26353v1 Announce Type: new Abstract: While large models demonstrate the strong representational power of vanilla attention, this core mechanism cannot be directly applied to Dense Object Tr

Generalized Disguise Makeup Presentation Attack Detection Using an Attention-Guided Patch-Based Framework

TutorialsDGX agent

arXiv:2604.26025v1 Announce Type: new Abstract: Despite significant advances in facial recognition systems, they remain vulnerable to face presentation attacks. Among them, disguise makeup attacks are

GIFGuard: Proactive Forensics against Deepfakes in Facial GIFs via Spatiotemporal Watermarking

Model ReleasesDGX agent

arXiv:2604.26519v1 Announce Type: new Abstract: The rapid evolution of deepfake technology poses an unprecedented threat to the authenticity of Graphics Interchange Format (GIF) imagery, which serves

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

AgentsDGX agent

arXiv:2604.26752v1 Announce Type: new Abstract: We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environmen

GNC-Pose: Geometry-Aware GNC-PnP for Accurate 6D Pose Estimation

SafetyDGX agent

arXiv:2512.06565v2 Announce Type: replace Abstract: We present GNC-Pose, a fully learning-free monocular 6D object pose estimation pipeline for textured objects that combines rendering-based initializ

Graph-based Semantic Calibration Network for Unaligned UAV RGBT Image Semantic Segmentation and A Large-scale Benchmark

Model ReleasesDGX agent

arXiv:2604.26893v1 Announce Type: new Abstract: Fine-grained RGBT image semantic segmentation is crucial for all-weather unmanned aerial vehicle (UAV) scene understanding. However, UAV RGBT semantic s

Hearing the Room Through the Shape of the Drum: Modal-Guided Sound Recovery from Multi-Point Surface Vibrations

ResearchDGX agent

arXiv:2604.26678v1 Announce Type: new Abstract: Optical vibration sensing enables recovering the scene sound directly from the surface vibration of nearby objects, turning everyday objects into ``visu

High-Dimensional Noise to Low-Dimensional Manifolds: A Manifold-Space Diffusion Framework for Degraded Hyperspectral Image Classification

ApplicationsDGX agent

arXiv:2604.26279v1 Announce Type: new Abstract: Recently, Hyperspectral Image (HSI) classification has attracted increasing attention in remote sensing. However, HSI data are inherently high-dimension

HOI-aware Adaptive Network for Weakly-supervised Action Segmentation

TutorialsDGX agent

arXiv:2604.26227v1 Announce Type: new Abstract: In this paper, we propose an HOI-aware adaptive network named AdaAct for weakly-supervised action segmentation. Most existing methods learn a fixed netw

HumanOmni-Speaker: Identifying Who said What and When

Model ReleasesDGX agent

arXiv:2603.21664v2 Announce Type: replace Abstract: While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human intera

KAYRA: A Microservice Architecture for AI-Assisted Karyotyping with Cloud and On-Premise Deployment

ResearchDGX agent

arXiv:2604.26869v1 Announce Type: cross Abstract: We present KAYRA, an end-to-end karyotyping system that operates inside the operational constraints of a clinical cytogenetic laboratory. KAYRA is arc

Last-Layer-Centric Feature Recombination: Unleashing 3D Geometric Knowledge in DINOv3 for Monocular Depth Estimation

ResearchDGX agent

arXiv:2604.26454v1 Announce Type: new Abstract: Monocular depth estimation (MDE) is a fundamental yet inherently ill-posed task. Recent vision foundation models (VFMs), particularly DINO-based transfo

Learning Sparse BRDF Measurement Samples from Image

TutorialsDGX agent

arXiv:2604.26740v1 Announce Type: new Abstract: Accurate BRDF acquisition is important for realistic rendering, but dense gonioreflectometer measurements are slow and expensive. We study how to select

Learning Vision-Based Omnidirectional Navigation: A Teacher-Student Approach Using Monocular Depth Estimation

SafetyDGX agent

arXiv:2603.01999v2 Announce Type: replace-cross Abstract: Reliable obstacle avoidance in industrial settings demands 3D scene understanding, but widely used 2D LiDAR sensors perceive only a single hor

MesonGS++: Post-training Compression of 3D Gaussian Splatting with Hyperparameter Searching

HardwareDGX agent

arXiv:2604.26799v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) achieves high-quality novel view synthesis with real-time rendering, but its storage cost remains prohibitive for practical

MixerCA: An Efficient and Accurate Model for High-Performance Hyperspectral Image Classification

Model ReleasesDGX agent

arXiv:2604.26138v1 Announce Type: new Abstract: Over the past decade, hyperspectral image (HSI) classification has drawn considerable interest due to HSIs' ability to effectively distinguish terrestri

Motion-Driven Multi-Object Tracking of Model Organisms in Space Science Experiments

ResearchDGX agent

arXiv:2604.26321v1 Announce Type: new Abstract: Automated animal behavior analysis relies on long-term, interpretable individual trajectories; however, multi-animal tracking in space science experimen

← Previous
1…164165166167168…209
Next →