AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
21 Apr 2026

MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline

Model ReleasesDGX agent

arXiv:2604.18418v1 Announce Type: new Abstract: Recent advances in deep research systems enable large language models to retrieve, synthesize, and reason over large-scale external knowledge. In medici

Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation

TutorialsDGX agent

arXiv:2604.18215v1 Announce Type: new Abstract: Spatially consistent long-horizon video generation aims to maintain temporal and spatial consistency along predefined camera trajectories. Existing meth

mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.17054v1 Announce Type: new Abstract: Scalable Vector Graphics (SVGs) function both as visual images and as structured code that encode rich geometric and layout information, yet most method

MESA: A Training-Free Multi-Exemplar Deep Framework for Restoring Ancient Inscription Textures

SafetyDGX agent

arXiv:2604.17390v1 Announce Type: new Abstract: Ancient inscriptions frequently suffer missing or corrupted regions from fragmentation, erosion, or other damage, hindering reading, and analysis. We re

MetaCloak-JPEG: JPEG-Robust Adversarial Perturbation for Preventing Unauthorized DreamBooth-Based Deepfake Generation

ResearchDGX agent

arXiv:2604.18537v1 Announce Type: new Abstract: The rapid progress of subject-driven text-to-image synthesis, and in particular DreamBooth, has enabled a consent-free deepfake pipeline: an adversary n

Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs

Model ReleasesDGX agent

arXiv:2603.02618v3 Announce Type: replace Abstract: Out-of-distribution (OOD) detection seeks to identify samples from unknown classes, a critical capability for deploying machine learning models in o

Missing Pattern Tree based Decision Grouping and Ensemble for Enhancing Pair Utilization in Deep Incomplete Multi-View Clustering

Model ReleasesDGX agent

arXiv:2512.21510v2 Announce Type: replace-cross Abstract: Real-world multi-view data often exhibit highly inconsistent missing patterns, posing significant challenges for incomplete multi-view cluster

MLE-UVAD: Minimal Latent Entropy Autoencoder for Fully Unsupervised Video Anomaly Detection

ResearchDGX agent

arXiv:2603.23868v2 Announce Type: replace Abstract: In this paper, we address the challenging problem of single-scene, fully unsupervised video anomaly detection (VAD), where raw videos containing bot

MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models

Model ReleasesDGX agent

arXiv:2601.03331v2 Announce Type: replace Abstract: Recent advances in Vision-Language Models (VLMs) have improved performance in multi-modal learning, raising the question of whether these models tru

MobileAgeNet: Lightweight Facial Age Estimation for Mobile Deployment

Model ReleasesDGX agent

arXiv:2604.17007v1 Announce Type: new Abstract: Mobile deployment of facial age estimation requires models that balance predictive accuracy with low latency and compact size. In this work, we present

Modeling Biomechanical Constraint Violations for Language-Agnostic Lip-Sync Deepfake Detection

ResearchDGX agent

arXiv:2604.16808v1 Announce Type: new Abstract: Current lip-sync deepfake detectors rely on pixel-level artifacts or audio-visual correspondence, failing to generalize across languages because these c

MODEST: Multi-Optics Depth-of-Field Stereo Dataset

AgentsDGX agent

arXiv:2511.20853v3 Announce Type: replace Abstract: Reliable depth estimation under real optical conditions remains a core challenge for camera vision in systems such as autonomous robotics and augmen

Motif-Video 2B: Technical Report

Model ReleasesDGX agent

arXiv:2604.16503v1 Announce Type: new Abstract: Training strong video generation models usually requires massive datasets, large parameter counts, and substantial compute. In this work, we ask whether

Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition

SafetyDGX agent

arXiv:2604.17062v1 Announce Type: new Abstract: Zero-shot action recognition is challenging due to the semantic gap between seen and unseen classes. We present a novel framework that enhances CLIP wit

MU-GeNeRF: Multi-view Uncertainty-guided Generalizable Neural Radiance Fields for Distractor-aware Scene

ApplicationsDGX agent

arXiv:2604.17965v1 Announce Type: new Abstract: Generalizable Neural Radiance Fields (GeNeRFs) enable high-quality scene reconstruction from sparse views and can generalize to unseen scenes. However,

MUA: Mobile Ultra-detailed Animatable Avatars

Local AiDGX agent

arXiv:2604.18583v1 Announce Type: new Abstract: Building photorealistic, animatable full-body digital humans remains a longstanding challenge in computer graphics and vision. Recent advances in animat

Multi-Camera Self-Calibration in Sports Motion Capture: Leveraging Human and Stick Poses

Model ReleasesDGX agent

arXiv:2604.17567v1 Announce Type: new Abstract: Multi-camera systems are widely employed in sports to capture the 3D motion of athletes and equipment, yet calibrating their extrinsic parameters remain

Multi-View Hierarchical Graph Neural Network for Sketch-Based 3D Shape Retrieval

Local AiDGX agent

arXiv:2604.18019v1 Announce Type: new Abstract: Sketch-based 3D shape retrieval (SBSR) aims to retrieve 3D shapes that are consistent with the category of the input hand-drawn sketch. The core challen

Multilevel neural networks with dual-stage feature fusion for human activity recognition

Model ReleasesDGX agent

arXiv:2604.16577v1 Announce Type: new Abstract: Human activity recognition (HAR) refers to the process of identifying human actions and activities using data collected from sensors. Neural networks, s

Multimodal Fusion of Histopathology Images and Electronic Health Records for Early Breast Cancer Diagnosis

ResearchDGX agent

arXiv:2604.17122v1 Announce Type: new Abstract: Breast cancer is a leading cause of cancer-related mortality worldwide, and timely accurate diagnosis is critical to improving survival outcomes. While

MultiWorld: Scalable Multi-Agent Multi-View Video World Models

AgentsDGX agent

arXiv:2604.18564v1 Announce Type: new Abstract: Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as

MuSteerNet: Human Reaction Generation from Videos via Observation-Reaction Mutual Steering

ResearchDGX agent

arXiv:2603.20187v2 Announce Type: replace Abstract: Video-driven human reaction generation aims to synthesize 3D human motions that directly react to observed video sequences, which is crucial for bui

Navigating Distribution Shifts in Medical Image Analysis: A Survey

SafetyDGX agent

arXiv:2411.05824v3 Announce Type: replace-cross Abstract: Medical Image Analysis (MedIA) has become indispensable in modern healthcare, enhancing clinical diagnostics and personalized treatment. Despi

New Fourth-Order Grayscale Indicator-Based Telegraph Diffusion Model for Image Despeckling

ResearchDGX agent

arXiv:2509.26010v2 Announce Type: replace Abstract: Second-order PDE models have been widely used for suppressing multiplicative noise, but they often introduce blocky artifacts in the early stages of

Noise-Adaptive Diffusion Sampling for Inverse Problems Without Task-Specific Tuning

Local AiDGX agent

arXiv:2604.16919v1 Announce Type: cross Abstract: Diffusion models (DMs) have recently shown remarkable performance on inverse problems (IPs). Optimization-based methods can fast solve IPs using DMs a

Noise Injection: Improving Out-of-Distribution Generalization for Limited Size Datasets

TutorialsDGX agent

arXiv:2511.03855v2 Announce Type: replace Abstract: Deep learned (DL) models for image recognition have been shown to fail to generalize to data from different devices, populations, etc. COVID-19 dete

NOOUGAT: Towards Unified Online and Offline Multi-Object Tracking

ApplicationsDGX agent

arXiv:2509.02111v2 Announce Type: replace Abstract: The long-standing division between extit{online} and extit{offline} Multi-Object Tracking (MOT) has led to fragmented solutions that fail to address

NTIRE 2026 Rip Current Detection and Segmentation (RipDetSeg) Challenge Report

Model ReleasesDGX agent

arXiv:2604.17070v1 Announce Type: new Abstract: This report presents the NTIRE 2026 Rip Current Detection and Segmentation (RipDetSeg) Challenge, which targets automatic rip current understanding in i

NullFace: Training-Free Localized Face Anonymization

ApplicationsDGX agent

arXiv:2503.08478v2 Announce Type: replace Abstract: Privacy concerns around ever increasing number of cameras are increasing in today's digital age. Although existing anonymization methods are able to

NVGS: Neural Visibility for Occlusion Culling in 3D Gaussian Splatting

Local AiDGX agent

arXiv:2511.19202v2 Announce Type: replace Abstract: 3D Gaussian Splatting can exploit frustum culling and level-of-detail strategies to accelerate rendering of scenes containing a large number of prim

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning

ResearchDGX agent

arXiv:2604.17052v1 Announce Type: new Abstract: Streaming video reasoning requires models to operate in a setting where history grows without bound while meaningful evidence remains scarce. In such a

OD3: Optimization-free Dataset Distillation for Object Detection

ResearchDGX agent

arXiv:2506.01942v2 Announce Type: replace Abstract: Training large neural networks on large-scale datasets requires substantial computational resources, particularly for dense prediction tasks such as

OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation

Model ReleasesDGX agent

arXiv:2604.18326v1 Announce Type: new Abstract: Recent advancements in audio-video joint generation models have demonstrated impressive capabilities in content creation. However, generating high-fidel

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models

ResearchDGX agent

arXiv:2511.14582v2 Announce Type: replace Abstract: Omnimodal large language models (OmniLLMs) have attracted increasing research attention of late towards unified audio-video understanding. However,

One-Step Diffusion with Inverse Residual Fields for Unsupervised Industrial Anomaly Detection

ResearchDGX agent

arXiv:2604.18393v1 Announce Type: new Abstract: Diffusion models have achieved outstanding performance in unsupervised industrial anomaly detection (uIAD) by learning a manifold of normal data under t

OneDrive: Unified Multi-Paradigm Driving with Vision-Language-Action Models

AgentsDGX agent

arXiv:2604.17915v1 Announce Type: new Abstract: Vision-Language Models(VLMs) excel at autoregressive text generation, yet end-to-end autonomous driving requires multi-task learning with structured out

Operationalizing Fairness in Text-to-Image Models: A Survey of Bias, Fairness Audits and Mitigation Strategies

SafetyDGX agent

arXiv:2604.16516v1 Announce Type: new Abstract: Text-to-Image (T2I) generation models have been widely adopted across various industries, yet are criticized for frequently exhibiting societal stereoty

Optimally Bridging Semantics and Data: Generative Semantic Communication via Schrodinger Bridge

TutorialsDGX agent

arXiv:2604.17802v1 Announce Type: cross Abstract: Generative Semantic Communication (GSC) is a promising solution for image transmission over narrow-band and high-noise channels. However, existing GSC

OptiMVMap: Offline Vectorized Map Construction via Optimal Multi-vehicle Perspectives

Model ReleasesDGX agent

arXiv:2604.17135v1 Announce Type: new Abstract: Offline vectorized maps constitute critical infrastructure for high-precision autonomous driving and mapping services. Existing approaches rely predomin

ORSIFlow: Saliency-Guided Rectified Flow for Optical Remote Sensing Salient Object Detection

ResearchDGX agent

arXiv:2603.28584v2 Announce Type: replace Abstract: Optical Remote Sensing Image Salient Object Detection (ORSI-SOD) remains challenging due to complex backgrounds, low contrast, irregular object shap

Ouroboros: Single-step Diffusion Models for Cycle-consistent Forward and Inverse Rendering

ResearchDGX agent

arXiv:2508.14461v3 Announce Type: replace Abstract: While multi-step diffusion models have advanced both forward and inverse rendering, existing approaches often treat these problems independently, le

OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection

SafetyDGX agent

arXiv:2511.21064v2 Announce Type: replace-cross Abstract: Open-Vocabulary Object Detection (OVOD) aims to enable detectors to generalize across categories by leveraging semantic information. Although

PA-TCNet: Pathology-Aware Temporal Calibration with Physiology-Guided Target Refinement for Cross-Subject Motor Imagery EEG Decoding in Stroke Patients

ResearchDGX agent

arXiv:2604.16554v1 Announce Type: new Abstract: Stroke patient cross-subject electroencephalography (EEG) decoding of motor imagery (MI) brain-computer interface (BCI) is essential for motor rehabilit

PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation

Model ReleasesDGX agent

arXiv:2604.17570v1 Announce Type: new Abstract: Peripheral Blood Smear (PBS) is a critical microscopic examination in hematopathology that yields whole-slide imaging (WSI). Unlike solid tissue patholo

PCM-NeRF: Probabilistic Camera Modeling for Neural Radiance Fields under Pose Uncertainty

ResearchDGX agent

arXiv:2604.17831v1 Announce Type: new Abstract: Neural surface reconstruction methods typically treat camera poses as fixed values, assuming perfect accuracy from Structure-from-Motion (SfM) systems.

Penny Wise, Pixel Foolish: Bypassing Price Constraints in Multimodal Agents via Visual Adversarial Perturbations

Model ReleasesDGX agent

arXiv:2604.16515v1 Announce Type: new Abstract: The rapid proliferation of Multimodal Large Language Models (MLLMs) has enabled mobile agents to execute high-stakes financial transactions, but their a

PEPR: Privileged Event-based Predictive Regularization for Domain Generalization

SafetyDGX agent

arXiv:2602.04583v2 Announce Type: replace Abstract: Deep neural networks for visual perception are highly susceptible to domain shift, which poses a critical challenge for real-world deployment under

Perspective-Equivariant Fine-tuning for Multispectral Demosaicing without Ground Truth

AgentsDGX agent

arXiv:2603.01332v2 Announce Type: replace Abstract: Multispectral demosaicing is crucial to reconstruct full-resolution spectral images from snapshot mosaiced measurements, enabling real-time imaging

PestVL-Net: Enabling Multimodal Pest Learning via Fine-grained Vision-Language Interaction

ApplicationsDGX agent

arXiv:2604.17278v1 Announce Type: new Abstract: Effective pest recognition and management are crucial for sustainable agricultural development. However, collecting pest data in real scenarios is often

Physics-Informed Tracking (PIT)

ResearchDGX agent

arXiv:2604.16895v1 Announce Type: new Abstract: We propose Physics-Informed Tracking (PIT), a video-based framework for tracking a single particle from video, where a neural network autoencoder locali

PlankFormer: Robust Plankton Instance Segmentation via MAE-Pretrained Vision Transformers and Pseudo Community Image Generation

ApplicationsDGX agent

arXiv:2604.17856v1 Announce Type: new Abstract: Plankton monitoring is essential for assessing aquatic ecosystems but is limited by the labor-intensive nature of manual microscopic analysis. Automatin

PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks

Model ReleasesDGX agent

arXiv:2602.06663v2 Announce Type: replace Abstract: Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their

PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems

ResearchDGX agent

arXiv:2604.16540v1 Announce Type: new Abstract: Poisoning input views of 3D reconstruction systems has been recently studied. However, we identify that existing studies simply backpropagate adversaria

Positioning radiata pine branches requiring pruning by drone stereo vision

AgentsDGX agent

arXiv:2604.16480v1 Announce Type: new Abstract: This paper presents a stereo-vision-based system mounted on a drone for detecting and localising radiata pine branches to support autonomous pruning. Th

PowerCLIP: Powerset Alignment for Contrastive Pre-Training

SafetyDGX agent

arXiv:2511.23170v5 Announce Type: replace Abstract: Contrastive vision-language pre-training frameworks such as CLIP have demonstrated impressive zero-shot performance across a range of vision-languag

PPEDCRF: Dynamic-CRF-Guided Selective Perturbation for Background-Based Location Privacy in Video Sequences

Model ReleasesDGX agent

arXiv:2604.17163v1 Announce Type: new Abstract: We propose PPEDCRF, a calibrated selective perturbation framework that protects background-based location privacy in released video frames against galle

Predicting Blastocyst Formation in IVF: Integrating DINOv2 and Attention-Based LSTM on Time-Lapse Embryo Images

ResearchDGX agent

arXiv:2604.16505v1 Announce Type: new Abstract: The selection of the optimal embryo for transfer is a critical yet challenging step in in vitro fertilization (IVF), primarily due to its reliance on th

Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

Model ReleasesDGX agent

arXiv:2508.20751v2 Announce Type: replace Abstract: Recent advancements highlight the importance of GRPO-based reinforcement learning methods and benchmarking in enhancing text-to-image (T2I) generati

Preparation of Fractal-Inspired Computational Architectures for Automated Neural Design Exploration

ResearchDGX agent

arXiv:2511.07329v3 Announce Type: replace-cross Abstract: It introduces FractalNet, a fractal-inspired computational architectures for advanced large language model analysis that mainly challenges mod

Privacy-Preserving Semantic Segmentation without Key Management

Local AiDGX agent

arXiv:2604.16523v1 Announce Type: new Abstract: This paper proposes a novel privacy-preserving semantic segmentation method that can use independent keys for each client and image. In the proposed met

← Previous
1…183184185186187…209
Next →