AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
20 Apr 2026

Hero-Mamba: Mamba-based Dual Domain Learning for Underwater Image Enhancement

Model ReleasesDGX agent

arXiv:2604.16266v1 Announce Type: new Abstract: Underwater images often suffer from severe degradation, such as color distortion, low contrast, and blurred details, due to light absorption and scatter

Hierarchical Codec Diffusion for Video-to-Speech Generation

SafetyDGX agent

arXiv:2604.15923v1 Announce Type: cross Abstract: Video-to-Speech (VTS) generation aims to synthesize speech from a silent video without auditory signals. However, existing VTS methods disregard the h

HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images

TutorialsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2603.02210v3 Announce Type: replace Abstract: Human-product images, which showcase the integration of humans and products, play a vital role in advertising, e-commerce, and digital marketing. Th

HyCal: A Training-Free Prototype Calibration Method for Cross-Discipline Few-Shot Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2604.15678v1 Announce Type: new Abstract: Pretrained Vision-Language Models (VLMs) like CLIP show promise in continual learning, but existing Few-Shot Class-Incremental Learning (FSCIL) methods

IA-CLAHE: Image-Adaptive Clip Limit Estimation for CLAHE

Model ReleasesDGX agent

arXiv:2604.16010v1 Announce Type: new Abstract: This paper proposes image-adaptive contrast limited adaptive histogram equalization (IA-CLAHE). Conventional CLAHE is widely used to boost the performan

Information Router for Mitigating Modality Dominance in Vision-Language Models

ApplicationsDGX agent

arXiv:2604.16264v1 Announce Type: new Abstract: Vision Language models (VLMs) have demonstrated strong performance across a wide range of benchmarks, yet they often suffer from modality dominance, whe

InstructTable: Improving Table Structure Recognition Through Instructions

Model ReleasesDGX agent

arXiv:2604.02880v2 Announce Type: replace Abstract: Table structure recognition (TSR) holds widespread practical importance by parsing tabular images into structured representations, yet encounters si

Learning Affine-Equivariant Proximal Operators

TutorialsDGX agent

arXiv:2604.15556v1 Announce Type: cross Abstract: Proximal operators are fundamental across many applications in signal processing and machine learning, including solving ill-posed inverse problems. R

Learning to Look before Learning to Like: Incorporating Human Visual Cognition into Aesthetic Quality Assessment

SafetyDGX agent

arXiv:2604.15853v1 Announce Type: new Abstract: Automated Aesthetic Quality Assessment (AQA) treats images primarily as static pixel vectors, aligning predictions with human-rating scores largely thro

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens

ResearchDGX agent

arXiv:2602.12370v2 Announce Type: replace Abstract: Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of mode

LP^{2}DH: A Locality-Preserving Pixel-Difference Hashing Framework for Dynamic Texture Recognition

ResearchDGX agent

arXiv:2604.15707v1 Announce Type: new Abstract: Spatiotemporal Local Binary Pattern (STLBP) is a widely used dynamic texture descriptor, but it suffers from extremely high dimensionality. To tackle th

M3R: Localized Rainfall Nowcasting with Meteorology-Informed MultiModal Attention

SafetyDGX agent

arXiv:2604.15377v1 Announce Type: cross Abstract: Accurate and timely rainfall nowcasting is crucial for disaster mitigation and water resource management. Despite recent advances in deep learning, pr

Making Image Editing Easier via Adaptive Task Reformulation with Agentic Executions

Model ReleasesDGX agent

arXiv:2604.15917v1 Announce Type: new Abstract: Instruction guided image editing has advanced substantially with recent generative models, yet it still fails to produce reliable results across many se

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

ResearchDGX agent

arXiv:2510.09065v2 Announce Type: replace-cross Abstract: We introduce MMAudioSep, a generative model for video/text-queried sound separation that is founded on a pretrained video-to-audio model. By l

MMGait: Towards Multi-Modal Gait Recognition

Model ReleasesDGX agent

arXiv:2604.15979v1 Announce Type: new Abstract: Gait recognition has emerged as a powerful biometric technique for identifying individuals at a distance without requiring user cooperation. Most existi

Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions

ApplicationsDGX agent

arXiv:2604.16135v1 Announce Type: new Abstract: Recent advances in generative motion synthesis have enabled the production of realistic human motions from diverse input modalities. However, synthesizi

Neural Gabor Splatting: Enhanced Gaussian Splatting with Neural Gabor for High-frequency Surface Reconstruction

ResearchDGX agent

arXiv:2604.15941v1 Announce Type: new Abstract: Recent years have witnessed the rapid emergence of 3D Gaussian splatting (3DGS) as a powerful approach for 3D reconstruction and novel view synthesis. I

neuralCAD-Edit: An Expert Benchmark for Multimodal-Instructed 3D CAD Model Editing

Model ReleasesDGX agent

arXiv:2604.16170v1 Announce Type: new Abstract: We introduce neuralCAD-Edit, the first benchmark for editing 3D CAD models collected from expert CAD engineers. Instead of text conditioning as in prior

P3T: Prototypical Point-level Prompt Tuning with Enhanced Generalization for 3D Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.15703v1 Announce Type: new Abstract: With the rise of pre-trained models in the 3D point cloud domain for a wide range of real-world applications, adapting them to downstream tasks has beco

PILOT: A Promptable Interleaved Layout-aware OCR Transformer

Model ReleasesDGX agent

arXiv:2504.03621v2 Announce Type: replace Abstract: Classical OCR pipelines decompose document reading into detection, segmentation, and recognition stages, which makes them sensitive to localization

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

Model ReleasesDGX agent

arXiv:2604.15670v1 Announce Type: new Abstract: Reasoning segmentation has recently expanded from ground-level scenes to remote-sensing imagery, yet UAV data poses distinct challenges, including obliq

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding

SafetyDGX agent

arXiv:2604.15770v1 Announce Type: new Abstract: Accurate open-vocabulary 3D scene understanding requires semantic representations that are both language-aligned and spatially precise at the pixel leve

PolarMAE: Efficient Fetal Ultrasound Pre-training via Semantic Screening and Polar-Guided Masking

Local AiDGX agent

arXiv:2604.15893v1 Announce Type: new Abstract: Intelligent fetal ultrasound (US) interpretation is crucial for prenatal diagnosis, but high annotation costs and operator-induced variance make unsuper

Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation

ApplicationsDGX agent

arXiv:2604.16108v1 Announce Type: new Abstract: Speech-Driven Facial Animation (SDFA) has gained significant attention due to its applications in movies, video games, and virtual reality. However, mos

Proper Body Landmark Subset Enables More Accurate and 5X Faster Recognition of Isolated Signs in LIBRAS

ResearchDGX agent

arXiv:2510.24887v4 Announce Type: replace Abstract: This paper examines the feasibility of utilizing lightweight body landmark detection for recognizing isolated signs in Brazilian Sign Language (LIBR

ProtoTTA: Prototype-Guided Test-Time Adaptation

ApplicationsDGX agent

arXiv:2604.15494v1 Announce Type: cross Abstract: Deep networks that rely on prototypes-interpretable representations that can be related to the model input-have gained significant attention for balan

Ranking XAI Methods for Head and Neck Cancer Outcome Prediction

ResearchDGX agent

arXiv:2604.16034v1 Announce Type: new Abstract: For head and neck cancer (HNC) patients, prognostic outcome prediction can support personalized treatment strategy selection. Improving prediction perfo

Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence

Model ReleasesDGX agent

arXiv:2603.13091v2 Announce Type: replace Abstract: The growing interest in embodied agents increases the demand for spatiotemporal video understanding, yet existing benchmarks largely emphasize extra

Repurposing 3D Generative Model for Autoregressive Layout Generation

Model ReleasesDGX agent

arXiv:2604.16299v1 Announce Type: new Abstract: We introduce LaviGen, a framework that repurposes 3D generative models for 3D layout generation. Unlike previous methods that infer object layouts from

Saturation-Aware Space-Variant Blind Image Deblurring

ApplicationsDGX agent

arXiv:2604.16200v1 Announce Type: new Abstract: This paper presents a novel saturation aware space variant blind image deblurring framework designed to address challenges posed by saturated pixels in

Scalable spatial point process models for forensic footwear analysis

ResearchDGX agent

arXiv:2602.07006v2 Announce Type: replace Abstract: Shoe print evidence recovered from crime scenes plays a key role in forensic investigations. By examining shoe prints, investigators can determine d

Scalable Unseen Objects 6-DoF Absolute Pose Estimation with Robotic Integration

SafetyDGX agent

arXiv:2503.05578v4 Announce Type: replace Abstract: Pose estimation-guided unseen object 6-DoF robotic manipulation is a key task in robotics. However, the scalability of current pose estimation metho

Self-Supervised Angular Deblurring in Photoacoustic Reconstruction via Noisier2Inverse

ResearchDGX agent

arXiv:2604.15681v1 Announce Type: new Abstract: Photoacoustic tomography (PAT) is an emerging imaging modality that combines the complementary strengths of optical contrast and ultrasonic resolution.

SENSE: Stereo OpEN Vocabulary SEmantic Segmentation

AgentsDGX agent

arXiv:2604.15946v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation enables models to segment objects or image regions beyond fixed class sets, offering flexibility in dynamic enviro

Splats in Splats++: Robust and Generalizable 3D Gaussian Splatting Steganography

ResearchDGX agent

arXiv:2604.15862v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently redefined the paradigm of 3D reconstruction, striking an unprecedented balance between visual fidelity and com

SPLIT: Self-supervised Partitioning for Learned Inversion in Nonlinear Tomography

ResearchDGX agent

arXiv:2604.15651v1 Announce Type: new Abstract: Machine learning has achieved impressive performance in tomographic reconstruction, but supervised training requires paired measurements and ground-trut

SSFT: A Lightweight Spectral-Spatial Fusion Transformer for Generic Hyperspectral Classification

Model ReleasesDGX agent

arXiv:2604.15828v1 Announce Type: new Abstract: Hyperspectral imaging enables fine-grained recognition of materials by capturing rich spectral signatures, but learning robust classifiers is challengin

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

ResearchDGX agent

arXiv:2602.05638v3 Announce Type: replace Abstract: While foundation models have advanced surgical video analysis, current approaches rely predominantly on pixel-level reconstruction objectives that w

TableSeq: Unified Generation of Structure, Content, and Layout

Local AiDGX agent

arXiv:2604.16070v1 Announce Type: new Abstract: We present TableSeq, an image-only, end-to-end framework for joint table structure recognition, content recognition, and cell localization. The model fo

The Amazing Stability of Flow Matching

ResearchDGX agent

arXiv:2604.16079v1 Announce Type: new Abstract: The success of deep generative models in generating high-quality and diverse samples is often attributed to particular architectures and large training

Topology-Driven Fusion of nnU-Net and MedNeXt for Accurate Brain Tumor Segmentation on Sub-Saharan Africa Dataset

ResearchDGX agent

arXiv:2604.15964v1 Announce Type: cross Abstract: Accurate automatic brain tumor segmentation in Low and Middle-Income (LMIC) countries is challenging due to the lack of defined national imaging proto

Towards In-Context Tone Style Transfer with A Large-Scale Triplet Dataset

ResearchDGX agent

arXiv:2604.16114v1 Announce Type: new Abstract: Tone style transfer for photo retouching aims to adapt the stylistic tone of the reference image to a given content image. However, the lack of high-qua

Towards Realistic Open-Vocabulary Remote Sensing Segmentation: Benchmark and Baseline

Model ReleasesDGX agent

arXiv:2604.15652v1 Announce Type: new Abstract: Open-vocabulary remote sensing image segmentation (OVRSIS) remains underexplored due to fragmented datasets, limited training diversity, and the lack of

Training Flow Matching: The Role of Weighting and Parameterization

ResearchDGX agent

arXiv:2603.06454v2 Announce Type: replace Abstract: We study the training objectives of denoising-based generative models, with a particular focus on loss weighting and output parameterization, includ

Two-Stage Framework for Efficient UAV-Based Wildfire Video Analysis with Adaptive Compression and Fire Source Detection

Local AiDGX agent

arXiv:2508.16739v2 Announce Type: replace Abstract: Unmanned Aerial Vehicles (UAVs) have become increasingly important in disaster emergency response by facilitating aerial video analysis. Due to the

TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models

Model ReleasesDGX agent

arXiv:2604.15967v1 Announce Type: cross Abstract: Despite the remarkable synthesis capabilities of text-to-image (T2I) models, safeguarding them against content violations remains a persistent challen

UA-Net: Uncertainty-Aware Network for TRISO Image Semantic Segmentation

ResearchDGX agent

arXiv:2604.15542v1 Announce Type: new Abstract: Tristructural isotropic (TRISO)-coated particle fuels undergo dimensional changes and chemical reactions during high-temperature neutron irradiation. Po

VeRVE: Versatile Retrieval for Videos via Unified Embeddings

Local AiDGX agent

arXiv:2601.12193v3 Announce Type: replace Abstract: Modern video retrieval systems are expected to handle diverse tasks ranging from corpus-level retrieval, fine-grained moment localization to flexibl

Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools

ResearchDGX agent

arXiv:2510.08480v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable potential in bridging visual and textual reasoning, yet their reliance on text

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions

Model ReleasesDGX agent

arXiv:2604.15823v1 Announce Type: new Abstract: Embodied robotic agents often perceive movies through an egocentric screen-view interface rather than native cinematic footage, introducing domain shift

Weak-to-Strong Knowledge Distillation Accelerates Visual Learning

ResearchDGX agent

arXiv:2604.15451v1 Announce Type: new Abstract: Large-scale visual learning is increasingly limited by training cost. Existing knowledge distillation methods transfer from a stronger teacher to a weak

When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.27759v3 Announce Type: replace Abstract: Visual-Language Models (VLMs) have demonstrated exceptional cross-modal understanding across various tasks, including zero-shot classification, imag

Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization

Local AiDGX agent

arXiv:2604.16248v1 Announce Type: new Abstract: Image geolocalization has traditionally been addressed through retrieval-based place recognition or geometry-based visual localization pipelines. Recent

Winner of CVPR2026 NTIRE Challenge on Image Shadow Removal: Semantic and Geometric Guidance for Shadow Removal via Cascaded Refinement

ResearchDGX agent

arXiv:2604.16177v1 Announce Type: new Abstract: We present a three-stage progressive shadow-removal pipeline for the CVPR2026 NTIRE WSRD+ challenge. Built on OmniSR, our method treats deshadowing as i

17 Apr 2026

3AM: 3egment Anything with Geometric Consistency in Videos

ResearchDGX agent

arXiv:2601.08831v5 Announce Type: replace Abstract: Video object segmentation methods like SAM2 achieve strong performance through memory-based architectures but struggle under large viewpoint changes

3D Conditional Image Synthesis of Left Atrial LGE MRI from Composite Semantic Masks

ResearchDGX agent

arXiv:2601.04588v2 Announce Type: replace Abstract: Segmentation of the left atrial (LA) wall and endocardium from late gadolinium-enhanced (LGE) MRI is essential for quantifying atrial fibrosis in pa

A deep learning framework for glomeruli segmentation with boundary attention

ResearchDGX agent

arXiv:2604.14263v1 Announce Type: cross Abstract: Accurate detection and segmentation of glomeruli in kidney tissue are essential for diagnostic applications. Traditional deep learning methods primari

AD4AD: Benchmarking Visual Anomaly Detection Models for Safer Autonomous Driving

Model ReleasesDGX agent

arXiv:2604.15291v1 Announce Type: new Abstract: The reliability of a machine vision system for autonomous driving depends heavily on its training data distribution. When a vehicle encounters significa

AgentIAD: Agentic Industrial Anomaly Detection via Adaptive Memory Augmentation

Model ReleasesDGX agent

arXiv:2512.13671v2 Announce Type: replace Abstract: Industrial anomaly detection (IAD) is challenging due to the subtle and highly localized nature of many defects, which single-pass vision--language

All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept Reproduction

TutorialsDGX agent

arXiv:2601.04567v2 Announce Type: replace Abstract: Harmful memes are ever-shifting in the Internet communities, which are difficult to analyze due to their type-shifting and temporal-evolving nature.

← Previous
1…187188189190191…209
Next →