AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
28 Apr 2026

ShapeUP: Scalable Image-Conditioned 3D Editing

ResearchDGX agent

arXiv:2602.05676v2 Announce Type: replace Abstract: Recent advancements in 3D foundation models have enabled the generation of high-fidelity assets, yet precise 3D manipulation remains a significant c

Shared-kernel Wavelet Neural Networks for Poisson Image Reconstruction

ResearchDGX agent

arXiv:2604.24000v1 Announce Type: cross Abstract: The Laplacian operator transforms the image into its Laplacian field, which usually is sparse and satisfies a stable distribution. On the other hand,

ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation

SafetyDGX agent

arXiv:2506.18493v2 Announce Type: replace Abstract: Customizing image generation remains a core challenge in controllable image synthesis. For single-concept generation, maintaining both identity pres


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs

TutorialsDGX agent

arXiv:2604.23996v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) has become a prevalent backbone for large vision-language models (VLMs), yet how modality-specific signals should guide expert

Smoothing Slot Attention Iterations and Recurrences

ResearchDGX agent

arXiv:2508.05417v3 Announce Type: replace Abstract: Slot Attention (SA) lies at the heart of mainstream Object-Centric Learning (OCL). Image features can be aggregated into object-level representation

SolarFCD: A Large-Scale Dataset and Benchmark for Solar Fault Classification in Photovoltaic Systems

Model ReleasesDGX agent

arXiv:2604.23662v1 Announce Type: new Abstract: The increasing global deployment of solar photovoltaic (PV) systems needs robust, scalable, and automated inspection technologies capable of detecting a

SPAGS: Sparse-View Articulated Object Reconstruction from Single State via Planar Gaussian Splatting

Model ReleasesDGX agent

arXiv:2511.17092v4 Announce Type: replace Abstract: Articulated objects are ubiquitous in daily environments, and their 3D reconstruction holds great significance across various fields. However, exist

Spatiotemporal Degradation-Aware 3D Gaussian Splatting for Realistic Underwater Scene Reconstruction

Model ReleasesDGX agent

arXiv:2604.23551v1 Announce Type: new Abstract: Reconstructing realistic underwater scenes from underwater video remains a meaningful yet challenging task in the multimedia domain. The inherent spatio

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels

TutorialsDGX agent

arXiv:2401.07669v2 Announce Type: replace Abstract: Adapting CLIP for videos has gained popularity due to its semantic and rich representation. While CLIP is a good starting point, it typically underg

STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning

ResearchDGX agent

arXiv:2604.23309v1 Announce Type: new Abstract: Remote sensing image change captioning (RSICC) aims to describe the difference between two remote sensing images. While recent methods have explored vid

Statistical Test for Diffusion-Based Anomaly Localization via Selective Inference

Local AiDGX agent

arXiv:2402.11789v5 Announce Type: replace-cross Abstract: Anomaly localization in images -- identifying regions that deviate from normal patterns -- is vital in applications such as medical diagnosis

StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space

TutorialsDGX agent

arXiv:2512.10959v2 Announce Type: replace Abstract: We introduce StereoSpace, a diffusion-based framework for monocular-to-stereo synthesis that models geometry purely through viewpoint conditioning,

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval

Model ReleasesDGX agent

arXiv:2601.20597v2 Announce Type: replace Abstract: Continual Text-to-Video Retrieval (CTVR) is a challenging multimodal continual learning setting, where models must incrementally learn new semantic

Text-Guided Multimodal Unified Industrial Anomaly Detection

SafetyDGX agent

arXiv:2604.22899v1 Announce Type: new Abstract: Industrial anomaly detection based on RGB-3D multimodal data has emerged as a mainstream paradigm for intelligent quality inspection. However, existing

TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering

Model ReleasesDGX agent

arXiv:2604.24459v1 Announce Type: new Abstract: Despite recent advances in text-to-image generation, models still struggle to accurately render prompt-specified text with correct spatial layout -- esp

TokenTrace: Multi-Concept Attribution through Watermarked Token Recovery

TutorialsDGX agent

arXiv:2602.19019v2 Announce Type: replace Abstract: Generative AI models pose a significant challenge to intellectual property (IP), as they can replicate unique artistic styles and concepts without a

TOL: Textual Localization with OpenStreetMap

Model ReleasesDGX agent

arXiv:2604.01644v2 Announce Type: replace Abstract: Natural language provides an intuitive way to express spatial intent in geospatial applications. While existing localization methods often rely on d

TopoHR: Hierarchical Centerline Representation for Cyclic Topology Reasoning in Driving Scenes with Point-to-Instance Relations

Model ReleasesDGX agent

arXiv:2604.24119v1 Announce Type: new Abstract: Topology reasoning is crucial for autonomous driving. Current methods primarily focus on instance-level learning for centerline detection, followed by a

Touchless Intraoperative Image Access System Based on Vision-Based Hand Tracking

ResearchDGX agent

arXiv:2604.24235v1 Announce Type: new Abstract: Touchless interaction with medical images is becoming increasingly important in the surgical field, where sterility and continuity of the operational wo

Toward Real-World Adoption of Portrait Relighting via Hybrid Domain Knowledge Fusion

ApplicationsDGX agent

arXiv:2604.23094v1 Announce Type: new Abstract: The real-world adoption of portrait relighting is hindered by dataset domain gaps, camera sensitivity, and computational costs. We address these challen

Towards Any-Quality Image Segmentation via Generative and Adaptive Latent Space Enhancement

SafetyDGX agent

arXiv:2601.02018v2 Announce Type: replace Abstract: Segment Anything Models (SAMs), known for their exceptional zero-shot segmentation performance, have garnered significant attention in the research

Towards Fair and Robust Volumetric CT Classification via KL-Regularised Group Distributionally Robust Optimisation

SafetyDGX agent

arXiv:2603.15941v2 Announce Type: replace Abstract: Automated diagnosis from chest computed tomography (CT) scans faces two persistent challenges in clinical deployment: distribution shift across acqu

Towards High-Fidelity CAD Generation via LLM-Driven Program Generation and Text-Based B-Rep Primitive Grounding

ApplicationsDGX agent

arXiv:2603.11831v2 Announce Type: replace Abstract: The field of Computer-Aided Design (CAD) generation has made significant progress in recent years. Existing methods typically fall into two separate

Transferable Physical-World Adversarial Patches Against Object Detection in Autonomous Driving

SafetyDGX agent

arXiv:2604.23105v1 Announce Type: new Abstract: Deep learning drives major advances in autonomous driving (AD), where object detectors are central to perception. However, adversarial attacks pose sign

Triple-Phase Sequential Fusion Network for Hepatobiliary Phase Liver MRI Synthesis

ResearchDGX agent

arXiv:2604.22904v1 Announce Type: cross Abstract: Gadoxetate disodium-enhanced MRI is essential for the detection and characterization of hepatocellular carcinoma. However, acquisition of the hepatobi

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

ResearchDGX agent

arXiv:2604.24763v1 Announce Type: new Abstract: Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creatin

U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and Registration

Local AiDGX agent

arXiv:2507.04503v2 Announce Type: replace Abstract: Accurate localization using visual information is a critical yet challenging task, especially in urban environments where nearby buildings and const

Understanding Representation Gaps Across Scales in Tropical Tree Species Classification from Drone Imagery

SafetyDGX agent

arXiv:2604.23019v1 Announce Type: new Abstract: Accurate classification of tropical tree species from unoccupied aerial vehicle (UAV) imagery remains challenging due to high species diversity and stro

Unified Multi-Foundation-Model Slide Representation for Pan-Cancer Recognition and Text-Guided Tumor Localization

Local AiDGX agent

arXiv:2604.22846v1 Announce Type: new Abstract: The expanding ecosystem of pathology foundation models has produced powerful but fragmented tile-level representations, limiting their use in clinical t

Urban Flood Observations (UFO): A hand-labeled training and validation dataset of post-flood inundation

ResearchDGX agent

arXiv:2604.23066v1 Announce Type: new Abstract: Urban flooding affects lives and infrastructure worldwide. Mapping inundation in complex urban environments from satellite imagery remains challenging d

V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think

SafetyDGX agent

arXiv:2604.23380v1 Announce Type: cross Abstract: Aligning denoising generative models with human preferences or verifiable rewards remains a key challenge. While policy-gradient online reinforcement

VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models

Model ReleasesDGX agent

arXiv:2510.08618v2 Announce Type: replace-cross Abstract: Omni-modal large language models (OLLMs) offer a promising end-to-end solution for slide-enhanced speech recognition due to their inherent mul

VDLF-Net: Variational Feature Fusion for Adaptive and Few-Shot Visual Learning

ResearchDGX agent

arXiv:2604.23641v1 Announce Type: new Abstract: This paper introduces VDLF-Net, which attaches a compact VAE to a multi-scale CNN backbone. Latent vectors and softmax-gate support the backbone feature

Vision-Based Lane Following and Traffic Sign Recognition for Resource-Constrained Autonomous Vehicles

Local AiDGX agent

arXiv:2604.22872v1 Announce Type: new Abstract: Autonomous vehicles (AVs) rely on real-time perception systems to understand road environments and ensure safe navigation. However, implementing reliabl

VitaminP: cross-modal learning enables whole-cell segmentation from routine histology

ResearchDGX agent

arXiv:2604.23799v1 Announce Type: new Abstract: Accurate whole-cell and nuclear segmentation is essential for precision pathology and spatial omics, yet routine hematoxylin and eosin (H&E) staining pr

Voxify3D: Pixel Art Meets Volumetric Rendering

SafetyDGX agent

arXiv:2512.07834v2 Announce Type: replace Abstract: Voxel art is a distinctive stylization widely used in games and digital media, yet automated generation from 3D meshes remains challenging due to co

Weakly Supervised Multicenter Nancy Index Scoring in Ulcerative Colitis Using Foundation Models

ResearchDGX agent

arXiv:2604.23706v1 Announce Type: new Abstract: Histologic assessment of ulcerative colitis (UC) activity is an important endpoint in clinical trials and routine care, but manual grading with indices

WebSerial Vision Training for Microcontrollers: A Browser-Based Companion to On-Device CNN Training

Model ReleasesDGX agent

arXiv:2604.22834v1 Announce Type: new Abstract: This paper presents webmcu-vision-web, a single-file, zero-install browser application for end-to-end TinyML vision model training and deployment on the

WildLIFT: Lifting monocular drone video to 3D for species-agnostic wildlife monitoring

ResearchDGX agent

arXiv:2604.24718v1 Announce Type: new Abstract: Monocular RGB cameras mounted on drones are widely used for wildlife monitoring, yet most analytical pipelines remain confined to two-dimensional image

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation

SafetyDGX agent

arXiv:2604.24764v1 Announce Type: new Abstract: Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods atte

Z^2-Sampling: Zero-Cost Zigzag Trajectories for Semantic Alignment in Diffusion Models

SafetyDGX agent

arXiv:2604.23536v1 Announce Type: new Abstract: Diffusion models have achieved unprecedented success in text-aligned generation, largely driven by Classifier-Free Guidance (CFG). However, standard CFG

Zero-to-CAD: Agentic Synthesis of Interpretable CAD Programs at Million-Scale Without Real Data

Model ReleasesDGX agent

arXiv:2604.24479v1 Announce Type: new Abstract: Computer-Aided Design (CAD) models are defined by their construction history: a parametric recipe that encodes design intent. However, existing large-sc

ZID-Net: Zero-Inference Diffusion Prior Decoupling Network for Single Image Dehazing

TutorialsDGX agent

arXiv:2604.23709v1 Announce Type: new Abstract: Single image dehazing is often constrained by a trade-off between restoration quality and computational efficiency. While efficient, CNN networks strugg

27 Apr 2026

3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale

Local AiDGX agent

arXiv:2511.13211v2 Announce Type: replace Abstract: Despite recent advancements in 3D-text cross-modal alignment, existing state-of-the-art methods still struggle to align fine-grained textual semanti

A Non-Invasive Alternative to RFID: Self-Sufficient 3D Identification of Group-Housed Livestock

AgentsDGX agent

arXiv:2604.22657v1 Announce Type: new Abstract: Accurate identification of individual farm animals in group-housed environments is a cornerstone of precision livestock management. However, current ind

Adapting MLLMs for Nuanced Video Retrieval

ResearchDGX agent

arXiv:2512.13511v2 Announce Type: replace Abstract: Our objective is to build an embedding model that captures the nuanced relationship between a search query and candidate videos. We cover three aspe

All Eyes on the Workflow: Automated and Efficient Event Discovery from Video Streams

ResearchDGX agent

arXiv:2604.22476v1 Announce Type: new Abstract: Disciplines such as business process management and process mining aid organizations by discovering insights about processes on the basis of recorded ev

Altitude-Adaptive Vision-Only Geo-Localization for UAVs in GPS-Denied Environments

ResearchDGX agent

arXiv:2602.23872v2 Announce Type: replace Abstract: To address the scale mismatch caused by large altitude variations in UAV visual place recognition, we propose a monocular vision-only altitude-adapt

Anatomy-Aware Unsupervised Detection and Localization of Retinal Abnormalities in Optical Coherence Tomography

Model ReleasesDGX agent

arXiv:2604.22139v1 Announce Type: new Abstract: Reliable automated analysis of Optical Coherence Tomography (OCT) imaging is crucial for diagnosing retinal disorders but faces a critical barrier: the

ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild

Model ReleasesDGX agent

arXiv:2604.22202v1 Announce Type: new Abstract: Symmetry detection is a fundamental problem in computer vision, and symmetries serve as powerful priors for downstream tasks. However, existing learning

Are Natural-Domain Foundation Models Effective for Accelerated Cardiac MRI Reconstruction?

TutorialsDGX agent

arXiv:2604.22557v1 Announce Type: cross Abstract: The emergence of large-scale pretrained foundation models has transformed computer vision, enabling strong performance across diverse downstream tasks

Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?

ResearchDGX agent

arXiv:2510.10254v2 Announce Type: replace Abstract: Recent advances in large generative models have shown that simple autoregressive formulations, when scaled appropriately, can exhibit strong zero-sh

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings

SafetyDGX agent

arXiv:2604.22280v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have emerged as a promising foundation for universal multimodal embeddings. Recent studies have shown that reas

Breaking Watermarks in the Frequency Domain: A Modulated Diffusion Attack Framework

ResearchDGX agent

arXiv:2604.22220v1 Announce Type: new Abstract: Digital image watermarking has advanced rapidly for copyright protection of generative AI, yet the comparatively limited progress in watermark attack te

CAGE-SGG: Counterfactual Active Graph Evidence for Open-Vocabulary Scene Graph Generation

ResearchDGX agent

arXiv:2604.22274v1 Announce Type: new Abstract: Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible and fine-grained relation phrases beyond a fixed predicate voc

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution

Model ReleasesDGX agent

arXiv:2604.22192v1 Announce Type: new Abstract: Chart-to-code generation demands strict visual precision and syntactic correctness from Vision-Language Models (VLMs). However, existing approaches are

Conditional Diffusion Posterior Alignment for Sparse-View CT Reconstruction

SafetyDGX agent

arXiv:2604.21960v1 Announce Type: cross Abstract: Computed Tomography (CT) is a widely used imaging modality in medical and industrial applications. To limit radiation exposure and measurement time, t

Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples

ApplicationsDGX agent

arXiv:2604.22477v1 Announce Type: new Abstract: Neuron labeling assigns textual descriptions to internal units of deep networks. Existing approaches typically rely on highly activating examples, often

Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation

ResearchDGX agent

arXiv:2510.19592v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) demonstrate strong video understanding by attending to visual tokens relevant to textual queries. To direct

Depth-Aware Rover: A Study of Edge AI and Monocular Vision for Real-World Implementation

Local AiDGX agent

arXiv:2604.22331v1 Announce Type: new Abstract: This study analyses simulated and real-world implementations of depth-aware rover navigation, highlighting the transition from stereo vision to monocula

← Previous
1…171172173174175…209
Next →