AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

ShapeUP: Scalable Image-Conditioned 3D Editing

DGX agent

arXiv:2602.05676v2 Announce Type: replace Abstract: Recent advancements in 3D foundation models have enabled the generation of high-fidelity assets, yet precise 3D manipulation remains a significant c

researcharxiv-cs-cv
28 Apr 2026
Research

Shared-kernel Wavelet Neural Networks for Poisson Image Reconstruction

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2604.24000v1 Announce Type: cross Abstract: The Laplacian operator transforms the image into its Laplacian field, which usually is sparse and satisfies a stable distribution. On the other hand,

researcharxiv-cs-cv
28 Apr 2026
Safety

ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation

DGX agent

arXiv:2506.18493v2 Announce Type: replace Abstract: Customizing image generation remains a core challenge in controllable image synthesis. For single-concept generation, maintaining both identity pres

safetyarxiv-cs-cv
28 Apr 2026
Tutorials

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs

DGX agent

arXiv:2604.23996v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) has become a prevalent backbone for large vision-language models (VLMs), yet how modality-specific signals should guide expert

tutorialsarxiv-cs-cv
28 Apr 2026
Research

Smoothing Slot Attention Iterations and Recurrences

DGX agent

arXiv:2508.05417v3 Announce Type: replace Abstract: Slot Attention (SA) lies at the heart of mainstream Object-Centric Learning (OCL). Image features can be aggregated into object-level representation

researcharxiv-cs-cv
28 Apr 2026
Model Releases

SolarFCD: A Large-Scale Dataset and Benchmark for Solar Fault Classification in Photovoltaic Systems

DGX agent

arXiv:2604.23662v1 Announce Type: new Abstract: The increasing global deployment of solar photovoltaic (PV) systems needs robust, scalable, and automated inspection technologies capable of detecting a

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

SPAGS: Sparse-View Articulated Object Reconstruction from Single State via Planar Gaussian Splatting

DGX agent

arXiv:2511.17092v4 Announce Type: replace Abstract: Articulated objects are ubiquitous in daily environments, and their 3D reconstruction holds great significance across various fields. However, exist

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

Spatiotemporal Degradation-Aware 3D Gaussian Splatting for Realistic Underwater Scene Reconstruction

DGX agent

arXiv:2604.23551v1 Announce Type: new Abstract: Reconstructing realistic underwater scenes from underwater video remains a meaningful yet challenging task in the multimedia domain. The inherent spatio

model-releasesarxiv-cs-cv
28 Apr 2026
Tutorials

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels

DGX agent

arXiv:2401.07669v2 Announce Type: replace Abstract: Adapting CLIP for videos has gained popularity due to its semantic and rich representation. While CLIP is a good starting point, it typically underg

tutorialsarxiv-cs-cv
28 Apr 2026
Research

STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning

DGX agent

arXiv:2604.23309v1 Announce Type: new Abstract: Remote sensing image change captioning (RSICC) aims to describe the difference between two remote sensing images. While recent methods have explored vid

researcharxiv-cs-cv
28 Apr 2026
Local Ai

Statistical Test for Diffusion-Based Anomaly Localization via Selective Inference

DGX agent

arXiv:2402.11789v5 Announce Type: replace-cross Abstract: Anomaly localization in images -- identifying regions that deviate from normal patterns -- is vital in applications such as medical diagnosis

local-aiarxiv-cs-cv
28 Apr 2026
Tutorials

StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space

DGX agent

arXiv:2512.10959v2 Announce Type: replace Abstract: We introduce StereoSpace, a diffusion-based framework for monocular-to-stereo synthesis that models geometry purely through viewpoint conditioning,

tutorialsarxiv-cs-cv
28 Apr 2026
Model Releases

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval

DGX agent

arXiv:2601.20597v2 Announce Type: replace Abstract: Continual Text-to-Video Retrieval (CTVR) is a challenging multimodal continual learning setting, where models must incrementally learn new semantic

model-releasesarxiv-cs-cv
28 Apr 2026
Safety

Text-Guided Multimodal Unified Industrial Anomaly Detection

DGX agent

arXiv:2604.22899v1 Announce Type: new Abstract: Industrial anomaly detection based on RGB-3D multimodal data has emerged as a mainstream paradigm for intelligent quality inspection. However, existing

safetyarxiv-cs-cv
28 Apr 2026
Model Releases

TextGround4M: A Prompt-Aligned Dataset for Layout-Aware Text Rendering

DGX agent

arXiv:2604.24459v1 Announce Type: new Abstract: Despite recent advances in text-to-image generation, models still struggle to accurately render prompt-specified text with correct spatial layout -- esp

model-releasesarxiv-cs-cv
28 Apr 2026
Tutorials

TokenTrace: Multi-Concept Attribution through Watermarked Token Recovery

DGX agent

arXiv:2602.19019v2 Announce Type: replace Abstract: Generative AI models pose a significant challenge to intellectual property (IP), as they can replicate unique artistic styles and concepts without a

tutorialsarxiv-cs-cv
28 Apr 2026
Model Releases

TOL: Textual Localization with OpenStreetMap

DGX agent

arXiv:2604.01644v2 Announce Type: replace Abstract: Natural language provides an intuitive way to express spatial intent in geospatial applications. While existing localization methods often rely on d

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

TopoHR: Hierarchical Centerline Representation for Cyclic Topology Reasoning in Driving Scenes with Point-to-Instance Relations

DGX agent

arXiv:2604.24119v1 Announce Type: new Abstract: Topology reasoning is crucial for autonomous driving. Current methods primarily focus on instance-level learning for centerline detection, followed by a

model-releasesarxiv-cs-cv
28 Apr 2026
Research

Touchless Intraoperative Image Access System Based on Vision-Based Hand Tracking

DGX agent

arXiv:2604.24235v1 Announce Type: new Abstract: Touchless interaction with medical images is becoming increasingly important in the surgical field, where sterility and continuity of the operational wo

researcharxiv-cs-cv
28 Apr 2026
Applications

Toward Real-World Adoption of Portrait Relighting via Hybrid Domain Knowledge Fusion

DGX agent

arXiv:2604.23094v1 Announce Type: new Abstract: The real-world adoption of portrait relighting is hindered by dataset domain gaps, camera sensitivity, and computational costs. We address these challen

applicationsarxiv-cs-cv
28 Apr 2026
Safety

Towards Any-Quality Image Segmentation via Generative and Adaptive Latent Space Enhancement

DGX agent

arXiv:2601.02018v2 Announce Type: replace Abstract: Segment Anything Models (SAMs), known for their exceptional zero-shot segmentation performance, have garnered significant attention in the research

safetyarxiv-cs-cv
28 Apr 2026
Safety

Towards Fair and Robust Volumetric CT Classification via KL-Regularised Group Distributionally Robust Optimisation

DGX agent

arXiv:2603.15941v2 Announce Type: replace Abstract: Automated diagnosis from chest computed tomography (CT) scans faces two persistent challenges in clinical deployment: distribution shift across acqu

safetyarxiv-cs-cv
28 Apr 2026
Applications

Towards High-Fidelity CAD Generation via LLM-Driven Program Generation and Text-Based B-Rep Primitive Grounding

DGX agent

arXiv:2603.11831v2 Announce Type: replace Abstract: The field of Computer-Aided Design (CAD) generation has made significant progress in recent years. Existing methods typically fall into two separate

applicationsarxiv-cs-cv
28 Apr 2026
Safety

Transferable Physical-World Adversarial Patches Against Object Detection in Autonomous Driving

DGX agent

arXiv:2604.23105v1 Announce Type: new Abstract: Deep learning drives major advances in autonomous driving (AD), where object detectors are central to perception. However, adversarial attacks pose sign

safetyarxiv-cs-cv
28 Apr 2026
Research

Triple-Phase Sequential Fusion Network for Hepatobiliary Phase Liver MRI Synthesis

DGX agent

arXiv:2604.22904v1 Announce Type: cross Abstract: Gadoxetate disodium-enhanced MRI is essential for the detection and characterization of hepatocellular carcinoma. However, acquisition of the hepatobi

researcharxiv-cs-cv
28 Apr 2026
Research

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

DGX agent

arXiv:2604.24763v1 Announce Type: new Abstract: Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creatin

researcharxiv-cs-cv
28 Apr 2026
Local Ai

U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and Registration

DGX agent

arXiv:2507.04503v2 Announce Type: replace Abstract: Accurate localization using visual information is a critical yet challenging task, especially in urban environments where nearby buildings and const

local-aiarxiv-cs-cv
28 Apr 2026
Safety

Understanding Representation Gaps Across Scales in Tropical Tree Species Classification from Drone Imagery

DGX agent

arXiv:2604.23019v1 Announce Type: new Abstract: Accurate classification of tropical tree species from unoccupied aerial vehicle (UAV) imagery remains challenging due to high species diversity and stro

safetyarxiv-cs-cv
28 Apr 2026
Local Ai

Unified Multi-Foundation-Model Slide Representation for Pan-Cancer Recognition and Text-Guided Tumor Localization

DGX agent

arXiv:2604.22846v1 Announce Type: new Abstract: The expanding ecosystem of pathology foundation models has produced powerful but fragmented tile-level representations, limiting their use in clinical t

local-aiarxiv-cs-cv
28 Apr 2026
Research

Urban Flood Observations (UFO): A hand-labeled training and validation dataset of post-flood inundation

DGX agent

arXiv:2604.23066v1 Announce Type: new Abstract: Urban flooding affects lives and infrastructure worldwide. Mapping inundation in complex urban environments from satellite imagery remains challenging d

researcharxiv-cs-cv
28 Apr 2026
Safety

V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think

DGX agent

arXiv:2604.23380v1 Announce Type: cross Abstract: Aligning denoising generative models with human preferences or verifiable rewards remains a key challenge. While policy-gradient online reinforcement

safetyarxiv-cs-cv
28 Apr 2026
Model Releases

VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models

DGX agent

arXiv:2510.08618v2 Announce Type: replace-cross Abstract: Omni-modal large language models (OLLMs) offer a promising end-to-end solution for slide-enhanced speech recognition due to their inherent mul

model-releasesarxiv-cs-cv
28 Apr 2026
Research

VDLF-Net: Variational Feature Fusion for Adaptive and Few-Shot Visual Learning

DGX agent

arXiv:2604.23641v1 Announce Type: new Abstract: This paper introduces VDLF-Net, which attaches a compact VAE to a multi-scale CNN backbone. Latent vectors and softmax-gate support the backbone feature

researcharxiv-cs-cv
28 Apr 2026
Local Ai

Vision-Based Lane Following and Traffic Sign Recognition for Resource-Constrained Autonomous Vehicles

DGX agent

arXiv:2604.22872v1 Announce Type: new Abstract: Autonomous vehicles (AVs) rely on real-time perception systems to understand road environments and ensure safe navigation. However, implementing reliabl

local-aiarxiv-cs-cv
28 Apr 2026
Research

VitaminP: cross-modal learning enables whole-cell segmentation from routine histology

DGX agent

arXiv:2604.23799v1 Announce Type: new Abstract: Accurate whole-cell and nuclear segmentation is essential for precision pathology and spatial omics, yet routine hematoxylin and eosin (H&E) staining pr

researcharxiv-cs-cv
28 Apr 2026
Safety

Voxify3D: Pixel Art Meets Volumetric Rendering

DGX agent

arXiv:2512.07834v2 Announce Type: replace Abstract: Voxel art is a distinctive stylization widely used in games and digital media, yet automated generation from 3D meshes remains challenging due to co

safetyarxiv-cs-cv
28 Apr 2026
Research

Weakly Supervised Multicenter Nancy Index Scoring in Ulcerative Colitis Using Foundation Models

DGX agent

arXiv:2604.23706v1 Announce Type: new Abstract: Histologic assessment of ulcerative colitis (UC) activity is an important endpoint in clinical trials and routine care, but manual grading with indices

researcharxiv-cs-cv
28 Apr 2026
Model Releases

WebSerial Vision Training for Microcontrollers: A Browser-Based Companion to On-Device CNN Training

DGX agent

arXiv:2604.22834v1 Announce Type: new Abstract: This paper presents webmcu-vision-web, a single-file, zero-install browser application for end-to-end TinyML vision model training and deployment on the

model-releasesarxiv-cs-cv
28 Apr 2026
Research

WildLIFT: Lifting monocular drone video to 3D for species-agnostic wildlife monitoring

DGX agent

arXiv:2604.24718v1 Announce Type: new Abstract: Monocular RGB cameras mounted on drones are widely used for wildlife monitoring, yet most analytical pipelines remain confined to two-dimensional image

researcharxiv-cs-cv
28 Apr 2026
Safety

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation

DGX agent

arXiv:2604.24764v1 Announce Type: new Abstract: Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods atte

safetyarxiv-cs-cv
28 Apr 2026
Safety

Z^2-Sampling: Zero-Cost Zigzag Trajectories for Semantic Alignment in Diffusion Models

DGX agent

arXiv:2604.23536v1 Announce Type: new Abstract: Diffusion models have achieved unprecedented success in text-aligned generation, largely driven by Classifier-Free Guidance (CFG). However, standard CFG

safetyarxiv-cs-cv
28 Apr 2026
Model Releases

Zero-to-CAD: Agentic Synthesis of Interpretable CAD Programs at Million-Scale Without Real Data

DGX agent

arXiv:2604.24479v1 Announce Type: new Abstract: Computer-Aided Design (CAD) models are defined by their construction history: a parametric recipe that encodes design intent. However, existing large-sc

model-releasesarxiv-cs-cv
28 Apr 2026
Tutorials

ZID-Net: Zero-Inference Diffusion Prior Decoupling Network for Single Image Dehazing

DGX agent

arXiv:2604.23709v1 Announce Type: new Abstract: Single image dehazing is often constrained by a trade-off between restoration quality and computational efficiency. While efficient, CNN networks strugg

tutorialsarxiv-cs-cv
28 Apr 2026
Local Ai

3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale

DGX agent

arXiv:2511.13211v2 Announce Type: replace Abstract: Despite recent advancements in 3D-text cross-modal alignment, existing state-of-the-art methods still struggle to align fine-grained textual semanti

local-aiarxiv-cs-cv
27 Apr 2026
Agents

A Non-Invasive Alternative to RFID: Self-Sufficient 3D Identification of Group-Housed Livestock

DGX agent

arXiv:2604.22657v1 Announce Type: new Abstract: Accurate identification of individual farm animals in group-housed environments is a cornerstone of precision livestock management. However, current ind

agentsarxiv-cs-cv
27 Apr 2026
Research

Adapting MLLMs for Nuanced Video Retrieval

DGX agent

arXiv:2512.13511v2 Announce Type: replace Abstract: Our objective is to build an embedding model that captures the nuanced relationship between a search query and candidate videos. We cover three aspe

researcharxiv-cs-cv
27 Apr 2026
Research

All Eyes on the Workflow: Automated and Efficient Event Discovery from Video Streams

DGX agent

arXiv:2604.22476v1 Announce Type: new Abstract: Disciplines such as business process management and process mining aid organizations by discovering insights about processes on the basis of recorded ev

researcharxiv-cs-cv
27 Apr 2026
Research

Altitude-Adaptive Vision-Only Geo-Localization for UAVs in GPS-Denied Environments

DGX agent

arXiv:2602.23872v2 Announce Type: replace Abstract: To address the scale mismatch caused by large altitude variations in UAV visual place recognition, we propose a monocular vision-only altitude-adapt

researcharxiv-cs-cv
27 Apr 2026
← Previous
1…214215216217218…261
Next →