AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
4 May 2026

Unsupervised Denoising of Real Clinical Low Dose Liver CT with Perceptual Attention Networks

ResearchDGX agent

arXiv:2605.00793v1 Announce Type: cross Abstract: With the development of deep learning, medical image processing has been widely used to assist clinical research. This paper focuses on the denoising

VecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single Image

ResearchDGX agent

arXiv:2602.04349v2 Announce Type: replace Abstract: 3D editing has emerged as a critical research area to provide users with flexible control over 3D assets. While current editing approaches predomina

Vesselpose: Vessel Graph Reconstruction from Learned Voxel-wise Direction Vectors in 3D Vascular Images

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.00538v1 Announce Type: new Abstract: Blood vessel segmentation and -tracing are essential tasks in many medical imaging applications. Although numerous methods exist, the prevailing segment

VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding

ResearchDGX agent

arXiv:2603.22285v2 Announce Type: replace Abstract: Long video understanding remains challenging for multimodal large language models (MLLMs) due to limited context windows, which necessitate identify

VkSplat: High-Performance 3DGS Training in Vulkan Compute

Local AiDGX agent

arXiv:2605.00219v1 Announce Type: new Abstract: We present VkSplat, a high-performance, cross-vendor 3D Gaussian Splatting (3DGS) training pipeline implemented fully in Vulkan compute, addressing perf

When Do Diffusion Models learn to Generate Multiple Objects?

TutorialsDGX agent

arXiv:2605.00273v1 Announce Type: new Abstract: Text-to-image diffusion models achieve impressive visual fidelity, yet they remain unreliable in multi-object generation. Despite extensive empirical ev

WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery

ResearchDGX agent

arXiv:2602.13305v2 Announce Type: replace Abstract: Wildfires are a growing threat to ecosystems, human lives, and infrastructure, with their frequency and intensity rising due to climate change and h

World Model for Robot Learning: A Comprehensive Survey

SafetyDGX agent

arXiv:2605.00080v1 Announce Type: cross Abstract: World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They s

1 May 2026

3D Reconstruction Techniques in the Manufacturing Domain: Applications, Research Opportunities and Use Cases

ApplicationsDGX agent

arXiv:2604.28064v1 Announce Type: new Abstract: This comprehensive review examines the evolution and the current state of the art in three-dimensional (3D) reconstruction techniques in manufacturing a

3D-ReGen: A Unified 3D Geometry Regeneration Framework

ResearchDGX agent

arXiv:2604.28134v1 Announce Type: new Abstract: We consider the problem of regenerating 3D objects from 2D images and initial 3D shapes. Most 3D generators operate in a one-shot fashion, converting te

A generalised pre-training strategy for deep learning networks in semantic segmentation of remotely sensed images

TutorialsDGX agent

arXiv:2604.27704v1 Announce Type: new Abstract: In the segmentation of remotely sensed images, deep learning models are typically pre-trained using large image databases like ImageNet before fine-tune

A Real-time Scale-robust Network for Glottis Segmentation in Nasal Transnasal Intubation

ResearchDGX agent

arXiv:2604.27383v1 Announce Type: cross Abstract: Nasotracheal intubation (NTI) is a critical clinical procedure for establishing and maintaining patient airway patency. Machine-assisted NTI has emerg

A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion

ResearchDGX agent

arXiv:2501.07451v4 Announce Type: replace Abstract: Model compression is essential in the deployment of large Computer Vision models on embedded devices. However, static optimization techniques (e.g.

Action Motifs: Self-Supervised Hierarchical Representation of Human Body Movements

TutorialsDGX agent

arXiv:2604.28173v1 Announce Type: new Abstract: Effective human behavior modeling requires a representation of the human body movement that capitalizes on its compositionality. We propose a hierarchic

Adjoint Inversion Reveals Holographic Superposition and Destructive Interference in CNN Classifiers

ResearchDGX agent

arXiv:2604.27529v1 Announce Type: new Abstract: A foundational assumption in CNN interpretability -- that deep encoders suppress background pixels while classifiers merely select from a cleaned featur

AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images

Model ReleasesDGX agent

arXiv:2604.28177v1 Announce Type: new Abstract: We introduce AEGIS, A holistic benchmark for Evaluating forensic analysis of AI-Generated academic ImageS. Compared to existing benchmarks, AEGIS featur

AesRM: Improving Video Aesthetics with Expert-Level Feedback

Model ReleasesDGX agent

arXiv:2604.28078v1 Announce Type: new Abstract: Despite rapid advances in photorealistic video generation, real-world applications such as filmmaking require video aesthetics, e.g., harmonious colors

AG-TAL: Anatomically-Guided Topology-Aware Loss for Multiclass Segmentation of the Circle of Willis Using Large-Scale Multi-Center Datasets

ResearchDGX agent

arXiv:2604.27357v1 Announce Type: cross Abstract: Accurate multiclass segmentation of the Circle of Willis (CoW) is essential for neurovascular disease management but remains challenging due to comple

An Extended Evaluation Split for DeepSpaceYoloDataset

ResearchDGX agent

arXiv:2604.27593v1 Announce Type: cross Abstract: Recent technological advances in astronomy, particularly the growing popularity of smart telescopes for the general public, make it possible to develo

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge

ResearchDGX agent

arXiv:2604.28022v1 Announce Type: new Abstract: Current DeepFake detection scenarios are mostly binary, yet data manipulation can vary across audio, video, or both, whose variability is not captured i

Assessing Pancreatic Ductal Adenocarcinoma Vascular Invasion: the PDACVI Benchmark

Model ReleasesDGX agent

arXiv:2604.27582v1 Announce Type: new Abstract: Surgical resection remains the only potentially curative treatment for pancreatic ductal adenocarcinoma (PDAC), and eligibility depends on accurate asse

AttriBE: Quantifying Attribute Expressivity in Body Embeddings for Recognition and Identification

SafetyDGX agent

arXiv:2604.27218v1 Announce Type: new Abstract: Person re-identification (ReID) systems that match individuals across images or video frames are essential in many real-world applications. However, exi

Automated Detection of Mutual Gaze and Joint Attention in Dual-Camera Settings via Dual-Stream Transformers

ResearchDGX agent

arXiv:2604.27105v1 Announce Type: new Abstract: Analyzing mutual gaze (MG) and joint attention (JA) is critical in developmental psychology but traditionally relies on labor-intensive manual coding. A

Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models

AgentsDGX agent

arXiv:2512.22046v2 Announce Type: replace Abstract: Prompt-driven Video Segmentation Foundation Models (VSFMs), such as SAM2, are increasingly used in applications including autonomous driving and dig

Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces

ResearchDGX agent

arXiv:2604.28122v1 Announce Type: new Abstract: Modern visual world modeling systems increasingly rely on high-capacity architectures and large-scale data to produce plausible motion, yet they often f

Beyond Pixel Fidelity: Minimizing Perceptual Distortion and Color Bias in Night Photography Rendering

SafetyDGX agent

arXiv:2604.28136v1 Announce Type: new Abstract: Night Photography Rendering (NPR) poses a significant challenge due to the extreme contrast between dark and illuminated areas in scenes, stemming from

CasLayout: Cascaded 3D Layout Diffusion for Indoor Scene Synthesis with Implicit Relation Modeling

Local AiDGX agent

arXiv:2604.27361v1 Announce Type: new Abstract: Synthesizing realistic 3D indoor scenes remains challenging due to data scarcity and the difficulty of simultaneously enforcing global architectural con

CBEN -- A Multimodal Machine Learning Dataset for Cloud Robust Remote Sensing Image Understanding

ResearchDGX agent

arXiv:2602.12652v2 Announce Type: replace Abstract: Clouds are a common phenomenon that distorts optical satellite imagery, which poses a challenge for remote sensing. However, in the literature cloud

ClimateVID -- Social Media Videos Analysis and Challenges Involved

ResearchDGX agent

arXiv:2604.27968v1 Announce Type: new Abstract: The pervasive growth of digital content, specifically short videos on social media platforms, has significantly altered how topics are discussed and und

Context as Prior: Bayesian-Inspired Intent Inference for Non-Speaking Agents with a Household Cat Testbed

ApplicationsDGX agent

arXiv:2604.27445v1 Announce Type: new Abstract: Many agents in real-world environments cannot reliably communicate their goals through language, including household pets, pre-verbal infants, and other

Continuous-tone Simple Points: An ell_0-Norm of Cyclic Gradient for Topology-Preserving Data-Driven Image Segmentation

ResearchDGX agent

arXiv:2604.28159v1 Announce Type: new Abstract: Topological features play an essential role in ensuring geometric plausibility and structural consistency in image analysis tasks such as segmentation a

Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning

Model ReleasesDGX agent

arXiv:2604.27604v1 Announce Type: new Abstract: We introduce SPUR, a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answe

Detecting Malicious Concepts without Image Generation in AI-Generated Content (AIGC)

ResearchDGX agent

arXiv:2502.08921v2 Announce Type: replace-cross Abstract: The task of text-to-image generation has achieved tremendous success in practice, with emerging concept generation models capable of producing

DOT-Sim: Differentiable Optical Tactile Simulation with Precise Real-to-Sim Physical Calibration

SafetyDGX agent

arXiv:2604.27367v1 Announce Type: cross Abstract: Simulating optical tactile sensors presents significant challenges due to their high deformability and intricate optical properties. To address these

Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training

ResearchDGX agent

arXiv:2604.27932v1 Announce Type: new Abstract: The computational cost of training a vision-language model (VLM) can be reduced by sampling the training data. Previous work on efficient VLM pre-traini

EAG-PT: Emission-Aware Gaussians and Path Tracing for Diffuse Indoor Scene Reconstruction and Editing

ResearchDGX agent

arXiv:2601.23065v2 Announce Type: replace-cross Abstract: Recent radiance-field-based reconstruction methods, such as NeRF and 3DGS, achieve high visual fidelity for indoor scenes, but often break dow

Echo-{alpha}: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation

Local AiDGX agent

arXiv:2604.28011v1 Announce Type: new Abstract: Ultrasound interpretation requires both precise lesion localization and holistic clinical reasoning, yet existing methods typically excel at only one of

EdgeFM: Efficient Edge Inference for Vision-Language Models

HardwareDGX agent

arXiv:2604.27476v1 Announce Type: new Abstract: Vision-language models (VLMs) have demonstrated strong applicability in edge industrial applications, yet their deployment remains severely constrained

Effective Prompt Pool Learning for Continual Category Discovery

Local AiDGX agent

arXiv:2407.19001v3 Announce Type: replace Abstract: This paper studies effective prompt pool learning for Continual Category Discovery (CCD), a challenging open-world setting where a model must discov

ELiC: Efficient LiDAR Geometry Compression via Cross-Bit-depth Feature Propagation and Bag-of-Encoders

Local AiDGX agent

arXiv:2511.14070v3 Announce Type: replace-cross Abstract: Hierarchical LiDAR geometry compression encodes voxel occupancies from low to high bit-depths, yet prior methods treat each depth independentl

Energy-Efficient Plant Monitoring via Knowledge Distillation

ApplicationsDGX agent

arXiv:2604.27178v1 Announce Type: new Abstract: Recent advances in large-scale visual representation learning have significantly improved performance in plant species and plant disease recognition tas

Fake3DGS: A Benchmark for 3D Manipulation Detection in Neural Rendering

Model ReleasesDGX agent

arXiv:2604.27590v1 Announce Type: new Abstract: Recent advances in 3D reconstruction and neural rendering,particularly 3D Gaussian Splatting, make it feasible and simple to edit 3D scenes and re-rende

Faster 3D Gaussian Splatting Convergence via Structure-Aware Densification

ResearchDGX agent

arXiv:2604.28016v1 Announce Type: new Abstract: 3D Gaussian Splatting has emerged as a powerful scene representation for real-time novel-view synthesis. However, its standard adaptive density control

FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting

Model ReleasesDGX agent

arXiv:2604.27974v1 Announce Type: new Abstract: Despite the rapid progress of large vision-language models (LVLMs), fine-grained, state-conditioned GUI interaction remains challenging. Current evaluat

FMCL: Class-Aware Client Clustering with Foundation Model Representations for Heterogeneous Federated Learning

ResearchDGX agent

arXiv:2604.27510v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative model training across distributed clients without sharing raw data, yet its performance deteriorates und

FreeOcc: Training-Free Embodied Open-Vocabulary Occupancy Prediction

Model ReleasesDGX agent

arXiv:2604.28115v1 Announce Type: cross Abstract: Existing learning-based occupancy prediction methods rely on large-scale 3D annotations and generalize poorly across environments. We present FreeOcc,

Frequency-Adaptive Discrete Cosine-ViT-ResNet Architecture for Sparse-Data Vision

ResearchDGX agent

arXiv:2505.22701v3 Announce Type: replace Abstract: A major challenge in rare animal image classification is the scarcity of data, as many species usually have only a small number of labeled samples.

Frequency-Aware Semantic Fusion with Gated Injection for AI-generated Image Detection

SafetyDGX agent

arXiv:2604.27875v1 Announce Type: new Abstract: AI-generated images are becoming increasingly realistic and diverse, posing significant challenges for generalizable detection. While Vision Foundation

FUN: A Focal U-Net Combining Reconstruction and Object Detection for Snapshot Spectral Imaging

ResearchDGX agent

arXiv:2604.27653v1 Announce Type: new Abstract: Conventional push-broom hyperspectral imaging suffers from slow acquisition speeds, precluding real-time object detection; in contrast, snapshot spectra

Gait Recognition via Deep Residual Networks and Multi-Branch Feature Fusion

Model ReleasesDGX agent

arXiv:2604.27353v1 Announce Type: new Abstract: Gait recognition has emerged as a compelling biometric modality for surveillance and security applications, offering inherent advantages such as non-int

Generalizable Sparse-View 3D Reconstruction from Unconstrained Images

Model ReleasesDGX agent

arXiv:2604.28193v1 Announce Type: new Abstract: Reconstructing 3D scenes from sparse, unposed images remains challenging under real-world conditions with varying illumination and transient occlusions.

Generate Your Talking Avatar from Video Reference

Model ReleasesDGX agent

arXiv:2604.27918v1 Announce Type: new Abstract: Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target g

Generative Human Geometry Distribution

ResearchDGX agent

arXiv:2503.01448v5 Announce Type: replace Abstract: Realistic human geometry generation is an important yet challenging task, requiring both the preservation of fine clothing details and the accurate

GourNet: A CNN-Based Model for Mango Leaf Disease Detection

ApplicationsDGX agent

arXiv:2604.27764v1 Announce Type: new Abstract: Mango cultivation is crucial in the agricultural sector, significantly contributing to economic development and food security. However, diseases affecti

GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance

Model ReleasesDGX agent

arXiv:2503.12844v2 Announce Type: replace Abstract: For people affected by blindness and low vision (BLV), safe and independent navigation remains a major challenge, impacting over 2.2 billion individ

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

Model ReleasesDGX agent

arXiv:2604.28196v1 Announce Type: new Abstract: Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existing approaches predominant

HiMix: Hierarchical Artifact-aware Mixup for Generalized Synthetic Image Detection

ResearchDGX agent

arXiv:2604.27903v1 Announce Type: new Abstract: The rapid evolution of generative models has enabled the creation of highly realistic and diverse synthetic images, posing significant challenges to rel

HQ-UNet: A Hybrid Quantum-Classical U-Net with a Quantum Bottleneck for Remote Sensing Image Segmentation

Model ReleasesDGX agent

arXiv:2604.27206v1 Announce Type: new Abstract: Semantic segmentation in remote sensing is commonly addressed using classical deep learning architectures such as U-Net, which require a large number of

Hyperspectral Image Classification via Efficient Global Spectral Supertoken Clustering

ResearchDGX agent

arXiv:2604.27364v1 Announce Type: new Abstract: Hyperspectral image classification demands spatially coherent predictions and precise boundary delineation. Yet prevailing superpixel-based methods face

Improving Calibration in Test-Time Prompt Tuning for Vision-Language Models via Data-Free Flatness-Aware Prompt Pretraining

ApplicationsDGX agent

arXiv:2604.27715v1 Announce Type: new Abstract: Test-time prompt tuning (TPT) has emerged as a promising technique for enhancing the adaptability of vision-language models by optimizing textual prompt

← Previous
1…162163164165166…209
Next →