AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
5 May 2026

Zero-Shot Interpretable Image Steganalysis for Invertible Image Hiding

Model ReleasesDGX agent

arXiv:2605.01331v1 Announce Type: new Abstract: Image steganalysis, which aims at detecting secret information concealed within images, has become a critical countermeasure for assessing the security

4 May 2026

2D-SuGaR: Surface-Aware Gaussian Splatting for Geometrically Accurate Mesh Reconstruction

ResearchDGX agent

arXiv:2605.00569v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful technique for generating photorealistic renderings of a scene in real-time. However, the volumetr

A Deep Learning-Based CCTV System for Automatic Smoking Detection in Fire Exit Zones


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety
DGX agent

arXiv:2508.11696v3 Announce Type: replace Abstract: A deep learning real-time smoking detection system for CCTV surveillance of fire exit areas is proposed due to critical safety requirements. The dat

A Model-based Visual Contact Localization and Force Sensing System for Compliant Robotic Grippers

Local AiDGX agent

arXiv:2605.00307v1 Announce Type: cross Abstract: Grasp force estimation can help prevent robots from damaging delicate objects during manipulation and improve learning-based robotic control. Integrat

A Novel Patch-Based TDA Approach for Computed Tomography Imaging

ResearchDGX agent

arXiv:2512.12108v5 Announce Type: replace Abstract: The development of machine learning models based on computed tomography (CT) imaging has been a major focus due to the promise that imaging holds fo

A Unified Deep Learning Framework for Motion Correction in Medical Imaging

Local AiDGX agent

arXiv:2409.14204v4 Announce Type: replace-cross Abstract: Deep learning has shown significant value in medical image registration for motion correction, however, current techniques are either limited

Adapting Large VLMs with Iterative and Manual Instructions for Generative Low-light Enhancement

TutorialsDGX agent

arXiv:2507.18064v2 Announce Type: replace Abstract: Most existing low-light image enhancement (LLIE) methods rely on pre-trained model priors, low-light inputs, or both, while neglecting the semantic

Adaptive Dual-Teacher Distillation with Subnetwork Rectification for Bridging Semantic Gaps in Black-Box Domain Adaptation

Model ReleasesDGX agent

arXiv:2603.22908v3 Announce Type: replace Abstract: Assuming that neither source data nor source model parameters are accessible, black-box domain adaptation (BBDA) represents a highly practical yet c

Adaptive Equilibrium: Dynamic Weighting Framework for Generalized Interruption of DeepFake Models

SafetyDGX agent

arXiv:2605.00443v1 Announce Type: cross Abstract: The advancement of generalized deepfake disruption is constrained by the interruption imbalance, a fundamental bottleneck inherent to the generation o

Adaptive Geodesic Conformal Prediction for Egocentric Camera Pose Estimation

ResearchDGX agent

arXiv:2605.00233v1 Announce Type: new Abstract: Egocentric pose estimation for Augmented Reality (AR) and assistive devices requires not just accurate predictions but guaranteed uncertainty regions. C

Affordance Agent Harness: Verification-Gated Skill Orchestration

AgentsDGX agent

arXiv:2605.00663v1 Announce Type: cross Abstract: Affordance grounding requires identifying where and how an agent should interact in open-world scenes, where actionable regions are often small, occlu

AIDA-ReID: Adaptive Intermediate Domain Adaptation for Generalizable and Source-Free Person Re-Identification

ResearchDGX agent

arXiv:2605.00111v1 Announce Type: new Abstract: Person re-identification (Re-ID) aims to match images of the same individual across non-overlapping camera views and remains challenging due to domain s

An End-to-End Decision-Aware Multi-Scale Attention-Based Model for Explainable Autonomous Driving

AgentsDGX agent

arXiv:2605.00291v1 Announce Type: new Abstract: The application of computer vision is gradually increasing across various domains. They employ deep learning models with a black-box nature. Without the

Being-H0.7: A Latent World-Action Model from Egocentric Videos

ApplicationsDGX agent

arXiv:2605.00078v1 Announce Type: cross Abstract: Visual-Language-Action models (VLAs) have advanced generalist robot control by mapping multimodal observations and language instructions directly to a

Beyond Heuristics: Learnable Density Control for 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2605.00408v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has demonstrated impressive real-time rendering performance, its efficacy remains constrained by a reliance on heuris

Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration

Model ReleasesDGX agent

arXiv:2605.00310v1 Announce Type: new Abstract: Super-resolution (SR) techniques have made major advances in reconstructing high-resolution images from low-resolution inputs. The increased resolution

BlenderRAG: High-Fidelity 3D Object Generation via Retrieval-Augmented Code Synthesis

SafetyDGX agent

arXiv:2605.00632v1 Announce Type: new Abstract: Automatic generation of executable Blender code from natural language remains challenging, with state-of-the-art LLMs producing frequent syntactic error

BOLT: Online Lightweight Adaptation for Preparation-Free Heterogeneous Cooperative Perception

SafetyDGX agent

arXiv:2605.00405v1 Announce Type: new Abstract: Most existing heterogeneous cooperative perception methods depend on prior preparation like offline joint training or tailored collaborator-model adapta

Brain MR Image Synthesis with 3D Multi-Contrast Self-Attention GAN

ResearchDGX agent

arXiv:2604.00070v2 Announce Type: replace-cross Abstract: Complete and high-quality multi-modal Magnetic Resonance Imaging (MRI) is essential for accurate neuro-oncological assessment, as each contras

Broadband Wide Field of View Imaging with Computational Mirrors

ResearchDGX agent

arXiv:2605.00029v1 Announce Type: cross Abstract: Traditional glass-based optics are typically optimized for narrow spectral bands, such as the visible (400-700nm) or shortwave infrared (1000-1800nm).

Certifiable Factor Graph Optimization

ResearchDGX agent

arXiv:2603.01267v2 Announce Type: replace-cross Abstract: We show that the factor graph and certifiable estimation paradigms, which have thus far been treated as essentially independent in the literat

ClustViT: Clustering-based Token Merging for Semantic Segmentation

ApplicationsDGX agent

arXiv:2510.01948v2 Announce Type: replace Abstract: Vision Transformers can achieve high accuracy and strong generalization across various contexts, but their practical applicability on real-world rob

CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection

Model ReleasesDGX agent

arXiv:2605.00630v1 Announce Type: new Abstract: The proliferation of advanced AI video synthesis techniques poses an unprecedented challenge to digital video authenticity. Existing AI-generated video

CollaFuse: Collaborative Diffusion Models

ResearchDGX agent

arXiv:2406.14429v3 Announce Type: replace-cross Abstract: In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic ima

Color Conditional Generation with Sliced Wasserstein Guidance

ResearchDGX agent

arXiv:2503.19034v2 Announce Type: replace Abstract: We propose SW-Guidance, a training-free approach for image generation conditioned on the color distribution of a reference image. While it is possib

Colorful-Noise: Training-Free Low-Frequency Noise Manipulation for Color-Based Conditional Image Generation

ResearchDGX agent

arXiv:2605.00548v1 Announce Type: new Abstract: Text-to-image diffusion models generate images by gradually converting white Gaussian noise into a natural image. White Gaussian noise is well suited fo

Combined Dictionary Unfolding Network with Gradient-Adaptive Fidelity for Transferable Multi-Source Fusion

ResearchDGX agent

arXiv:2605.00461v1 Announce Type: cross Abstract: Deep Unfolding Network-based methods have emerged as effective solutions for multi-source image fusion by combining model-driven iterative optimizatio

Copula-enhanced Vision Transformer for high myopia diagnosis through OU UWF fundus images

ResearchDGX agent

arXiv:2501.06540v2 Announce Type: replace Abstract: The advancement of AI-assisted myopia screening necessitates the joint diagnosis of both-eye (OU) high myopia (HM) status and the prediction of axia

CURE-OOD: Benchmarking Out-of-Distribution Detection for Survival Prediction

Model ReleasesDGX agent

arXiv:2605.00350v1 Announce Type: new Abstract: ``How long can I live and remain free of cancer?'' is often the first question a patient asks after receiving a cancer diagnosis and treatment. Accurate

Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations

SafetyDGX agent

arXiv:2512.20260v5 Announce Type: replace Abstract: Weakly-Supervised Camouflaged Object Detection (WSCOD) aims to locate and segment objects that are visually concealed within their surrounding scene

Deepfakes: we need to re-think the concept of 'real' images

Model ReleasesDGX agent

arXiv:2509.21864v2 Announce Type: replace Abstract: The wide availability and low usability barrier of modern image generation models has triggered the reasonable fear of criminal misconduct and negat

Depth-Guided Privacy-Preserving Visual Localization Using 3D Sphere Clouds

Local AiDGX agent

arXiv:2605.00562v1 Announce Type: new Abstract: The emergence of deep neural networks capable of revealing high-fidelity scene details from sparse 3D point clouds has raised significant privacy concer

DiffMI: Breaking Face Recognition Privacy via Diffusion-Driven Training-Free Model Inversion

ResearchDGX agent

arXiv:2504.18015v4 Announce Type: replace-cross Abstract: Face recognition poses serious privacy risks due to its reliance on sensitive and immutable biometric data. While modern systems mitigate priv

Diffusion Models are Secretly Zero-Shot 3DGS Harmonizers

ApplicationsDGX agent

arXiv:2503.06740v2 Announce Type: replace Abstract: Gaussian Splatting has become a popular technique for various 3D Computer Vision tasks, including novel view synthesis, scene reconstruction, and dy

Diffusion Models for Solving Inverse Problems via Posterior Sampling with Piecewise Guidance

ResearchDGX agent

arXiv:2507.18654v2 Announce Type: replace-cross Abstract: Diffusion models are powerful tools for sampling from high-dimensional distributions by progressively transforming pure noise into structured

Discrete Cosine Transform Based Decorrelated Attention for Vision Transformers

ResearchDGX agent

arXiv:2405.13901v4 Announce Type: replace Abstract: Self-attention is central to the success of Transformer architectures; however, learning the query, key, and value projections from random initializ

DMDSC: A Dynamic-Margin Deep Simplex Classifier for Open-Set Recognition on Medical Image Datasets

ResearchDGX agent

arXiv:2605.00675v1 Announce Type: new Abstract: Medical imaging datasets are often characterized by extreme class imbalances, where rare pathologies are significantly underrepresented compared to comm

DPU or GPU for Accelerating Neural Networks Inference -- Why not both? Split CNN Inference

HardwareDGX agent

arXiv:2605.00174v1 Announce Type: cross Abstract: Video and image streaming on edge devices requires low latency. To address this, Neural Networks (NNs) are widely used, and prior work mainly focuses

Driving with A Thousand Faces: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving

Model ReleasesDGX agent

arXiv:2602.18757v2 Announce Type: replace Abstract: Human driving behavior is inherently diverse, yet most end-to-end autonomous driving (E2E-AD) systems learn a single average driving style, neglecti

Efficient Spatio-Temporal Vegetation Pixel Classification with Vision Transformers

Model ReleasesDGX agent

arXiv:2605.00296v1 Announce Type: new Abstract: Plant phenology-the study of recurrent life cycle events-is essential for understanding ecosystem dynamics and their responses to climate change impacts

Elimination Templates in Macaulay2

ResearchDGX agent

arXiv:2605.00278v1 Announce Type: cross Abstract: We introduce the package exttt{EliminationTemplates} for the Macaulay2 computer algebra system, which provides tools for constructing automatic solver

End-to-End Autoregressive Image Generation with 1D Semantic Tokenizer

ResearchDGX agent

arXiv:2605.00503v1 Announce Type: new Abstract: Autoregressive image modeling relies on visual tokenizers to compress images into compact latent representations. We design an end-to-end training pipel

Event-based Civil Infrastructure Visual Defect Detection: ev-CIVIL Dataset and Benchmark

Model ReleasesDGX agent

arXiv:2504.05679v2 Announce Type: replace Abstract: Small unmanned aerial vehicle (UAV)-based visual inspections are a more efficient alternative to manual methods for examining civil structural defec

Exploring the Limits of End-to-End Feature-Affinity Propagation for Single-Point Supervised Infrared Small Target Detection

ResearchDGX agent

arXiv:2605.00722v1 Announce Type: new Abstract: Single-point supervised infrared small target detection (IRSTD) drastically reduces dense annotation costs. Current state-of-the-art (SOTA) methods achi

Faithful Extreme Image Rescaling with Learnable Reversible Transformation and Semantic Priors

Model ReleasesDGX agent

arXiv:2605.00605v1 Announce Type: new Abstract: Most recent extreme rescaling methods struggle to preserve semantically consistent structures and produce realistic details, due to the severely ill-pos

Federated Distillation for Whole Slide Image via Gaussian-Mixture Feature Alignment and Curriculum Integration

Local AiDGX agent

arXiv:2605.00578v1 Announce Type: new Abstract: Federated learning (FL) offers a promising framework for collaborative digital pathology by enabling model training across institutions. However, real-w

FieryGS: In-the-Wild Fire Synthesis with Physics-Integrated Gaussian Splatting

ApplicationsDGX agent

arXiv:2605.00177v1 Announce Type: cross Abstract: We consider the problem of synthesizing photorealistic, physically plausible combustion effects in in-the-wild 3D scenes. Traditional CFD and graphics

Flow matching for Sentinel-2 super-resolution: implementation, application, and implications

ResearchDGX agent

arXiv:2605.00367v1 Announce Type: new Abstract: Developing robust techniques for super-resolution of satellite imagery involves navigating commonly observed trade-offs between spectral fidelity and pe

Foundation AI Models for Aerosol Optical Depth Estimation from PACE Satellite Data

SafetyDGX agent

arXiv:2605.00678v1 Announce Type: new Abstract: Aerosol Optical Depth (AOD) retrieval is essential for Earth observation, supporting applications from air quality monitoring to climate studies. Conven

FreeRet: MLLMs as Training-Free Retrievers

SafetyDGX agent

arXiv:2509.24621v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are emerging as versatile foundations for mixed-modality retrieval. Yet, they often require heavy post-hoc

From Images2Mesh: A 3D Surface Reconstruction Pipeline for Non-Cooperative Space Objects

Model ReleasesDGX agent

arXiv:2605.00147v1 Announce Type: new Abstract: On-orbit inspection imagery is crucial as it enables characterization of non-cooperative resident space objects, providing the geometry and structural c

From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models

ResearchDGX agent

arXiv:2605.00474v1 Announce Type: new Abstract: Modern vision models achieve remarkable accuracy, but explaining where evidence arises, what the model encodes, and how internal computations assemble t

GAFSV-Net: A Vision Framework for Online Signature Verification

ResearchDGX agent

arXiv:2605.00120v1 Announce Type: new Abstract: Online signature verification (OSV) requires distinguishing skilled forgeries from genuine samples under high intra-class variability and with very few

Gated Differential Linear Attention: A Linear-Time Decoder for High-Fidelity Medical Segmentation

Local AiDGX agent

arXiv:2603.02727v4 Announce Type: replace Abstract: Medical image segmentation requires models that preserve fine anatomical boundaries while remaining practical for clinical deployment. Transformers

GMGaze: MoE-Based Context-Aware Gaze Estimation with CLIP and Multiscale Transformer

ResearchDGX agent

arXiv:2605.00799v1 Announce Type: new Abstract: Gaze estimation methods commonly use facial appearances to predict the direction of a person gaze. However, previous studies show three major challenges

GOR-IS: 3D Gaussian Object Removal in the Intrinsic Space

ApplicationsDGX agent

arXiv:2605.00498v1 Announce Type: new Abstract: Recent advances in Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have made it standard practice to reconstruct 3D scenes from multi-vie

High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions

ResearchDGX agent

arXiv:2605.00496v1 Announce Type: new Abstract: Understanding human actions from visual observations is essential for human--robot interaction, particularly when semantic interpretation of unfamiliar

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

Model ReleasesDGX agent

arXiv:2507.01955v3 Announce Type: replace Abstract: Multimodal foundation models (MFMs), such as GPT-4o, have recently made remarkable progress. However, their detailed visual understanding beyond que

IdentiFace: Multi-Modal Iterative Diffusion Framework for Identifiable Suspect Face Generation in Crime Investigations

ApplicationsDGX agent

arXiv:2605.00526v1 Announce Type: new Abstract: Suspect face generation remains a technical challenge in crime investigations. Traditional sketch-drawing workflows suffer from low efficiency and quali

Image Score: Learning and Evaluating Human Preferences for Mercari Search

ResearchDGX agent

arXiv:2408.11349v2 Announce Type: replace Abstract: Mercari is the largest C2C e-commerce marketplace in Japan, having more than 20 million active monthly users. Search being the fundamental way to di

← Previous
1…160161162163164…209
Next →