AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
29 Apr 2026

Personalization Toolkit: Training Free Personalization of Large Vision Language Models

Model ReleasesDGX agent

arXiv:2502.02452v4 Announce Type: replace Abstract: Personalization of Large Vision-Language Models (LVLMs) involves customizing models to recognize specific users or object instances and to generate

Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation

TutorialsDGX agent

arXiv:2604.25255v1 Announce Type: new Abstract: Speech-preserving facial expression manipulation (SPFEM) aims to enhance human expressiveness without altering mouth movements tied to the original spee

PhyloSDF: Phylogenetically-Conditioned Neural Generation of 3D Skull Morphology via Residual Flow Matching

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.25371v1 Announce Type: cross Abstract: Generating novel, biologically plausible three-dimensional morphological structures is a fundamental challenge in computational evolutionary biology,

PortraVec: Image-Based Portrait Vectorization with Text-Guided Manipulation

Model ReleasesDGX agent

arXiv:2410.04182v2 Announce Type: replace Abstract: While portrait sketch generation is a special task in sketch synthesis, most existing methods are pixel-based, limiting their interpretability and e

Power Foam: Unifying Real-Time Differentiable Ray Tracing and Rasterization

ResearchDGX agent

arXiv:2604.24994v1 Announce Type: cross Abstract: We introduce a differentiable 3D representation that unifies the ray tracing capabilities of foam-based ray tracing with the efficiency of modern rast

Practical exposure correction via compensation

HardwareDGX agent

arXiv:2212.14245v2 Announce Type: replace Abstract: In computer vision, correcting the exposure level is a fundamental task for enhancing the visual quality of observations with inappropriate lightnes

Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models

ResearchDGX agent

arXiv:2604.25642v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual-textual understanding, yet their reliability is critically undermined b

QB-LIF: Learnable-Scale Quantized Burst Neurons for Efficient SNNs

Model ReleasesDGX agent

arXiv:2604.25688v1 Announce Type: new Abstract: Binary spike coding enables sparse and event-driven computation in spiking neural networks (SNNs), yet its 1-bit-per-timestep representation fundamental

QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding

Model ReleasesDGX agent

arXiv:2604.25884v1 Announce Type: cross Abstract: Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representatio

Quantum-Inspired Robust and Scalable SAR Object Classification

ResearchDGX agent

arXiv:2604.25755v1 Announce Type: cross Abstract: SAR image classification naturally has to deal with huge noise and a high dynamic range particularly requiring robust classification models. Additiona

RABC-Net: Reliability-Aware Annotation-Free Skin Lesion Segmentation for Low-Resource Dermoscopy

Local AiDGX agent

arXiv:2604.05594v2 Announce Type: replace Abstract: Pixel-level annotation is costly in low-resource dermoscopy. We present RABC-Net, a reliability-aware annotation-free segmentation system that combi

Rapid tracking through strongly scattering media with physics-informed neuromorphic speckle analysis

ResearchDGX agent

arXiv:2604.25310v1 Announce Type: new Abstract: This work addresses the critical problem of tracking fast-moving objects through strongly scattering media in a low-light environment. Different from ex

Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models

SafetyDGX agent

arXiv:2604.25636v1 Announce Type: new Abstract: Unified multimodal models (UMMs) integrate visual understanding and generation within a single framework. For text-to-image (T2I) tasks, this unified ca

Representation Paradigms in AI-based 3D Radiological Image Reconstruction: A Systematic Review

Model ReleasesDGX agent

arXiv:2504.11349v3 Announce Type: replace Abstract: The demand for high-quality medical imaging in clinical practice and assisted diagnosis has made 3D image reconstruction in radiological imaging a k

ResetEdit: Precise Text-guided Editing of Generated Image via Resettable Starting Latent

SafetyDGX agent

arXiv:2604.25128v1 Announce Type: new Abstract: Recent advances in diffusion models have enabled high-quality image generation, leading to increasing demand for post-generation editing that modifies l

ReSim: Reliable World Simulation for Autonomous Driving

SafetyDGX agent

arXiv:2506.09981v2 Announce Type: replace Abstract: How can we reliably simulate future driving scenarios under a wide range of ego driving behaviors? Recent driving world models, developed exclusivel

Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles

ApplicationsDGX agent

arXiv:2604.25889v1 Announce Type: new Abstract: Current deepfake detection models achieve state-of-the-art performance on pristine academic datasets but suffer severe spatial attention drift under rea

Robustness Evaluation of a Foundation Segmentation Model Under Simulated Domain Shifts in Abdominal CT: Implications for Health Digital Twin Deployment

ApplicationsDGX agent

arXiv:2604.25685v1 Announce Type: cross Abstract: Foundation segmentation models such as the Segment Anything Model (SAM) have demonstrated strong generalization across natural images; however, their

SaliencyDecor: Enhancing Neural Network Interpretability through Feature Decorrelation

ResearchDGX agent

arXiv:2604.25315v1 Announce Type: new Abstract: Gradient-based saliency methods are widely used to interpret deep neural networks, yet they often produce noisy and unstable explanations that poorly al

SAMe: A Semantic Anatomy Mapping Engine for Robotic Ultrasound

Local AiDGX agent

arXiv:2604.25646v1 Announce Type: new Abstract: Robotic ultrasound has advanced local image-driven control, contact regulation, and view optimization, yet current systems lack the anatomical understan

SARE: Sample-wise Adaptive Reasoning for Training-free Fine-grained Visual Recognition

Model ReleasesDGX agent

arXiv:2603.17729v3 Announce Type: replace Abstract: Recent advances in Large Vision-Language Models (LVLMs) have enabled training-free Fine-Grained Visual Recognition (FGVR). However, effectively expl

SARU: A Shadow-Aware and Removal Unified Framework for Remote Sensing Images with New Benchmarks

Model ReleasesDGX agent

arXiv:2604.25432v1 Announce Type: new Abstract: Shadows are a prevalent problem in remote sensing imagery (RSI), degrading visual quality and severely limiting the performance of downstream tasks like

Scalable Secure Biometric Authentication without Auxiliary Identifiers

Local AiDGX agent

arXiv:2604.25071v1 Announce Type: cross Abstract: The prevalence of biometric authentication has been on the rise due to its ease of use and elimination of weak passwords. To date, most biometric auth

SecureScan: An AI-Driven Multi-Layer Framework for Malware and Phishing Detection Using Logistic Regression and Threat Intelligence Integration

Model ReleasesDGX agent

arXiv:2602.10750v2 Announce Type: replace-cross Abstract: The growing sophistication of modern malware and phishing campaigns has diminished the effectiveness of traditional signature-based intrusion

Self-DACE++: Robust Low-Light Enhancement via Efficient Adaptive Curve Estimation

Model ReleasesDGX agent

arXiv:2604.25367v1 Announce Type: new Abstract: In this paper, we present Self-DACE++, an improved unsupervised and lightweight framework for Low-Light Image Enhancement (LLIE), building upon our prev

Semantic-aware Random Convolution and Source Matching for Domain Generalization in Medical Image Segmentation

ResearchDGX agent

arXiv:2512.01510v3 Announce Type: replace Abstract: We tackle the challenging problem of single-source domain generalization (DG) for medical image segmentation, where we train a network on one domain

ShapeY: A Principled Framework for Measuring Shape Recognition Capacity via Nearest-Neighbor Matching

ResearchDGX agent

arXiv:2604.25065v1 Announce Type: new Abstract: Object recognition (OR) in humans relies heavily on shape cues and the ability to recognize objects across varying 3D viewpoints. Unlike humans, deep ne

SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring

Model ReleasesDGX agent

arXiv:2604.25855v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve ever-stronger performance on visual-language tasks. Even as traditional visual question answering bench

Sketch2Arti: Sketch-based Articulation Modeling of CAD Objects

ResearchDGX agent

arXiv:2604.25781v1 Announce Type: new Abstract: Articulation modeling aims to infer movable parts and their motion parameters for a 3D object, enabling interactive animation, simulation, and shape edi

Soft-TransFormers for Continual Learning

Model ReleasesDGX agent

arXiv:2411.16073v3 Announce Type: replace-cross Abstract: Inspired by the Well-initialized Lottery Ticket Hypothesis (WLTH), we introduce Soft-Transformer (Soft-TF), a parameter-efficient framework fo

Splatent: Splatting Diffusion Latents for Novel View Synthesis

ResearchDGX agent

arXiv:2512.09923v2 Announce Type: replace Abstract: Radiance field representations have recently been explored in the latent space of VAEs that are commonly used by diffusion models. This direction of

Subjective Portrait Region Cropping in Landscape Videos with Temporal Annotation Smoothing

Model ReleasesDGX agent

arXiv:2604.24947v1 Announce Type: new Abstract: With the rise of mobile video consumption on diverse handheld display resolutions and orientation modes, altering videos to aspect ratios poses challeng

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation

Model ReleasesDGX agent

arXiv:2506.23690v2 Announce Type: replace Abstract: Diffusion-based video motion customization facilitates the acquisition of human motion representations from a few video samples, while achieving arb

Task-Driven Prompt Learning: A Joint Framework for Multi-modal Cloud Removal and Segmentation

Model ReleasesDGX agent

arXiv:2601.12052v2 Announce Type: replace Abstract: Optical remote sensing imagery is indispensable for Earth observation, yet persistent cloud occlusion limits its downstream utility. Most cloud remo

The Forensic Cost of Watermark Removal

Model ReleasesDGX agent

arXiv:2604.25491v1 Announce Type: new Abstract: Current watermark removal methods are evaluated on two axes: attack success rate and perceptual quality. We show this is insufficient. While state-of-th

The Surprising Effectiveness of Canonical Knowledge Distillation for Semantic Segmentation

TutorialsDGX agent

arXiv:2604.25530v1 Announce Type: new Abstract: Recent knowledge distillation (KD) methods for semantic segmentation introduce increasingly complex hand-crafted objectives, yet are typically evaluated

The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents

Model ReleasesDGX agent

arXiv:2604.25299v1 Announce Type: new Abstract: Diffusion models have achieved success in high-fidelity data synthesis, yet their capacity for more complex, structured reasoning like text following ta

TopoMamba: Topology-Aware Scanning and Fusion for Segmenting Heterogeneous Medical Visual Media

ResearchDGX agent

arXiv:2604.25545v1 Announce Type: new Abstract: Visual state-space models (SSMs) have shown strong potential for medical image segmentation, yet their effectiveness is often limited by two practical i

Towards Robust Deep Learning-based Rumex Obtusifolius Detection from Drone Images

ResearchDGX agent

arXiv:2604.25316v1 Announce Type: new Abstract: Domain adaptation (DA) addresses the challenge of transferring a machine learning model trained on a source domain to a target domain with a different d

Towards Seamless Lunar Mosaics: Deep Radiometric Normalization for Cross-Sensor Orbital Imagery Using Chandrayaan-2 TMC Data

TutorialsDGX agent

arXiv:2604.25208v1 Announce Type: new Abstract: Radiometric inconsistencies remain a major challenge in generating seamless lunar mosaics from multi-mission orbital imagery due to variability in illum

UltraGS: Real-Time Physically-Decoupled Gaussian Splatting for Ultrasound Novel View Synthesis

HardwareDGX agent

arXiv:2511.07743v3 Announce Type: replace Abstract: Ultrasound imaging is a cornerstone of non-invasive clinical diagnostics, yet its limited field of view poses challenges for novel view synthesis. W

UniSER: A Foundation Model for Unified Soft Effects Removal

TutorialsDGX agent

arXiv:2511.14183v3 Announce Type: replace Abstract: Digital images are often degraded by soft effects such as lens flare, haze, shadows, and reflections, which reduce aesthetics even though the underl

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations

ApplicationsDGX agent

arXiv:2604.24885v1 Announce Type: new Abstract: We introduce an efficient, resolution-agnostic autoregressive (AR) image synthesis approach that generalizes to arbitrary resolutions and aspect ratios,

ViPO: Visual Preference Optimization at Scale

TutorialsDGX agent

arXiv:2604.24953v1 Announce Type: new Abstract: While preference optimization is crucial for improving visual generative models, how to effectively scale this paradigm remains largely unexplored. Curr

VISION-SLS: Safe Perception-Based Control from Learned Visual Representations via System Level Synthesis

SafetyDGX agent

arXiv:2604.24894v1 Announce Type: cross Abstract: We propose VISION-SLS, a method for nonlinear output-feedback control from high-resolution RGB images which provides robust constraint satisfaction gu

Vision SmolMamba: Spike-Guided Token Pruning for Energy-Efficient Spiking State-Space Vision Models

ResearchDGX agent

arXiv:2604.25570v1 Announce Type: new Abstract: Spiking Transformers have shown strong potential for long-range visual modeling through spike-driven self-attention. However, their quadratic token inte

When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents

Model ReleasesDGX agent

arXiv:2604.25213v1 Announce Type: new Abstract: OpenAI's GPT-Image-2 has effectively erased the visual boundary between authentic and AI-edited document images: a single number on a receipt can be rep

28 Apr 2026

2D Pre-Training for 3D Pose Estimation

TutorialsDGX agent

arXiv:2604.22830v1 Announce Type: new Abstract: Pre-training is a general method that is used in a range of deep learning tasks. By first training a model on one task, and then further training on the

2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA

Local AiDGX agent

arXiv:2604.23935v1 Announce Type: new Abstract: Audio-based video object segmentation aims to locate and segment objects in videos conditioned on audio cues, requiring precise understanding of both ap

6thGrid-Net: Unified Remote Sensing Image Dehazing Based on Color Restoration and Edge-Preserving

Model ReleasesDGX agent

arXiv:2604.24149v1 Announce Type: new Abstract: Remote sensing images are frequently degraded by adverse weather conditions, particularly clouds and haze, which severely impair downstream applications

A Digital Pathology Resource for Liver Cancer Quantification with Datasets, Benchmarks, and Tools

Local AiDGX agent

arXiv:2604.22858v1 Announce Type: new Abstract: Liver cancer, especially hepatocellular carcinoma (HCC), imposes a substantial global disease burden. Accurate diagnosis and prognostic assessment direc

A Graph-Augmented knowledge Distillation based Dual-Stream Vision Transformer with Region-Aware Attention for Gastrointestinal Disease Classification with Explainable AI

Local AiDGX agent

arXiv:2512.21372v2 Announce Type: replace-cross Abstract: The accurate classification of gastrointestinal diseases from endoscopic and histopathological imagery remains a significant challenge in medi

A Heterogeneous Two-Stream Framework for Video Action Recognition with Comparative Fusion Analysis

ResearchDGX agent

arXiv:2604.23415v1 Announce Type: new Abstract: Most two-stream action recognition networks apply the same convolutional backbone to both RGB and optical flow streams, ignoring the fact that the two m

A Hierarchical Ensemble Inference Pipeline for Robust White Blood Cell Classification Under Domain Shifts

ApplicationsDGX agent

arXiv:2604.23271v1 Announce Type: new Abstract: Automated white blood cell (WBC) classification is essential for scalable leukaemia screening. However, real-world deployment is challenged by domain sh

A Hierarchical Self-Consistent Regularization Approach to Satellite Image Time Series Classification

Model ReleasesDGX agent

arXiv:2510.04916v2 Announce Type: replace Abstract: Deep learning has become increasingly important in remote sensing image classification due to its ability to extract semantic information from compl

A Pose-only Geometric Constraint for Multi-Camera Pose Adjustment

Model ReleasesDGX agent

arXiv:2604.23704v1 Announce Type: new Abstract: Multi-camera systems offer rich observation capabilities for visual navigation and 3D scene reconstruction; however, the resulting feature redundancy of

A satellite foundation model for improved wealth monitoring

Model ReleasesDGX agent

arXiv:2604.23166v1 Announce Type: cross Abstract: Poverty statistics guide social policy, but in many low- and middle-income countries, censuses and household surveys that collect these data are costl

A Synergistic CNN-Transformer Network with Pooling Attention Fusion for Hyperspectral Image Classification

ResearchDGX agent

arXiv:2604.23622v1 Announce Type: new Abstract: In the hyperspectral image (HSI) classification task, each pixel is categorized into a specific land-cover category or material. Convolutional neural ne

A Topology fixated Shape Gradient Framework for Non Simple Boundary Extraction for CIE Lab color images with Repulsive Energy

ResearchDGX agent

arXiv:2604.23167v1 Announce Type: new Abstract: A levelset free but a hybrid image segmentation approach based on a modified version of the piece wise constant shape gradient of an Mumford Shah shape

Accelerating New Product Introduction for Visual Quality Inspection via Few-Shot Diffusion-Based Defect Synthesis

ResearchDGX agent

arXiv:2604.22850v1 Announce Type: new Abstract: Industrial visual inspection systems often suffer from a severe scarcity of labeled defect data, particularly during the early stages of New Product Int

← Previous
1…167168169170171…209
Next →