AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
15 Apr 2026

Dual-Modality Anchor-Guided Filtering for Test-time Prompt Tuning

Model ReleasesDGX agent

arXiv:2604.12403v1 Announce Type: new Abstract: Test-Time Prompt Tuning (TPT) adapts vision-language models using augmented views, but its effectiveness is hindered by the challenge of determining whi

EDGE-Shield: Efficient Denoising-staGE Shield for Violative Content Filtering via Scalable Reference-Based Matching

Model ReleasesDGX agent

arXiv:2604.06063v2 Announce Type: replace Abstract: The advent of Text-to-Image generative models poses significant risks of copyright violation and deepfake generation. Since the rapid proliferation

EigenCoin: sassanid coins classification based on Bhattacharyya distance

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.11932v1 Announce Type: new Abstract: Solving pattern recognition problems using imbalanced databases is a hot topic, which entices researchers to bring it into focus. Therefore, we consider

ELoG-GS: Dual-Branch Gaussian Splatting with Luminance-Guided Enhancement for Extreme Low-light 3D Reconstruction

Model ReleasesDGX agent

arXiv:2604.12592v1 Announce Type: new Abstract: This paper presents our approach to the NTIRE 2026 3D Restoration and Reconstruction Challenge (Track 1), which focuses on reconstructing high-quality 3

Evolution-Inspired Sample Competition for Deep Neural Network Optimization

SafetyDGX agent

arXiv:2604.12568v1 Announce Type: new Abstract: Conventional deep network training generally optimizes all samples under a largely uniform learning paradigm, without explicitly modeling the heterogene

Evolution of Optimization Methods: Algorithms, Scenarios, and Evaluations

ResearchDGX agent

arXiv:2604.12968v1 Announce Type: cross Abstract: Balancing convergence speed, generalization capability, and computational efficiency remains a core challenge in deep learning optimization. First-ord

Fall Risk and Gait Analysis in Community-Dwelling Older Adults using World-Spaced 3D Human Mesh Recovery

ResearchDGX agent

arXiv:2604.11961v1 Announce Type: new Abstract: Gait assessment is a key clinical indicator of fall risk and overall health in older adults. However, standard clinical practice is largely limited to s

Fragile Reconstruction: Adversarial Vulnerability of Reconstruction-Based Detectors for Diffusion-Generated Images

SafetyDGX agent

arXiv:2604.12781v1 Announce Type: new Abstract: Recently, detecting AI-generated images produced by diffusion-based models has attracted increasing attention due to their potential threat to safety. A

From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception

ResearchDGX agent

arXiv:2604.12508v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding, they frequently falter in fine

Fundus Image-based Glaucoma Screening via Retinal Knowledge-Oriented Dynamic Multi-Level Feature Integration

Model ReleasesDGX agent

arXiv:2604.12351v1 Announce Type: new Abstract: Automated diagnosis based on color fundus photography is essential for large-scale glaucoma screening. However, existing deep learning models are typica

Generative Anonymization in Event Streams

Model ReleasesDGX agent

arXiv:2604.12803v1 Announce Type: new Abstract: Neuromorphic vision sensors offer low latency and high dynamic range, but their deployment in public spaces raises severe data protection concerns. Rece

Generative Refinement Networks for Visual Synthesis

Model ReleasesDGX agent

arXiv:2604.13030v1 Announce Type: new Abstract: While diffusion models dominate the field of visual generation, they are computationally inefficient, applying a uniform computational effort regardless

Grasp in Gaussians: Fast Monocular Reconstruction of Dynamic Hand-Object Interactions

SafetyDGX agent

arXiv:2604.12929v1 Announce Type: new Abstract: We present Grasp in Gaussians (GraG), a fast and robust method for reconstructing dynamic 3D hand-object interactions from a single monocular video. Unl

GroupKAN: Efficient Kolmogorov-Arnold Networks via Grouped Spline Modeling

Model ReleasesDGX agent

arXiv:2511.05477v2 Announce Type: replace Abstract: Medical image segmentation demands models that achieve high accuracy while maintaining computational efficiency and clinical interpretability. While

GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality

Model ReleasesDGX agent

arXiv:2604.12315v1 Announce Type: new Abstract: Agricultural parcel extraction plays an important role in remote sensing-based agricultural monitoring, supporting parcel surveying, precision managemen

Habitat Classification from Ground-Level Imagery Using Deep Neural Networks

Local AiDGX agent

arXiv:2507.04017v3 Announce Type: replace Abstract: Habitat assessment at local scales -- critical for enhancing biodiversity and guiding conservation priorities -- often relies on expert field survey

Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting

AgentsDGX agent

arXiv:2604.12626v1 Announce Type: cross Abstract: Training embodied AI agents depends critically on the visual fidelity of simulation environments and the ability to model dynamic humans. Current simu

HTDC: Hesitation-Triggered Differential Calibration for Mitigating Hallucination in Large Vision-Language Models

ResearchDGX agent

arXiv:2604.12115v1 Announce Type: new Abstract: Large vision-language models (LVLMs) achieve strong multimodal performance, but still suffer from hallucinations caused by unstable visual grounding and

Hypergraph-State Collaborative Reasoning for Multi-Object Tracking

ResearchDGX agent

arXiv:2604.12665v1 Announce Type: new Abstract: Motion reasoning serves as the cornerstone of multi-object tracking (MOT), as it enables consistent association of targets across frames. However, exist

HyperLiDAR: Adaptive Post-Deployment LiDAR Segmentation via Hyperdimensional Computing

Local AiDGX agent

arXiv:2604.12331v1 Announce Type: new Abstract: LiDAR semantic segmentation plays a pivotal role in 3D scene understanding for edge applications such as autonomous driving. However, significant challe

Image-to-Image Translation Framework Embedded with Rotation Symmetry Priors

ResearchDGX agent

arXiv:2604.12805v1 Announce Type: new Abstract: Image-to-image translation (I2I) is a fundamental task in computer vision, focused on mapping an input image from a source domain to a corresponding ima

IMU: Influence-guided Machine Unlearning

ResearchDGX agent

arXiv:2508.01620v3 Announce Type: replace-cross Abstract: Machine Unlearning (MU) aims to selectively erase the influence of specific data points from pretrained models. However, most existing MU meth

INST-Align: Implicit Neural Alignment for Spatial Transcriptomics via Canonical Expression Fields

Model ReleasesDGX agent

arXiv:2604.12084v1 Announce Type: new Abstract: Spatial transcriptomics (ST) measures mRNA expression while preserving spatial organization, but multi-slice analysis faces two coupled difficulties: la

Label-Efficient Cross-Modality Generalization for Liver Segmentation in Multi-Phase MRI

ApplicationsDGX agent

arXiv:2510.04705v4 Announce Type: replace Abstract: Accurate liver segmentation in multi-phase MRI is vital for liver fibrosis assessment, yet labeled data is often scarce and unevenly distributed acr

Latent Chain-of-Thought World Modeling for End-to-End Driving

Model ReleasesDGX agent

arXiv:2512.10226v2 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models for autonomous driving explore inference-time reasoning as a way to improve driving performance and safet

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA

Model ReleasesDGX agent

arXiv:2509.10026v4 Announce Type: replace Abstract: As large vision language models (VLMs) advance, their capabilities in multilingual visual question answering (mVQA) have significantly improved. Cha

Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis

TutorialsDGX agent

arXiv:2604.12650v1 Announce Type: new Abstract: Existing deepfake detection research has primarily focused on scenarios where the manipulated subject is actively speaking, i.e., generating fabricated

LiveMoments: Reselected Key Photo Restoration in Live Photos via Reference-guided Diffusion

SafetyDGX agent

arXiv:2604.12286v1 Announce Type: new Abstract: Live Photo captures both a high-quality key photo and a short video clip to preserve the precious dynamics around the captured moment. While users may c

Lyra 2.0: Explorable Generative 3D Worlds

ResearchDGX agent

arXiv:2604.13036v1 Announce Type: new Abstract: Recent advances in video generation enable a new paradigm for 3D scene creation: generating camera-controlled videos that simulate scene walkthroughs, t

M3D-Stereo: A Multiple-Medium and Multiple-Degradation Dataset for Stereo Image Restoration

Model ReleasesDGX agent

arXiv:2604.12917v1 Announce Type: new Abstract: Image restoration under adverse conditions, such as underwater, haze or fog, and low-light environments, remains a highly challenging problem due to com

MedConcept: Unsupervised Concept Discovery for Interpretability in Medical VLMs

Model ReleasesDGX agent

arXiv:2604.11868v1 Announce Type: new Abstract: While medical Vision-Language models (VLMs) achieve strong performance on tasks such as tumor or organ segmentation and diagnosis prediction, their opaq

MedGS: Gaussian Splatting for Multi-Modal 3D Medical Imaging

ApplicationsDGX agent

arXiv:2509.16806v4 Announce Type: replace Abstract: Endoluminal endoscopic procedures are essential for diagnosing colorectal cancer and other severe conditions in the digestive tract, urogenital syst

Mema: Memory-Augmented Adapter for Enhanced Vision-Language Understanding

Model ReleasesDGX agent

arXiv:2603.00655v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable performance by aligning pretrained visual representations with the linguistic know

Mitigating Shortcut Learning via Feature Disentanglement in Medical Imaging: A Benchmark Study

Model ReleasesDGX agent

arXiv:2602.18502v2 Announce Type: replace Abstract: Although deep learning models in medical imaging often achieve excellent classification performance, they can rely on shortcut learning, exploiting

Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection

Model ReleasesDGX agent

arXiv:2604.12380v1 Announce Type: new Abstract: Camouflaged Object Detection (COD) aims to segment objects that blend seamlessly into complex backgrounds, with growing interest in exploiting additiona

Navigating the Accuracy-Size Trade-Off with Flexible Model Merging

ResearchDGX agent

arXiv:2505.23209v3 Announce Type: replace Abstract: Model merging has emerged as an efficient method to combine multiple single-task fine-tuned models. The merged model can enjoy multi-task capabiliti

NoisePrints: Distortion-Free Watermarks for Authorship in Private Diffusion Models

TutorialsDGX agent

arXiv:2510.13793v2 Announce Type: replace Abstract: With the rapid adoption of diffusion models for visual content generation, proving authorship and protecting copyright have become critical. This ch

Nucleus-Image: Sparse MoE for Image Generation

Model ReleasesDGX agent

arXiv:2604.12163v1 Announce Type: new Abstract: We present Nucleus-Image, a text-to-image generation model that establishes a new Pareto frontier in quality-versus-efficiency by matching or exceeding

OFA-Diffusion Compression: Compressing Diffusion Model in One-Shot Manner

Model ReleasesDGX agent

arXiv:2604.12668v1 Announce Type: new Abstract: The Diffusion Probabilistic Model (DPM) achieves remarkable performance in image generation, while its increasing parameter size and computational overh

OmniFood8K: Single-Image Nutrition Estimation via Hierarchical Frequency-Aligned Fusion

ResearchDGX agent

arXiv:2604.12356v1 Announce Type: new Abstract: Accurate estimation of food nutrition plays a vital role in promoting healthy dietary habits and personalized diet management. Most existing food datase

On Efficient Variants of Segment Anything Model: A Survey

ApplicationsDGX agent

arXiv:2410.04960v5 Announce Type: replace Abstract: The Segment Anything Model (SAM) is a foundational model for image segmentation tasks, known for its strong generalization across diverse applicatio

One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion

ApplicationsDGX agent

arXiv:2508.04559v3 Announce Type: replace Abstract: Recent diffusion-based approaches have made significant advances in image-based virtual try-on, enabling more realistic and end-to-end garment synth

One View Is Enough! Monocular Training for In-the-Wild Novel View Generation

ResearchDGX agent

arXiv:2603.23488v2 Announce Type: replace Abstract: Monocular novel-view synthesis has long required multi-view image pairs for supervision, limiting training data scale and diversity. We argue it is

PC-MIL: Decoupling Feature Resolution from Supervision Scale in Whole-Slide Learning

Local AiDGX agent

arXiv:2604.12100v1 Announce Type: new Abstract: Whole-slide image (WSI) classification in computational pathology is commonly formulated as slide-level Multiple Instance Learning (MIL) with a single g

PDF-GS: Progressive Distractor Filtering for Robust 3D Gaussian Splatting

ApplicationsDGX agent

arXiv:2604.12580v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have enabled impressive real-time photorealistic rendering. However, conventional training pipelines inh

Physics-Grounded Monocular Vehicle Distance Estimation Using Standardized License Plate Typography

SafetyDGX agent

arXiv:2604.12239v1 Announce Type: new Abstract: Accurate inter-vehicle distance estimation is a cornerstone of Advanced Driver Assistance Systems (ADAS) and autonomous driving. While LiDAR and radar p

Pi-HOC: Pairwise 3D Human-Object Contact Estimation

ApplicationsDGX agent

arXiv:2604.12923v1 Announce Type: new Abstract: Resolving real-world human-object interactions in images is a many-to-many challenge, in which disentangling fine-grained concurrent physical contact is

PianoFlow: Music-Aware Streaming Piano Motion Generation with Bimanual Coordination

ResearchDGX agent

arXiv:2604.12856v1 Announce Type: new Abstract: Audio-driven bimanual piano motion generation requires precise modeling of complex musical structures and dynamic cross-hand coordination. However, exis

Point Prompting: Counterfactual Tracking with Video Diffusion Models

ResearchDGX agent

arXiv:2510.11715v2 Announce Type: replace Abstract: Trackers and video generators solve closely related problems: the former analyze motion, while the latter synthesize it. We show that this connectio

Privacy-Preserving Structureless Visual Localization via Image Obfuscation

ResearchDGX agent

arXiv:2604.12068v1 Announce Type: new Abstract: Visual localization is the task of estimating the camera pose of an image relative to a scene representation. In practice, visual localization systems a

Probabilistic Feature Imputation and Uncertainty-Aware Multimodal Federated Aggregation

SafetyDGX agent

arXiv:2604.12970v1 Announce Type: cross Abstract: Multimodal federated learning enables privacy-preserving collaborative model training across healthcare institutions. However, a fundamental challenge

QMC-Net: Data-Aware Quantum Representations for Remote Sensing Image Classification

ResearchDGX agent

arXiv:2604.11817v1 Announce Type: cross Abstract: Hybrid quantum-classical models offer a promising route for learning from complex data; however, their application to multi-band remote sensing imager

Radar-Camera BEV Multi-Task Learning with Cross-Task Attention Bridge for Joint 3D Detection and Segmentation

AgentsDGX agent

arXiv:2604.12918v1 Announce Type: new Abstract: Bird's-eye-view (BEV) representations are the dominant paradigm for 3D perception in autonomous driving, providing a unified spatial canvas where detect

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.12371v1 Announce Type: new Abstract: We study typographic prompt injection attacks on vision-language models (VLMs), where adversarial text is rendered as images to bypass safety mechanisms

Redefining Quality Criteria and Distance-Aware Score Modeling for Image Editing Assessment

SafetyDGX agent

arXiv:2604.12175v1 Announce Type: new Abstract: Recent advances in image editing have heightened the need for reliable Image Editing Quality Assessment (IEQA). Unlike traditional methods, IEQA require

ReefMapGS: Enabling Large-Scale Underwater Reconstruction by Closing the Loop Between Multimodal SLAM and Gaussian Splatting

ResearchDGX agent

arXiv:2604.11992v1 Announce Type: cross Abstract: 3D Gaussian Splatting is a powerful visual representation, providing high-quality and efficient 3D scene reconstruction, but it is crucially dependent

Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models

SafetyDGX agent

arXiv:2604.12582v1 Announce Type: new Abstract: Recent Video Large Language Models (Video-LLMs) have demonstrated strong capability in video understanding, yet they still suffer from hallucinations. E

Representing 3D Faces with Learnable B-Spline Volumes

ResearchDGX agent

arXiv:2604.12894v1 Announce Type: new Abstract: We present CUBE (Control-based Unified B-spline Encoding), a new geometric representation for human faces that combines B-spline volumes with learned fe

Retrievals Can Be Detrimental: Unveiling the Backdoor Vulnerability of Retrieval-Augmented Diffusion Models

ResearchDGX agent

arXiv:2501.13340v4 Announce Type: replace Abstract: Diffusion models (DMs) have recently demonstrated remarkable generation capability. However, their training generally requires huge computational re

Risk-Calibrated Learning: Minimizing Fatal Errors in Medical AI

SafetyDGX agent

arXiv:2604.12693v1 Announce Type: new Abstract: Deep learning models often achieve expert-level accuracy in medical image classification but suffer from a critical flaw: semantic incoherence. These hi

← Previous
1…192193194195196…207
Next →