AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
19 May 2026

CT-DegradBench: A Physics-Informed Benchmark for CT Degradation Detection and Severity Estimation

Model ReleasesDGX agent

arXiv:2605.16431v1 Announce Type: new Abstract: Computed tomography (CT) images are frequently degraded by acquisition artifacts, including noise, blur, streaking, aliasing, and metal artifacts. Yet C

Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection

ResearchDGX agent

arXiv:2603.01993v2 Announce Type: replace Abstract: Recent advances in generative AI have significantly enhanced the realism of multimodal media manipulation, thereby posing substantial challenges to

Dance Across Shifts: Forward-Facilitation Continual Test-Time Adaptation through Dynamic Style Bridging

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.18608v1 Announce Type: new Abstract: Continual Test-Time Adaptation (CTTA) aims to empower perception systems to handle dynamic distribution shifts encountered after deployment. Existing me

DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos

ApplicationsDGX agent

arXiv:2605.18102v1 Announce Type: new Abstract: Monocular video human mesh recovery is essential for digital humans, avatar animation, and embodied simulation, where both temporal stability and expres

DASH: A Meta-Attack Framework for Synthesizing Effective and Stealthy Adversarial Examples

ResearchDGX agent

arXiv:2508.13309v3 Announce Type: replace Abstract: Numerous techniques have been proposed for generating adversarial examples in white-box settings under strict Lp-norm constraints. However, such nor

DECODE: Domain-aware Continual Domain Expansion for Motion Prediction

AgentsDGX agent

arXiv:2411.17917v2 Announce Type: replace Abstract: Motion prediction is critical for autonomous vehicles to effectively navigate complex environments and accurately anticipate the behaviors of other

DecoRec: Decomposed 3D Scene Reconstruction from Single-View Images via Object-Level Diffusion

ResearchDGX agent

arXiv:2605.16807v1 Announce Type: new Abstract: In this paper, we introduce extit{DecoRec}, a novel system designed to elevate single-view 2D images to a decomposed 3D scene mesh. Current methods for

Decoupling Motion and Geometry in 4D Gaussian Splatting

ResearchDGX agent

arXiv:2603.00952v2 Announce Type: replace Abstract: High-fidelity reconstruction of dynamic scenes is an important yet challenging problem. While recent 4D Gaussian Splatting (4DGS) has demonstrated t

Deep learning-based compression of giga-resolution whole slide images

ResearchDGX agent

arXiv:2605.17668v1 Announce Type: new Abstract: Implementation of digital pathology leads to an increased number of whole slide images (WSIs). The large size of WSIs is challenging. Today, WSIs are co

Deep Learning for MRI Slice Interpolation: The Critical Role of Problem Formulation

ResearchDGX agent

arXiv:2605.16476v1 Announce Type: cross Abstract: Through-plane resolution in clinical MRI is typically much coarser than in-plane resolution, limiting diagnostic utility. This work investigates deep

Deepfake Detection in Social Media: A Temporal Artifact Analysis Using 3D Convolutional Neural Networks

ResearchDGX agent

arXiv:2605.17573v1 Announce Type: new Abstract: Synthetic facial videos have proliferated across social media faster than platform moderation can respond, raising the cost of disinformation and identi

Degradation Frequency Curve: An Explicit Frequency-Quantified Representation for All-in-One Image Restoration

TutorialsDGX agent

arXiv:2605.17506v1 Announce Type: new Abstract: A fundamental difficulty in all-in-one blind image restoration is that degradation is usually treated as an implicit factor hidden in degraded-to-clean

DepthPolyp: Pseudo-Depth Guided Lightweight Segmentation for Real-Time Colonoscopy

Model ReleasesDGX agent

arXiv:2605.16519v1 Announce Type: new Abstract: Accurate polyp segmentation in colonoscopy is essential for early colorectal cancer detection, yet real-world clinical environments pose persistent chal

Designing streetscapes from street-view imagery using diffusion models

Model ReleasesDGX agent

arXiv:2605.17527v1 Announce Type: new Abstract: Street-view imagery (SVI) is widely used to quantify key indicators of urban environment, such as green- ery, sky, or road view indices. However, existi

DeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied Tracking

Model ReleasesDGX agent

arXiv:2605.17451v1 Announce Type: new Abstract: Aerial object tracking has broad applications in public safety, emergency rescue, wildlife monitoring, and related fields. However, existing aerial trac

DEVIS-GRPO: Unleashing GRPO on Dynamic Extreme View Synthesis

SafetyDGX agent

arXiv:2605.16937v1 Announce Type: new Abstract: Trajectory-controlled video generation has become essential for controllable video generation. While current methods perform well under small-view camer

Diffeomorphic Cortical Alignment via Direct Warping of Streamline Endpoints

Local AiDGX agent

arXiv:2605.16742v1 Announce Type: new Abstract: Cortical surface registration is often driven by local geometric descriptors (e.g., sulcal depth and curvature). While this approach achieves geometric

Diffusion-Based sRGB Real Noise Generation via Prompt-Driven Noise Representation Learning

Model ReleasesDGX agent

arXiv:2603.04870v2 Announce Type: replace Abstract: Denoising in the sRGB image space is challenging due to large noise variability. Although end-to-end methods perform well, their effectiveness in re

Diffusion Models, Denoiser Architecture and Creativity

SafetyDGX agent

arXiv:2605.16415v1 Announce Type: new Abstract: The creativity of diffusion models refers to their ability to generate highly realistic images that are different from their training data. Creativity i

DiffWind: Physics-Informed Differentiable Modeling of Wind-Driven Object Dynamics

ApplicationsDGX agent

arXiv:2603.09668v2 Announce Type: replace Abstract: Modeling wind-driven object dynamics from video observations is highly challenging due to the invisibility and spatio-temporal variability of wind,

DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers

SafetyDGX agent

arXiv:2605.16732v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art image generation quality but incur substantial memory and computational costs at inference. While

DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes

Model ReleasesDGX agent

arXiv:2601.13839v2 Announce Type: replace Abstract: Social media imagery provides a low-latency source of situational information during natural and human-induced disasters, enabling rapid damage asse

Distribution Prototype Diffusion Learning for Open-set Supervised Anomaly Detection

TutorialsDGX agent

arXiv:2502.20981v2 Announce Type: replace Abstract: In Open-set Supervised Anomaly Detection (OSAD), the existing methods typically generate pseudo anomalies to compensate for the scarcity of observed

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting

ResearchDGX agent

arXiv:2605.18173v1 Announce Type: new Abstract: End-to-end scene text spotting, which unifies text detection and recognition within a single framework, has witnessed remarkable progress driven by deep

DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing

SafetyDGX agent

arXiv:2605.16990v1 Announce Type: new Abstract: While 2D diffusion models have achieved remarkable success in identity-preserving personalization, extending this capability to 3D assets remains a sign

DriveSafer: End-to-End Autonomous Driving with Safety Guidance

Model ReleasesDGX agent

arXiv:2605.16737v1 Announce Type: cross Abstract: End-to-End (E2E) autonomous driving models have shown growing capability in recent years, with performance improving on increasingly challenging bench

DSAA: Dual-Stage Attribute Activation for Fine-grained Open Vocabulary Detection

Model ReleasesDGX agent

arXiv:2605.18023v1 Announce Type: new Abstract: Open-Vocabulary Object Detection (OVD) models break the limitations of closed-set detection, enabling the iden- tification of unseen categories through

Dual-Rate Diffusion: Accelerating diffusion models with an interleaved heavy-light network

ResearchDGX agent

arXiv:2605.18190v1 Announce Type: cross Abstract: Diffusion models achieve state-of-the-art generative performance but suffer from high computational costs during inference due to the repeated evaluat

EchoSR: Efficient Context Harnessing for Lightweight Image Super-Resolution

ResearchDGX agent

arXiv:2605.17470v1 Announce Type: new Abstract: Image super-resolution (SR) aims to reconstruct high-quality, high-resolution (HR) images from low-resolution (LR) inputs and plays a critical role in v

Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing

SafetyDGX agent

arXiv:2605.16951v1 Announce Type: new Abstract: A fundamental challenge in image editing lies in preserving spatial locality: edits should improve targeted content without inadvertently altering surro

Efficient 3D Content Reconstruction and Generation

HardwareDGX agent

arXiv:2605.18052v1 Announce Type: new Abstract: Automatic 3D content creation seeks to replace labor-intensive modeling and scanning pipelines with systems that can synthesize or recover 3D assets dir

Efficient Sparse-to-Dense Visual Localization via Compact Gaussian Scene Representation and Accelerated Dense Pose Estimation

ResearchDGX agent

arXiv:2605.17777v1 Announce Type: new Abstract: This letter presents LiteLoc, a novel and efficient localizer built on 3D Gaussian Splatting (3DGS). The previous state-of-the-art (SoTA) sparse-to-dens

Efficient Spatially-Variant Convolution via Differentiable Sparse Kernel Complex

ResearchDGX agent

arXiv:2512.04556v2 Announce Type: replace-cross Abstract: Image convolution with complex kernels is a fundamental operation in photography, scientific imaging, and animation effects, yet direct dense

EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos

Model ReleasesDGX agent

arXiv:2605.18734v1 Announce Type: new Abstract: Egocentric memory is widely used in embodied intelligence, but it may be insufficient for comprehensive spatial-temporal reasoning. Inspired by human re

EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation

ApplicationsDGX agent

arXiv:2605.18214v1 Announce Type: new Abstract: Collecting large-scale egocentric video datasets with dense spatial and temporal annotations is costly, slow, and often constrained by environmental bia

EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning

Model ReleasesDGX agent

arXiv:2605.17262v1 Announce Type: new Abstract: Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI a

EgoKit: Towards Unified Low-Cost Egocentric Data Collection with Heterogeneous Devices

Local AiDGX agent

arXiv:2605.16797v1 Announce Type: new Abstract: Egocentric video is increasingly used as a data source for robot learning, activity understanding, and embodied AI research, but collecting it at scale

Embedded ConvNet Ensembles: A Lightweight Approach to Recognize Arabic Handwritten Characters

ResearchDGX agent

arXiv:2605.18060v1 Announce Type: new Abstract: Arabic Handwritten Character Recognition (AHCR) has recently advanced significantly with deep Convolutional Neural Networks (ConvNets). However, many mo

Employing Vision-Language Models for Face Image Quality Assessment

Model ReleasesDGX agent

arXiv:2605.17489v1 Announce Type: new Abstract: Face Image Quality Assessment (FIQA) is a crucial control step in biometric pipelines. It ensures only reliable samples are processed to maintain system

Enhancing Event-based Object Detection with Monocular Normal Maps

AgentsDGX agent

arXiv:2508.02127v2 Announce Type: replace Abstract: Object detection in autonomous driving is frequently compromised by complex illumination. While event cameras offer a robust solution, they are susc

Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos

SafetyDGX agent

arXiv:2605.18233v1 Announce Type: new Abstract: Without incurring significant computational overhead, train-free long video generation aims to enable foundation video generation models to produce long

EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.17070v1 Announce Type: new Abstract: While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on que

Error-Decomposed Class-Conditional Fusion for Statistically Guaranteed Hard-Category Robust Perception

Model ReleasesDGX agent

arXiv:2605.17591v1 Announce Type: new Abstract: Aggregate object detection metrics inherently mask catastrophic and repeatable failures in operationally critical, long-tail minority classes. This pape

EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers

ResearchDGX agent

arXiv:2605.16745v1 Announce Type: new Abstract: This paper addresses the challenge of integrating 3D meshes as a native modality within Multimodal Large Language Models (MLLMs). Diffusion-based large

Evidence-Guided Unknown Rejection for High-Confidence Near-Known Unknowns

ResearchDGX agent

arXiv:2605.17818v1 Announce Type: new Abstract: Open-set recognition systems face a neglected failure mode: high-confidence near-known unknowns, which lie outside the known label set but are close eno

Expandable, Compressible, Mineable: Open-World Thermal Image Restoration

Model ReleasesDGX agent

arXiv:2605.16967v1 Announce Type: new Abstract: In open-world settings, thermal infrared (TIR) image degradations continuously emerge and evolve, while most existing all-in-one restoration methods are

Explaining Object Detectors via Collective Contribution of Pixels

ResearchDGX agent

arXiv:2412.00666v4 Announce Type: replace Abstract: Visual explanations for object detectors are crucial for enhancing their reliability. Object detectors identify and localize instances by assessing

extit{Don't Guess, Just Ask}: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification

Model ReleasesDGX agent

arXiv:2605.17531v1 Announce Type: new Abstract: Referring segmentation aims to segment the target objects in images or videos based on the textual query. Despite remarkable progress over the past year

Face inpainting with Identity Preserving Latent Diffusion Models

SafetyDGX agent

arXiv:2605.16696v1 Announce Type: new Abstract: Face inpainting techniques recover missing or occluded facial regions in a visually realistic manner, but preserving the identity in the final output re

Factorized Latent Dynamics for Video JEPA: An Empirical Study of Auxiliary Objectives

ResearchDGX agent

arXiv:2605.17165v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPA) are a promising framework for self-supervised video representation learning, yet the behavior of auxilia

Fast Kernel-Space Diffusion for Remote Sensing Pansharpening

Model ReleasesDGX agent

arXiv:2505.18991v3 Announce Type: replace Abstract: Pansharpening seeks to fuse high-resolution panchromatic (PAN) and low-resolution multispectral (LRMS) images into a single image with both fine spa

FG-TreeSeg: Flow-Guided Tree Crown Segmentation without Instance Annotations

ResearchDGX agent

arXiv:2602.00470v2 Announce Type: replace Abstract: Individual tree crown segmentation is an important task in remote sensing for forest biomass estimation and ecological monitoring. However, accurate

Flow Matching with Optimized Subclass Priors for Medical Image Augmentation

SafetyDGX agent

arXiv:2605.16469v1 Announce Type: cross Abstract: Rare diseases dominate the diagnostic challenge in medical imaging yet are severely underrepresented in clinical datasets, causing classifiers to fail

Forget-It-All: Multi-Concept Machine Unlearning via Concept-Aware Neuron Masking

ResearchDGX agent

arXiv:2601.06163v2 Announce Type: replace Abstract: The widespread adoption of text-to-image (T2I) diffusion models has raised concerns about their potential to generate copyrighted, inappropriate, or

Forget Many, Forget Right: Scalable and Precise Concept Unlearning in Diffusion Models

SafetyDGX agent

arXiv:2601.06162v4 Announce Type: replace-cross Abstract: Text-to-image diffusion models have achieved remarkable progress, yet their use raises copyright and misuse concerns, prompting research into

FrequencyBooster: Full-Frequency Modeling for High-Fidelity Pixel Diffusion

Local AiDGX agent

arXiv:2605.17759v1 Announce Type: new Abstract: To circumvent the inherent fidelity bottlenecks and optimization misalignment of VAE-based latent diffusion, pixel-space diffusion models have emerged a

From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models

Model ReleasesDGX agent

arXiv:2508.01608v2 Announce Type: replace Abstract: Image geolocalization, the task of identifying the geographic location depicted in an image, is important for applications in crisis response, digit

Functionalization via Structure Completion and Motion Rectification

ResearchDGX agent

arXiv:2605.18010v1 Announce Type: new Abstract: Acquisition and creation of 3D assets have been largely view- or appearance-driven. As a result, existing digital 3D models often lack the requisite str

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation

SafetyDGX agent

arXiv:2512.23180v3 Announce Type: replace Abstract: Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding

GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic Guidance

ResearchDGX agent

arXiv:2605.18252v1 Announce Type: new Abstract: We introduce GaussianZoom, a generative zoom-in 3D reconstruction system with an iterative progressive framework that combines geometry-consistent scene

← Previous
1…127128129130131…211
Next →