AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
11 May 2026

A Unified Framework for the Detection and Classification of Fatty Pancreas in Ultrasound Images

ResearchDGX agent

arXiv:2605.07466v1 Announce Type: new Abstract: Non-alcoholic fatty pancreas disease (NAFPD) is an underdiagnosed condition associated with metabolic syndrome, insulin resistance, and increased risk o

A Unified Measure-Theoretic View of Diffusion, Score-Based, and Flow Matching Generative Models

ResearchDGX agent

arXiv:2605.06829v1 Announce Type: cross Abstract: We survey continuous-time generative modeling methods based on transporting a simple reference distribution to a data distribution via stochastic or d

Adaptive Subspace Projection for Generative Personalization

SafetyDGX agent

arXiv:2605.07257v1 Announce Type: new Abstract: Generative personalization often suffers from the semantic collapsing problem (SCP), where a learned personalized concept overpowers the rest of the tex


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

AdpSplit: Error-Driven Adaptive Splitting for Faster Geometry Discovery in 3D Gaussian Splatting

ResearchDGX agent

arXiv:2605.06876v1 Announce Type: new Abstract: Adaptive density control in 3D Gaussian Splatting (3DGS) repeatedly grows the Gaussian population through fixed-cardinality random splitting to discover

Advancing Reliable Synthetic Video Detection: Insights from the SAFE Challenge

ResearchDGX agent

arXiv:2605.06912v1 Announce Type: new Abstract: The proliferation of generative video technologies has intensified the need for reliable methods to detect and characterize synthetic media. To address

AGA3DNet: Anatomy-Guided Gaussian Priors with Multi-view xLSTM for 3D Brain MRI Subtype Classification

ResearchDGX agent

arXiv:2605.07142v1 Announce Type: new Abstract: Accurate 3D brain MRI subtype classification benefits from both localized anatomical cues and long-range contextual reasoning. We present AGA3DNet, a re

AGILE: Hand-Object Interaction Reconstruction from Video via Agentic Generation

AgentsDGX agent

arXiv:2602.04672v3 Announce Type: replace Abstract: Reconstructing dynamic hand-object interactions from monocular videos is critical for dexterous manipulation data collection and creating realistic

Anisotropic Modality Align

SafetyDGX agent

arXiv:2605.07825v1 Announce Type: cross Abstract: Training multimodal large language models has long been limited by the scarcity of high-quality paired multimodal data. Recent studies show that the s

Aquatic Neuromorphic Optical Flow

ResearchDGX agent

arXiv:2605.07653v1 Announce Type: new Abstract: Underwater environments impose severe constraints on conventional imaging systems and demand solutions that balance high-quality sensing with strict res

AsyncEvGS: Asynchronous Event-Assisted Gaussian Splatting for Handheld Motion-Blurred Scenes

ResearchDGX agent

arXiv:2605.07192v1 Announce Type: new Abstract: 3D reconstruction methods such as 3D Gaussian Splatting (3DGS) and Neural Radiance Fields (NeRF) achieve impressive photorealism but fail when input ima

Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering

ResearchDGX agent

arXiv:2603.18636v2 Announce Type: replace Abstract: Diffusion Transformers (DiTs) achieve strong video generation quality but suffer from high inference cost due to dense 3D attention, motivating spar

Attention Transfer Is Not Universally Effective for Vision Transformers

Model ReleasesDGX agent

arXiv:2605.07191v1 Announce Type: new Abstract: A recent work shows that Attention Transfer, which transfers only the attention patterns from a pre-trained teacher Vision Transformer (ViT) to a random

AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language Models

ApplicationsDGX agent

arXiv:2605.07478v1 Announce Type: new Abstract: Speech-driven facial animation requires accurate correspondence between acoustic signals and facial motion, especially for articulation-related mouth mo

Benchmarking Foundation Models for Renal Lesion Stratification in CT

Model ReleasesDGX agent

arXiv:2605.07749v1 Announce Type: new Abstract: The rapid proliferation of open-source medical foundation models (FMs) raises a practical question: how well do their pre-trained representations transf

Beyond Defenses: Manifold-Aligned Regularization for Intrinsic 3D Point Cloud Robustness

ResearchDGX agent

arXiv:2605.07590v1 Announce Type: new Abstract: Despite extensive progress in point cloud robustness, existing methods primarily improve performance through augmentation or defense mechanisms, while o

Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs

Model ReleasesDGX agent

arXiv:2605.07562v1 Announce Type: new Abstract: Remote sensing vision-language models (RS-VLMs) face a fundamental mismatch with natural-image counterparts: the same geographic object exhibits radical

Breaking Spatial Uniformity: Prior-Guided Mamba with Radial Serialization for Lens Flare Removal

Model ReleasesDGX agent

arXiv:2605.07650v1 Announce Type: new Abstract: Lens flares, caused by complex optical aberrations, severely degrade image quality especially in nighttime photography. Although recent restoration meth

BRIDGE: Background Routing and Isolated Discrete Gating for Coarse-Mask Local Editing

Model ReleasesDGX agent

arXiv:2605.07846v1 Announce Type: new Abstract: Coarse-mask local image editing asks a model to modify a user-indicated region while preserving the surrounding scene. In practice, however, rough masks

Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment

TutorialsDGX agent

arXiv:2605.06969v1 Announce Type: new Abstract: Infrared-Visible image fusion (IVIF) aims to integrate thermal information and detailed spatial structures into a single fused image to enhance percepti

CalexNet: Soft Cascade-Aligned Training and Calibration for Lightweight Early-Exit Branches

SafetyDGX agent

arXiv:2509.08318v2 Announce Type: replace Abstract: Early-exit cascades over a frozen convolutional backbone enable adaptive inference but suffer from three sources of train-inference mismatch: branch

Clinically Aware Synthetic Image Generation for Concept Coverage in Chest X-ray Models

Model ReleasesDGX agent

arXiv:2603.15525v2 Announce Type: replace Abstract: Deep learning models for chest X-ray diagnosis are constrained by limited coverage of clinically meaningful concept combinations in publicly availab

Cloud-top infrared observations reveal the four-dimensional precipitation structure

ResearchDGX agent

arXiv:2605.07499v1 Announce Type: new Abstract: Accurate four-dimensional (4D) precipitation information is essential for understanding the Earth's energy and water cycles, yet remains observationally

CONSIGN: Conformal Segmentation Informed by Spatial Groupings via Decomposition

ResearchDGX agent

arXiv:2505.14113v3 Announce Type: replace Abstract: Most machine learning-based image segmentation models produce pixel-wise confidence scores that represent the model's predicted probability for each

Consistency Regularised Gradient Flows for Inverse Problems

ResearchDGX agent

arXiv:2605.07907v1 Announce Type: cross Abstract: Vision-Language Latent Diffusion Models (LDMs) (Rombach et al., 2022) provide powerful generative priors for inverse problems. However, existing LDM-b

Contrast-X: A Multi-Modal Contrast Image Synthesis Benchmark and Universal Modality Flow Matching

Model ReleasesDGX agent

arXiv:2601.15884v2 Announce Type: replace Abstract: Contrast-enhanced imaging is central to oncologic diagnosis, but contrast agents can be contraindicated for many of the patients who need them most.

DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection

Local AiDGX agent

arXiv:2604.02753v2 Announce Type: replace Abstract: Open-vocabulary Object Detection (OVOD) enables models to recognize objects beyond predefined categories, but existing approaches remain limited in

Decoupling Semantics and Fingerprints: A Universal Representation for AI-Generated Image Detection

SafetyDGX agent

arXiv:2605.07074v1 Announce Type: new Abstract: Detecting AI-generated images across unseen architectures remains challenging, as existing models often overfit to generator-specific fingerprints and s

DeepFedNAS: Efficient Hardware-Aware Architecture Adaptation for Heterogeneous IoT Federations via Pareto-Guided Supernet Training

HardwareDGX agent

arXiv:2601.15127v3 Announce Type: replace-cross Abstract: Deploying federated learning across heterogeneous IoT device fleets requires tailored neural network architectures for each device class, yet

Deeply Dual Supervised learning for melanoma recognition

Model ReleasesDGX agent

arXiv:2508.01994v2 Announce Type: replace Abstract: As the application of deep learning in dermatology continues to grow, the recognition of melanoma has garnered significant attention, demonstrating

Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision

TutorialsDGX agent

arXiv:2605.07940v1 Announce Type: new Abstract: Exemplar-based image editing applies a transformation defined by a source-target image pair to a new query image. Existing methods rely on a pair-of-pai

Differentiable Ray Tracing with Gaussians for Unified Radio Propagation Simulation and View Synthesis

ResearchDGX agent

arXiv:2605.07781v1 Announce Type: new Abstract: Explicit neural representations such as 3D Gaussian Splatting (3DGS) enable high-fidelity and real-time novel view synthesis, yet optimize for alpha-com

Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers

SafetyDGX agent

arXiv:2605.07503v1 Announce Type: new Abstract: Efficiently aligning large-scale video diffusion models with human intent requires a scalable and trajectory-aware pathway that bridges the inherent dis

DIMoE-Adapters: Dynamic Expert Evolution for Continual Learning in Vision-Language Models

ResearchDGX agent

arXiv:2605.07494v1 Announce Type: new Abstract: Continual learning enables vision-language models to accumulate knowledge and adapt to evolving tasks without retraining from scratch. However, in multi

DINO-MVR: Multi-View Readout of Frozen DINOv3 for Annotation-Efficient Medical Segmentation

ResearchDGX agent

arXiv:2605.07221v1 Announce Type: new Abstract: Adapting foundation models to medical segmentation typically requires either backbone fine-tuning or high-capacity task-specific decoders, both of which

Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2508.20909v2 Announce Type: replace Abstract: Foundation models pre-trained on large-scale natural image datasets offer a powerful paradigm for medical image segmentation. However, effectively t

Disambiguating 2D-3D Correspondences in Gaussian Splatting-based Feature Fields for Visual Localization

ResearchDGX agent

arXiv:2605.07351v1 Announce Type: new Abstract: While Gaussian Splatting-based Feature Fields (GSFFs) have shown promise for visual localization, this paper highlights that photometrically optimized G

DKDS: A Benchmark Dataset of Degraded Kuzushiji Documents with Seals for Detection and Binarization

Model ReleasesDGX agent

arXiv:2511.09117v4 Announce Type: replace Abstract: Kuzushiji, a pre-modern Japanese cursive script, can currently be read and understood by only a few thousand trained experts in Japan. With the rapi

Dr-BA: Separable Optimization for Direct Radar Bundle Adjustment & Localization

SafetyDGX agent

arXiv:2605.07041v1 Announce Type: cross Abstract: This paper introduces Dr-BA, a first-of-its-kind radar bundle adjustment (BA) framework that operates directly on 2D spinning radar intensity images.

DualResolution Residual Architecture with Artifact Suppression for Melanocytic Lesion Segmentation

Local AiDGX agent

arXiv:2508.06816v3 Announce Type: replace Abstract: Lesion segmentation, in contrast to natural scene segmentation, requires handling subtle variations in texture and color, frequent imaging artifacts

DVD: Discrete Voxel Diffusion for 3D Generation and Editing

ResearchDGX agent

arXiv:2605.07971v1 Announce Type: new Abstract: We introduce Discrete Voxel Diffusion (DVD), a discrete diffusion framework to generate, assess, and edit sparse voxels for SLat (Structured LATent) bas

Dynamic Mode Decomposition along Depth in Vision Transformers

AgentsDGX agent

arXiv:2605.07556v1 Announce Type: new Abstract: Recent work has shown that contiguous vision transformer (ViT) blocks (a) can be replaced by a linear map and (b) organize into recurrent phases of comp

Edge Detection for Organ Boundaries via Top Down Refinement and SubPixel Upsampling

ResearchDGX agent

arXiv:2508.06805v2 Announce Type: replace Abstract: Accurate localization of organ boundaries is critical in medical imaging for segmentation, registration, surgical planning, and radiotherapy. While

EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement

Local AiDGX agent

arXiv:2605.07457v1 Announce Type: new Abstract: Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as

EditTransfer++: Toward Faithful and Efficient Visual-Prompt-Guided Image Editing

SafetyDGX agent

arXiv:2605.07455v1 Announce Type: new Abstract: Visual-prompt-guided edit transfer aims to learn image transformations directly from example pairs, offering more precise and controllable editing than

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting

ResearchDGX agent

arXiv:2605.07642v1 Announce Type: new Abstract: Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such a

Enhancing Eye Movement Biometrics for User Authentication via Continuous Gaze Offset Score Fusion

ResearchDGX agent

arXiv:2605.06810v1 Announce Type: cross Abstract: Eye movement biometrics (EMB) use subject-specific gaze dynamics for user authentication and identification. Recent deep learning-based EMB systems ac

Enhancing Federated Quadruplet Learning: Stochastic Client Selection and Embedding Stability Analysis

ResearchDGX agent

arXiv:2605.07888v1 Announce Type: cross Abstract: Federated Learning (FL) enables decentralised model training across distributed clients without requiring data centralisation. However, the generalisa

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency

Local AiDGX agent

arXiv:2605.06280v2 Announce Type: replace Abstract: Recent advancements in image animation have utilized diffusion models to breathe life into static images. However, existing controllable frameworks

Explainable Part-Based Vehicle Classifier with Spatial Awareness

ResearchDGX agent

arXiv:2605.07831v1 Announce Type: new Abstract: In the area of Intelligent Transportation Systems (ITS), fine-grained vehicle classification systems play an essential role. Recently, the authors have

EyeCue: Driver Cognitive Distraction Detection via Gaze-Empowered Egocentric Video Understanding

ResearchDGX agent

arXiv:2605.07859v1 Announce Type: new Abstract: Driver cognitive distraction is a major cause of road collisions and remains difficult to detect. Unlike manual or visual distraction, cognitive distrac

Fine-tuning a vision-language model for fracture-surface morphology recognition

Model ReleasesDGX agent

arXiv:2605.07145v1 Announce Type: cross Abstract: Vision-language models (VLMs) have shown strong potential for scientific image understanding, but general-purpose models often lack the domain-specifi

Flatness and Gradient Alignment Are Both Necessary: Spectral-Aware Gradient-Aligned Exploration for Multi-Distribution Learning

SafetyDGX agent

arXiv:2605.07914v1 Announce Type: cross Abstract: Sharpness-aware and gradient-alignment methods have been shown to improve generalization, however each family of methods targets a single geometric pr

From Pixels to Primitives: Scene Change Detection in 3D Gaussian Splatting

TutorialsDGX agent

arXiv:2605.07203v1 Announce Type: new Abstract: Scene change detection methods built on Gaussian splatting universally follow a render-then-compare paradigm: the pre-change scene is rendered into 2D a

From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data

Model ReleasesDGX agent

arXiv:2605.07861v1 Announce Type: new Abstract: Makeup transfer aims to apply the makeup style of a reference portrait to a source portrait while preserving identity and background. Early methods form

Frozen Backpropagation: Relaxing Weight Symmetry in Deep Spiking Neural Networks

HardwareDGX agent

arXiv:2505.13741v2 Announce Type: replace Abstract: Direct training of Spiking Neural Networks (SNNs) on neuromorphic hardware can greatly reduce energy costs compared to GPU-based training. However,

FS-I2P:A Hierarchical Focus-Sweep Registration Network with Dynamically Allocated Depth

ResearchDGX agent

arXiv:2605.07607v1 Announce Type: new Abstract: Image-to-point cloud registration is often challenged by viewpoint changes, cross-modal discrepancies, and repetitive textures, which induce scale ambig

GC-ART: Global Learnable Second-Order Rational Tone Curves for Illumination Robustness

Model ReleasesDGX agent

arXiv:2605.07329v1 Announce Type: new Abstract: We introduce GC-ART (Global Curve Adaptive Rational Tone-mapping), a lightweight differentiable pre-processing module for robust image classification. G

GEM: Generating LiDAR World Model via Deformable Mamba

AgentsDGX agent

arXiv:2605.07326v1 Announce Type: new Abstract: World models, which simulate environmental dynamics and generate sensor observations, are gaining increasing attention in autonomous driving. However, p

GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization

SafetyDGX agent

arXiv:2605.07399v1 Announce Type: new Abstract: Diffusion Vision-Language Models (dVLMs), built upon the non-causal foundations of Diffusion Large Language Models (dLLMs), have demonstrated remarkable

GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval

SafetyDGX agent

arXiv:2509.23370v2 Announce Type: replace Abstract: The CLIP model has established itself as a cornerstone of large-scale retrieval systems. However, its performance often degrades under distributiona

← Previous
1…147148149150151…209
Next →