AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
21 Apr 2026

HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models

ResearchDGX agent

arXiv:2508.00553v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) encode images and videos into abundant tokens, which contain substantial redundancy and computation cost. While visual

HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained Models

Model ReleasesDGX agent

arXiv:2604.16499v1 Announce Type: new Abstract: Black-box adversarial attack on vision-language pre-trained models is a practical and challenging task, as text and image perturbations need to be consi

HSG: Hyperbolic Scene Graph

TutorialsDGX agent

arXiv:2604.17454v1 Announce Type: new Abstract: Scene graph representations enable structured visual understanding by modeling objects and their relationships, and have been widely used for multiview


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Human Cognition in Machines: A Unified Perspective of World Models

AgentsDGX agent

arXiv:2604.16592v1 Announce Type: cross Abstract: This comprehensive report distinguishes prior works by the cognitive functions they innovate. Many works claim an almost 'human-like' cognitive capabi

Hybrid Multi-Dimensional MRI Prostate Cancer Detection via Hadamard Network-Based Bias Correction and Residual Networks

SafetyDGX agent

arXiv:2604.17107v1 Announce Type: new Abstract: Magnetic Resonance Imaging (MRI) is vital for prostate cancer (PCa) diagnosis. While advanced techniques such as Hybrid Multi-dimensional MRI (HM-MRI) h

Hybrid Quantum Neural Networks for Enhanced Breast Cancer Thermographic Classification: A Novel Quantum-Classical Integration Approach

ApplicationsDGX agent

arXiv:2604.16953v1 Announce Type: cross Abstract: Breast cancer diagnosis through thermographic image analysis remains a critical challenge in medical AI, with classical deep learning approaches facin

Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy

Model ReleasesDGX agent

arXiv:2510.22215v2 Announce Type: replace-cross Abstract: Retrieval over visually rich documents is essential for tasks such as legal discovery, scientific search, and enterprise knowledge management.

HyKey: Hyperspectral Keypoint Detection and Matching in Minimally Invasive Surgery

ResearchDGX agent

arXiv:2604.17446v1 Announce Type: new Abstract: Purpose: 3D reconstruction in minimally invasive surgery (MIS) enables enhanced surgical guidance through improved visualisation, tool tracking, and aug

Hyperbolic Enhanced Representation Learning for Incomplete Multi-view Clustering

SafetyDGX agent

arXiv:2604.16959v1 Announce Type: cross Abstract: Incomplete Multi-View Clustering (IMVC) faces the challenge of learning discriminative representations from fragmentary observations while maintaining

Hyperspectral Unmixing Hierarchies

ResearchDGX agent

arXiv:2604.16969v1 Announce Type: new Abstract: Unmixing reveals the spatial distribution and spectral details of different constituents, called endmembers, in a hyperspectral image. Because unmixing

ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models

Model ReleasesDGX agent

arXiv:2604.16405v1 Announce Type: cross Abstract: Video-generative world models are increasingly used as neural simulators for embodied planning and policy learning, yet their ability to predict physi

Identifying Ethical Biases in Action Recognition Models

SafetyDGX agent

arXiv:2604.17971v1 Announce Type: new Abstract: Human Action Recognition (HAR) models are increasingly deployed in high-stakes environments, yet their fairness across different human appearances has n

iDocV2: Leveraging Self-Supervision and Open-Set Detection for Improving Pattern Spotting in Historical Documents

Model ReleasesDGX agent

arXiv:2604.16726v1 Announce Type: new Abstract: Considering the imminent massification of digital books, it has become critical to facilitate searching collections through graphical patterns. Current

IMA-MoE: An Interpretable Modality-Aware Mixture-of-Experts Framework for Characterizing the Neurobiological Signatures of Binge Eating Disorder

ResearchDGX agent

arXiv:2604.17028v1 Announce Type: new Abstract: Binge eating disorder (BED) is the most prevalent eating disorder. However, current diagnostic frameworks remain largely grounded in symptom-based crite

Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback

SafetyDGX agent

arXiv:2412.02617v2 Announce Type: replace-cross Abstract: Large text-to-video models hold immense potential for a wide range of downstream applications. However, they struggle to accurately depict dyn

Improving Radio Interferometry Imaging by Explicitly Modeling Cross-Domain Consistency in Reconstruction

ResearchDGX agent

arXiv:2604.16794v1 Announce Type: new Abstract: Radio astronomy plays a crucial role in understanding the universe, particularly within the realm of non-thermal astrophysics. Images of celestial objec

IncepDeHazeGAN: Novel Satellite Image Dehazing

ResearchDGX agent

arXiv:2604.16609v1 Announce Type: new Abstract: Dehazing is a technique in computer vision for enhancing the visual quality of images captured in cloudy or foggy conditions. Dehazing helps to recover

Incoherent Deformation, Not Capacity: Diagnosing and Mitigating Overfitting in Dynamic Gaussian Splatting

Model ReleasesDGX agent

arXiv:2604.16747v1 Announce Type: new Abstract: Dynamic 3D Gaussian Splatting methods achieve strong training-view PSNR on monocular video but generalize poorly: on the D-NeRF benchmark we measure an

IncreFA: Breaking the Static Wall of Generative Model Attribution

Model ReleasesDGX agent

arXiv:2604.17736v1 Announce Type: new Abstract: As AI generative models evolve at unprecedented speed, image attribution has become a moving target. New diffusion, adversarial and autoregressive gener

Inductive Convolution Nuclear Norm Minimization for Tensor Completion with Arbitrary Sampling

ResearchDGX agent

arXiv:2604.17001v1 Announce Type: new Abstract: The recently established Convolution Nuclear Norm Minimization (CNNM) addresses the problem of extit{tensor completion with arbitrary sampling} (TCAS),

Inference-Time Temporal Probability Smoothing for Stable Video Segmentation with SAM2 under Weak Prompts

ResearchDGX agent

arXiv:2604.17115v1 Announce Type: new Abstract: Interactive video segmentation models such as SAM2 have demonstrated strong generalization across diverse visual domains. However, under weak user super

Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception

SafetyDGX agent

arXiv:2604.17651v1 Announce Type: new Abstract: World models, generative AI systems that simulate how environments evolve, are transforming autonomous driving, yet all existing approaches adopt an ego

Instant Colorization of Gaussian Splats

ResearchDGX agent

arXiv:2604.17155v1 Announce Type: new Abstract: Gaussian Splatting has recently become one of the most popular frameworks for photorealistic 3D scene reconstruction and rendering. While current raster

Instinct vs. Reflection: Unifying Token and Verbalized Confidence in Multimodal Large Models

SafetyDGX agent

arXiv:2604.17274v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in various perception and reasoning tasks. Despite this success, ens

Instruction-as-State: Environment-Guided and State-Conditioned Semantic Understanding for Embodied Navigation

AgentsDGX agent

arXiv:2604.18223v1 Announce Type: new Abstract: Vision-and-Language Navigation requires agents to follow natural-language instructions in visually changing environments. A central challenge is the dyn

Integrating Feature Selection and Machine Learning for Nitrogen Assessment in Grapevine Leaves using In-Field Hyperspectral Imaging

ApplicationsDGX agent

arXiv:2507.17869v3 Announce Type: replace-cross Abstract: Nitrogen (N) is one of the most critical nutrients in winegrape production, influencing vine vigor, fruit composition, and wine quality. Becau

INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2604.18051v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting o

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts

Model ReleasesDGX agent

arXiv:2509.10813v3 Announce Type: replace Abstract: The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts.

Is SAM3 ready for pathology segmentation?

ResearchDGX agent

arXiv:2604.18225v1 Announce Type: new Abstract: Is Segment Anything Model 3 (SAM3) capable in segmenting Any Pathology Images? Digital pathology segmentation spans tissue-level and nuclei-level scales

Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models

ResearchDGX agent

arXiv:2512.02636v3 Announce Type: replace-cross Abstract: Log-likelihood evaluation enables important capabilities in generative models, including model comparison, certain fine-tuning objectives, and

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription

Model ReleasesDGX agent

arXiv:2502.20295v2 Announce Type: replace-cross Abstract: Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical t

KaLDeX: Kalman Filter based Linear Deformable Cross Attention for Retina Vessel Segmentation

ResearchDGX agent

arXiv:2410.21160v2 Announce Type: replace-cross Abstract: Background and Objective: In the realm of ophthalmic imaging, accurate vascular segmentation is paramount for diagnosing and managing various

KIRA: Knowledge-Intensive Image Retrieval and Reasoning Architecture for Specialized Visual Domains

Model ReleasesDGX agent

arXiv:2604.16915v1 Announce Type: new Abstract: Retrieval augmented generation (RAG) has transformed text based question answering, yet its extension to visual domains remains hindered by fundamental

LAGS: Low-Altitude Gaussian Splatting with Groupwise Heterogeneous Graph Learning

ApplicationsDGX agent

arXiv:2604.16910v1 Announce Type: new Abstract: Low-altitude Gaussian splatting (LAGS) facilitates 3D scene reconstruction by aggregating aerial images from distributed drones. However, as LAGS priori

Latent-Compressed Variational Autoencoder for Video Diffusion Models

ResearchDGX agent

arXiv:2604.16479v1 Announce Type: new Abstract: Video variational autoencoders (VAEs) used in latent diffusion models typically require a sufficiently large number of latent channels to ensure high-qu

LayerCache: Exploiting Layer-wise Velocity Heterogeneity for Efficient Flow Matching Inference

Model ReleasesDGX agent

arXiv:2604.16492v1 Announce Type: new Abstract: Flow Matching models achieve state-of-the-art image generation quality but incur substantial inference cost due to iterative denoising through large Tra

LBFTI: Layer-Based Facial Template Inversion for Identity-Preserving Fine-Grained Face Reconstruction

ResearchDGX agent

arXiv:2604.18358v1 Announce Type: new Abstract: In face recognition systems, facial templates are widely adopted for identity authentication due to their compliance with the data minimization principl

Learned Nonlocal Feature Matching and Filtering for RAW Image Denoising

Model ReleasesDGX agent

arXiv:2604.17453v1 Announce Type: cross Abstract: Being one of the oldest and most basic problems in image processing, image denoising has seen a resurgence spurred by rapid advances in deep learning.

LiquidTAD: An Efficient Method for Temporal Action Detection via Liquid Neural Dynamics

Model ReleasesDGX agent

arXiv:2604.18274v1 Announce Type: new Abstract: Temporal Action Detection (TAD) in untrimmed videos is currently dominated by Transformer-based architectures. While high-performing, their quadratic co

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length

Model ReleasesDGX agent

arXiv:2512.04677v5 Announce Type: replace Abstract: Audio-driven avatar interaction demands real-time, streaming, and infinite-length generation -- capabilities fundamentally at odds with the sequenti

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing

Model ReleasesDGX agent

arXiv:2604.17021v1 Announce Type: new Abstract: Video editing aims to modify input videos according to user intent. Recently, end-to-end training methods have garnered widespread attention, constructi

LLaMA-XR: A Novel Framework for Radiology Report Generation using LLaMA and QLoRA Fine Tuning

Model ReleasesDGX agent

arXiv:2506.03178v2 Announce Type: replace-cross Abstract: Automated radiology report generation holds significant potential to reduce radiologists' workload and enhance diagnostic accuracy. However, g

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

ResearchDGX agent

arXiv:2501.05067v3 Announce Type: replace Abstract: In this paper, we introduce LLaVA-Octopus, a novel video multimodal large language model. LLaVA-Octopus adaptively weights features from different v

LLM as a Tool, Not an Agent: Code-Mined Tree Transformations for Neural Architecture Search

AgentsDGX agent

arXiv:2604.16555v1 Announce Type: cross Abstract: Neural Architecture Search (NAS) aims to automatically discover high-performing deep neural network (DNN) architectures. However, conventional algorit

LOD-Net: Locality-Aware 3D Object Detection Using Multi-Scale Transformer Network

Local AiDGX agent

arXiv:2604.16696v1 Announce Type: new Abstract: 3D object detection in point cloud data remains a challenging task due to the sparsity and lack of global structure inherent in the input. In this work,

Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation

Model ReleasesDGX agent

arXiv:2604.17428v1 Announce Type: new Abstract: As video generation models achieve unprecedented capabilities, the demand for robust video evaluation metrics becomes increasingly critical. Traditional

Long-Text-to-Image Generation via Compositional Prompt Decomposition

Model ReleasesDGX agent

arXiv:2604.18258v1 Announce Type: new Abstract: While modern text-to-image (T2I) models excel at generating images from intricate prompts, they struggle to capture the key details when the inputs are

LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation

AgentsDGX agent

arXiv:2604.17190v1 Announce Type: new Abstract: Aerial Vision-and-Language Navigation (Aerial VLN) enables unmanned aerial vehicles (UAVs) to follow natural language instructions and navigate complex

Lorentz Framework for Semantic Segmentation

TutorialsDGX agent

arXiv:2604.16836v1 Announce Type: new Abstract: Semantic segmentation in hyperbolic space enables compact modeling of hierarchical structure while providing inherent uncertainty quantification. Prior

Low Light Image Enhancement Challenge at NTIRE 2026

ResearchDGX agent

arXiv:2604.17669v1 Announce Type: new Abstract: This paper presents a comprehensive review of the NTIRE 2026 Low Light Image Enhancement Challenge, highlighting the proposed solutions and final result

Lumos3D: A Single-Forward Framework for Low-Light 3D Scene Restoration

Model ReleasesDGX agent

arXiv:2511.09818v2 Announce Type: replace Abstract: Restoring 3D scenes with low-light conditions is challenging, and most existing methods depend on precomputed camera poses and scene-specific optimi

MambaKick: Early Penalty Direction Prediction from HAR Embeddings

ApplicationsDGX agent

arXiv:2604.16588v1 Announce Type: new Abstract: Penalty kicks in soccer are decided under extreme time constraints, where goalkeepers benefit from anticipating shot direction from the kickers motion b

Mammo-FM: Breast-specific foundational model for Integrated Mammographic Diagnosis, Prognosis, and Reporting

SafetyDGX agent

arXiv:2512.00198v2 Announce Type: replace Abstract: Breast cancer is one of the leading causes of death among women worldwide. We introduce Mammo-FM, the first foundation model specifically for mammog

MARCO: Navigating the Unseen Space of Semantic Correspondence

Model ReleasesDGX agent

arXiv:2604.18267v1 Announce Type: new Abstract: Recent advances in semantic correspondence rely on dual-encoder architectures, combining DINOv2 with diffusion backbones. While accurate, these billion-

Marrying Text-to-Motion Generation with Skeleton-Based Action Recognition

Model ReleasesDGX agent

arXiv:2604.17090v1 Announce Type: new Abstract: Human action recognition and motion generation are two active research problems in human-centric computer vision, both aiming to align motion with textu

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems

Model ReleasesDGX agent

arXiv:2503.16549v2 Announce Type: replace Abstract: Despite strong results on many tasks, multimodal large language models (MLLMs) still underperform on visual mathematical problem solving, especially

Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis

SafetyDGX agent

arXiv:2411.17690v3 Announce Type: replace-cross Abstract: Unified decoder-only transformers have shown promise for multimodal generation, yet the mechanisms by which they synchronize modalities with h

Medial Axis Aware Learning of Signed Distance Functions

ResearchDGX agent

arXiv:2604.16512v1 Announce Type: new Abstract: We propose a novel variational method to compute a highly accurate global signed distance function (SDF) to a given point cloud. To this end, the jump s

Medical Image Understanding Improves Survival Prediction via Visual Instruction Tuning

Model ReleasesDGX agent

arXiv:2604.18250v1 Announce Type: new Abstract: Accurate prognostication and risk estimation are essential for guiding clinical decision-making and optimizing patient management. While radiologist-ass

MEDN: Motion-Emotion Feature Decoupling Network for Micro-Expression Recognition

Model ReleasesDGX agent

arXiv:2604.17899v1 Announce Type: new Abstract: Unlike macro-expression, micro-expression does not follow a strictly consistent mapping rule between emotions and Action Units (AUs). As a result, some

← Previous
1…182183184185186…209
Next →