AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
5 May 2026

Certified vs. Empirical Adversarial Robust-ness via Hybrid Convolutions with Attention Stochasticity

ResearchDGX agent

arXiv:2605.01519v1 Announce Type: new Abstract: We introduce Hybrid Convolutions with Attention Stochasticity (HyCAS), an adversarial defense that narrows the long-standing gap between provable robust

CEZSAR: A Contrastive Embedding Method for Zero-Shot Action Recognition

Model ReleasesDGX agent

arXiv:2605.01165v1 Announce Type: new Abstract: This paper proposes a novel Zero-Shot Action Recognition~(ZSAR) method based on contrastive learning. In ZSAR, we aim to classify examples from classes

CGFformer: Cluster-Guidance Frequency Transformer for Pansharpening

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.01490v1 Announce Type: new Abstract: Pansharpening aims to generate high-resolution multispectral (HRMS) images by fusing low-resolution multispectral (LRMS) images with high-resolution pan

Channel-Level Relation to Attentive Aggregation with Neighborhood-Homogeneity Constraint for Point Cloud Analysis

Model ReleasesDGX agent

arXiv:2605.02357v1 Announce Type: new Abstract: In 3D point cloud understanding, the core challenge lies in accurately capturing discriminative features within complex neighborhoods, which directly af

Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts

Model ReleasesDGX agent

arXiv:2605.01882v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown considerable potential in chart understanding and reasoning tasks. However, they still struggle with

CHASE: Competing Hypotheses for Ambiguity-Aware Selective Prediction

Local AiDGX agent

arXiv:2605.01346v1 Announce Type: new Abstract: Standard selective prediction methods typically estimate uncertainty from the output of a single predictive branch. While effective for general uncertai

Checkerboard: A Simple, Effective, Efficient and Learning-free Clean Label Backdoor Attack with Low Poisoning Budget

Model ReleasesDGX agent

arXiv:2605.01298v1 Announce Type: cross Abstract: Backdoor attacks threaten the deep learning supply chain by poisoning a small fraction of the training data so that a model behaves normally on clean

Closed-Loop LLM Discovery of Non-Standard Channel Priors in Vision Models

TutorialsDGX agent

arXiv:2601.08517v2 Announce Type: replace Abstract: Channel-configuration search, the optimization of layer specifications such as channel widths in deep neural networks, presents a combinatorial chal

CNN-based Multi-In-Multi-Out Model for Efficient Spatiotemporal Prediction

Model ReleasesDGX agent

arXiv:2605.01277v1 Announce Type: new Abstract: Recently, Convolutional Neural Network (CNN) or Transformer architecture based models have been proposed to overcome the limitations of Recurrent Neural

Colinearity Decay: Training Quantization-Friendly ViTs with Outlier Decay

SafetyDGX agent

arXiv:2605.01330v1 Announce Type: new Abstract: Low-bit quantization is a practical route for efficiently deploying vision Transformers, yet activation outliers complicate fully quantized deployment.

Combining Facial Videos and Biosignals for Stress Estimation During Driving

SafetyDGX agent

arXiv:2601.04376v3 Announce Type: replace Abstract: Reliable stress recognition is critical in applications such as medical monitoring and safety-critical systems, including real-world driving. While

Comparative Evaluation of Convolutional and Transformer-Based Detectors for Automated Weed Detection in Precision Agriculture

ResearchDGX agent

arXiv:2605.00908v1 Announce Type: new Abstract: This paper presents a comparative evaluation of convolutional and transformer-based object detection architectures for early weed detection in realistic

Compression as Adaptation: Implicit Visual Representation with Diffusion Foundation Models

ResearchDGX agent

arXiv:2603.07615v2 Announce Type: replace-cross Abstract: Modern visual generative models acquire rich visual knowledge through large-scale training, yet existing visual representations (such as pixel

Continual Few-shot Adaptation for Synthetic Fingerprint Detection

ResearchDGX agent

arXiv:2603.14632v2 Announce Type: replace Abstract: The quality and realism of synthetically generated fingerprint images have increased significantly over the past decade fueled by advancements in ge

Continuous quantification of viral plaque dynamics using ultra-large-area label-free imaging enables rapid antiviral susceptibility testing

ResearchDGX agent

arXiv:2605.01738v1 Announce Type: cross Abstract: The plaque reduction assay (PRA) remains the gold standard for antiviral susceptibility testing, evaluating drug potency by measuring reductions in pl

Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses

ResearchDGX agent

arXiv:2509.01453v2 Announce Type: replace Abstract: Images vary in how memorable they are to humans. Inspired by findings from cognitive science and computer vision, we explore correlates of image mem

Cross-Domain Adversarial Augmentation: Stabilizing GANs for Medical and Handwriting Data Scarcity

ResearchDGX agent

arXiv:2605.01815v1 Announce Type: new Abstract: Generative Adversarial Networks (GANs) offer a pragmatic route to mitigate data scarcity in vision tasks. We study generative augmentation across two lo

Cross-Language Learning within Arabic Script for Low-Resource HTR

Model ReleasesDGX agent

arXiv:2605.02089v1 Announce Type: new Abstract: Handwritten Text Recognition (HTR) under limited labeled data remains a challenging problem, particularly for Arabic-script languages. Although modern s

Cross-Polarization Fusion of VV AND VH SAR Observations for Improved Flood Mapping

ResearchDGX agent

arXiv:2605.02153v1 Announce Type: new Abstract: Synthetic Aperture Radar (SAR) imagery is widely used for flood monitoring due to its all-weather and day-night imaging capability. However, flood mappi

Cross-Scale Pretraining: Enhancing Self-Supervised Learning for Low-Resolution Satellite Imagery for Semantic Segmentation

TutorialsDGX agent

arXiv:2601.12964v2 Announce Type: replace Abstract: Self-supervised pretraining in remote sensing is mostly done using mid-spatial resolution (MR) image datasets due to their high availability. Given

CSGuard: Toward Forgery-Resistant Watermarking in Diffusion Models via Compressed Sensing Constraint

ResearchDGX agent

arXiv:2605.01479v1 Announce Type: new Abstract: Latent-based diffusion model watermarking embeds watermarks into generated images' latent space to enable content attribution, offering a training-free

CUE: Concept-Aware Multi-Label Expansion to Mitigate Concept Confusion in Long-Tailed Learning

SafetyDGX agent

arXiv:2605.01309v1 Announce Type: new Abstract: Long-tailed distributions are common in real-world recognition tasks, where a few head classes have many samples while most tail classes have very few.

Decision Boundary-aware Generation for Long-tailed Learning

SafetyDGX agent

arXiv:2605.01468v1 Announce Type: new Abstract: Long-tailed data bias decision boundaries toward head classes and degrade tail class accuracy. Diffusion-based generative augmentation address this prob

Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation

Model ReleasesDGX agent

arXiv:2605.01448v1 Announce Type: cross Abstract: Cross-task generalization is a core challenge in open-world robotic manipulation, and the key lies in extracting transferable manipulation knowledge f

Decouple and Cache: KV Cache Construction for Streaming Video Understanding

TutorialsDGX agent

arXiv:2605.01858v1 Announce Type: new Abstract: Streaming video understanding requires processing unbounded video streams with limited memory and computation, posing two key challenges. First, continu

Deep neural networks with Fisher vector encoding for medical image classification

Model ReleasesDGX agent

arXiv:2605.01667v1 Announce Type: new Abstract: Orderless encoding methods have shown to improve Convolutional Neural Networks (CNNs) for image classification in the context of limited availability of

Degradation-Aware Adaptive Context Gating for Unified Image Restoration

TutorialsDGX agent

arXiv:2605.01236v1 Announce Type: new Abstract: Unified image restoration using a single model often faces task interference due to diverse degradations. To address this, we propose DACG-IR (Degradati

Developing a Strong Pre-Trained Base Model for Plant Leaf Disease Classification

Model ReleasesDGX agent

arXiv:2605.01283v1 Announce Type: new Abstract: Plants, crops and their yields are essential to our very existence, but diseases and pests cause large losses every year. As such it is vital to ensure

DGS-Net: Distillation-Guided Gradient Surgery for CLIP Fine-Tuning in AI-Generated Image Detection

ResearchDGX agent

arXiv:2511.13108v3 Announce Type: replace Abstract: The rapid progress of generative models such as GANs and diffusion models has led to the widespread proliferation of AI-generated images, raising co

Dino-NestedUNet: Unlocking Foundation Vision Encoders for Pathology Tumor Bulk Segmentation via Dense Decoding

ResearchDGX agent

arXiv:2605.00894v1 Announce Type: new Abstract: Vision foundation models (VFMs), such as DINOv3, provide rich semantic representations that are promising for computational pathology. However, many cur

DIPLI: Deep Image Prior Lucky Imaging for Blind Astronomical Image Restoration

ApplicationsDGX agent

arXiv:2503.15984v3 Announce Type: replace Abstract: Modern image restoration and super-resolution methods utilize deep learning due to its superior performance compared to traditional algorithms. Howe

DirectEdit: Step-Level Accurate Inversion for Flow-Based Image Editing

ResearchDGX agent

arXiv:2605.02417v1 Announce Type: new Abstract: With recent advancements in large-scale pre-trained text-to-image (T2I) models, training-free image editing methods have demonstrated remarkable success

Disciplined Diffusion: Text-to-Image Diffusion Model against NSFW Generation

SafetyDGX agent

arXiv:2605.01113v1 Announce Type: new Abstract: Text-to-image (T2I) diffusion models have the ability to build high-quality pictures from text prompts, but they pose safety concerns because they can g

Disentangled Anatomy-Disease Diffusion (DADD) for Controllable Ulcerative Colitis Progression Synthesis

ResearchDGX agent

arXiv:2605.01848v1 Announce Type: new Abstract: Synthesizing longitudinal medical images at controllable disease stages while preserving patient-specific anatomy is hindered by the entanglement of pat

DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video Generation

ResearchDGX agent

arXiv:2411.14295v3 Announce Type: replace Abstract: Generating high-quality stereo videos requires consistent depth perception and temporal coherence across frames. Despite advances in image and video

Divergence is Uncertainty: A Closed-Form Posterior Covariance for Flow Matching

ResearchDGX agent

arXiv:2605.00941v1 Announce Type: cross Abstract: Flow matching has become a leading framework for generative modeling, but quantifying the uncertainty of its samples remains an open problem. Existing

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models

SafetyDGX agent

arXiv:2605.01896v1 Announce Type: new Abstract: Emerging multi-modal world models attempt to jointly generate videos across diverse modalities (e.g., RGB, depth, and mask), yet they fail to fully expl

Does it Really Count? Assessing Semantic Grounding in Text-Guided Class-Agnostic Counting

ApplicationsDGX agent

arXiv:2605.02752v1 Announce Type: new Abstract: Open-world text-guided class-agnostic counting (CAC) has emerged as a flexible paradigm for counting arbitrary object classes by using natural language

DP-SfM: Dual-Pixel Structure-from-Motion without Scale Ambiguity

ResearchDGX agent

arXiv:2605.01852v1 Announce Type: new Abstract: Multi-view 3D reconstruction, namely, structure-from-motion followed by multi-view stereo, is a fundamental component of 3D computer vision. In general,

Dual-branch Robust Unlearnable Examples

Model ReleasesDGX agent

arXiv:2605.01718v1 Announce Type: new Abstract: Unlearnable examples (UEs) aim to compromise model training by injecting imperceptible perturbations to clean samples. However, existing UE schemes exhi

DuFal: Dual-Frequency-Aware Learning for High-Fidelity Extremely Sparse-view CBCT Reconstruction

Model ReleasesDGX agent

arXiv:2601.15416v2 Announce Type: replace Abstract: Sparse-view Cone-Beam Computed Tomography reconstruction from limited X-ray projections remains a challenging problem in medical imaging due to the

DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving

AgentsDGX agent

arXiv:2603.19675v2 Announce Type: replace Abstract: Recently, world models have been incorporated into the autonomous driving systems to improve the planning reliability. Existing approaches typically

DynoSLAM: Dynamic SLAM with Generative Graph Neural Networks for Real-World Social Navigation

Local AiDGX agent

arXiv:2605.02759v1 Announce Type: cross Abstract: Traditional Simultaneous Localization and Mapping (SLAM) algorithms rely heavily on the static environment assumption, which severely limits their app

EAPFusion: Intrinsic Evolving Auxiliary Prior Guidance for Infrared and Visible Image Fusion

AgentsDGX agent

arXiv:2605.01916v1 Announce Type: new Abstract: Infrared-visible image fusion aims to create an information-rich fused image by integrating the complementary thermal saliency from infrared sensing and

ECG-biometrics-bench: A Unified Framework for Reproducible Benchmarking of ECG Biometrics

ApplicationsDGX agent

arXiv:2605.01548v1 Announce Type: cross Abstract: Electrocardiogram (ECG) biometrics have emerged as a promising modality for continuous, liveness-aware authentication in wearable systems. However, ma

Edge-Efficient Image Restoration: Transformer Distillation into State-Space Models

ResearchDGX agent

arXiv:2605.02794v1 Announce Type: new Abstract: We propose a modular framework for hybrid image restoration that integrates transformer and state-space model (SSM) blocks with a focus on improving run

EdgeLPR: On the Deep Neural Network trade-off between Precision and Performance in LiDAR Place Recognition

Model ReleasesDGX agent

arXiv:2605.02275v1 Announce Type: new Abstract: Place recognition is essential for long-term autonomous navigation, enabling loop closure and consistent mapping. Although deep learning has improved pe

EduGage: Methods and Dataset for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning

ResearchDGX agent

arXiv:2605.01238v1 Announce Type: cross Abstract: Engagement, which links to attentional, emotional, and cognitive dimensions, plays an important role in learning. In online and video-based learning e

Embody4D: A Generalist 4D World Model for Embodied AI

ResearchDGX agent

arXiv:2605.01799v1 Announce Type: new Abstract: World models have made significant progress in modeling dynamic environments; however, most embodied world models are still restricted to 2D representat

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness

Model ReleasesDGX agent

arXiv:2605.01024v1 Announce Type: new Abstract: Multimodal Emotion Recognition (MER) is critical for interpreting real-world interactions. While Multimodal Large Language Models (MLLM) have shown prom

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning

ResearchDGX agent

arXiv:2605.02378v1 Announce Type: new Abstract: In-context learning (ICL) allows large models to adapt to tasks using a few examples, yet its extension to vision-language models (VLMs) remains fragile

Evolving Token Communication with Parametric Memory Network

Model ReleasesDGX agent

arXiv:2605.01869v1 Announce Type: cross Abstract: Token communication has emerged as a promising framework for efficient wireless transmission by representing source data as compact semantic tokens. H

Exploring Data-Free LoRA Transferability for Video Diffusion Models

SafetyDGX agent

arXiv:2605.01929v1 Announce Type: new Abstract: Video diffusion models leveraging step distillation or causal distillation have achieved remarkable performance. However, adapting existing LoRAs to the

Exploring Entropy-based Active Learning for Fair Brain Segmentation

SafetyDGX agent

arXiv:2605.01706v1 Announce Type: new Abstract: Active learning (AL) has emerged as a crucial strategy for reducing the prohibitive costs associated with medical image segmentation. However, standard

Exploring Prompt Alignment with Clinical Factors in Zero-Shot Segmentation VLMs for NSCLC Tumor Segmentation

SafetyDGX agent

arXiv:2605.01266v1 Announce Type: new Abstract: Zero-shot vision-language models (VLMs) offer a promptable alternative to task-specific training for gross tumor volume (GTV) delineation in non-small-c

ExpoCM: Exposure-Aware One-Step Generative Single-Image HDR Reconstruction

SafetyDGX agent

arXiv:2605.02464v1 Announce Type: new Abstract: Single-image HDR reconstruction aims to recover high dynamic range radiance from a single low dynamic range (LDR) input, but remains highly ill-posed du

FEAT: Fashion Editing and Try-On from Any Design

ResearchDGX agent

arXiv:2605.02393v1 Announce Type: new Abstract: Fashion design aims to express a designer's creative intent and to depict how garments interact with the human body. Recent methods condition on multimo

Fine-Grained Class-Conditional Distribution Balancing for Debiased Learning

SafetyDGX agent

arXiv:2505.06831v2 Announce Type: replace Abstract: Achieving group-robust generalization in the presence of spurious correlations remains a significant challenge, particularly when bias annotations a

Fine-Tuning Impairs the Balancedness of Foundation Models in Long-tailed Personalized Federated Learning

Model ReleasesDGX agent

arXiv:2605.02247v1 Announce Type: new Abstract: Personalized federated learning (PFL) with foundation models has emerged as a promising paradigm enabling clients to adapt to heterogeneous data distrib

FLoRA: Fusion-Latent for Optical Reconstruction and Flood Area Segmentation via Cross-Modal Multi-Task Distillation Network

SafetyDGX agent

arXiv:2605.02137v1 Announce Type: new Abstract: Accurate flood water mapping is critical for disaster management, yet current methods struggle to fully exploit the potential of spaceborne imagery. Opt

← Previous
1…155156157158159…209
Next →