AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
18 May 2026

Community-aware evaluation and threshold calibration for open-set plankton image recognition

ResearchDGX agent

arXiv:2605.15835v1 Announce Type: new Abstract: Automated plankton image recognition is increasingly used in aquatic ecosystem monitoring, but deployed classifiers inevitably encounter unseen taxa and

COPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection

Model ReleasesDGX agent

arXiv:2605.15325v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong performance in video anomaly detection (VAD) while providing interpretable predictions. However, existin

Cross-Modal Registration Between 3D and 2D Fingerprints via Pose-Aware Unwrapping and Point-Cloud Fusion

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.15796v1 Announce Type: new Abstract: Three-dimensional (3D) fingerprints preserve global finger geometry and local ridge structure while avoiding contact-induced deformation, but they remai

DealMaTe: Multi-Dimensional Material Transfer via Diffusion Transformer

ResearchDGX agent

arXiv:2605.15681v1 Announce Type: cross Abstract: Recently, diffusion-based material transfer methods rely on image fine-tuning or complex architectures with auxiliary networks but face challenges suc

Decentralized LoRA augmented transformer with multi-scale feature learning for secured eye diagnosis

Model ReleasesDGX agent

arXiv:2505.06982v3 Announce Type: replace Abstract: Accurate and privacy-preserving diagnosis of ophthalmic diseases remains a critical challenge in medical imaging, particularly given the limitations

Deep Pre-Alignment for VLMs

Model ReleasesDGX agent

arXiv:2605.15300v1 Announce Type: new Abstract: Most Vision Language Models (VLMs) directly map outputs from ViT encoders to the LLM via a lightweight projector. While effective, recent analysis sugge

Degradation-Aware Blur-Segmentation of Brain Tumor

ResearchDGX agent

arXiv:2605.15671v1 Announce Type: cross Abstract: Multimodal 3D MRI brain tumor segmentation is a pivotal step in radiotherapy target delineation, surgical planning and post-treatment assessment. Exis

DIPA: Distilled Preconditioned Algorithms for Solving Imaging Inverse Problems

ResearchDGX agent

arXiv:2605.15456v1 Announce Type: cross Abstract: Solving imaging inverse problems has usually been addressed by designing proper prior models of the underlying signal. However, minimizing the data fi

Discretizing Group-Convolutional Neural Networks for 3D Geometry in Feature Space

SafetyDGX agent

arXiv:2605.15368v1 Announce Type: new Abstract: Group-convolutional neural networks (GCNNs) are among the most important methods for introducing symmetry as an inductive bias in deep learning: In each

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?

SafetyDGX agent

arXiv:2605.15855v1 Announce Type: new Abstract: Despite strong image-generation performance, diffusion models' reconstruction objectives limit alignment with human preferences. RL enables such alignme

DreamSR: Towards Ultra-High-Resolution Image Super-Resolution via a Receptive-Field Enhanced Diffusion Transformer

Local AiDGX agent

arXiv:2605.15682v1 Announce Type: new Abstract: Large-scale pre-trained diffusion models have been extensively adopted for real-world image Super-Resolution because of their powerful generative priors

DualReg: Dual-Space Filtering and Reinforcement for Rigid Registration

SafetyDGX agent

arXiv:2508.17034v2 Announce Type: cross Abstract: Noisy, partially overlapping data and the need for real-time processing pose major challenges for rigid registration. Considering that feature-based m

Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation

Model ReleasesDGX agent

arXiv:2605.16003v1 Announce Type: new Abstract: Autoregressive video diffusion models enable open-ended generation through local attention and KV caching. However, existing training-free long-video op

EduVQA: Towards Concept-Aware Assessment of Educational AI-Generated Videos

Model ReleasesDGX agent

arXiv:2603.03066v2 Announce Type: replace Abstract: Existing AI-generated video quality assessment (AIGVQA) methods mainly focus on global perceptual realism and coarse text-video alignment, while ove

Efficient Image Synthesis with Sphere Latent Encoder

ResearchDGX agent

arXiv:2605.15592v1 Announce Type: new Abstract: Few-step image generation has seen rapid progress, with consistency and meanflow-based methods significantly reducing the number of sampling steps. Desp

EgoExo-WM: Unlocking Exo Video for Ego World Models

SafetyDGX agent

arXiv:2605.15477v1 Announce Type: new Abstract: Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited avail

ElasticDiT: Efficient Diffusion Transformers via Elastic Architecture and Sparse Attention for High-Resolution Image Generation on Mobile Devices

SafetyDGX agent

arXiv:2605.15684v1 Announce Type: new Abstract: The Diffusion Transformer (DiT) architecture is the state-of-the-art paradigm for high-fidelity image generation, underpinning models like Stable Diffus

ELDOR: A Dataset and Benchmark for Illegal Gold Mining in the Amazon Rainforest

Model ReleasesDGX agent

arXiv:2605.15397v1 Announce Type: new Abstract: Illegal gold mining in the Amazon rainforest causes deforestation, water contamination, and long-term ecosystem disruption, yet remains difficult to mon

Embedding-perturbed Exploration Preference Optimization for Flow Models

SafetyDGX agent

arXiv:2605.15803v1 Announce Type: new Abstract: Recent advancements have established Reinforcement Learning (RL) as a pivotal paradigm for aligning generative models with human intent. However, group-

End-to-end plaque counting and virus titration from laboratory plate images with deep learning

Model ReleasesDGX agent

arXiv:2605.16008v1 Announce Type: new Abstract: Plaque assays remain the gold standard readout of virus infectivity; however, plaque counting from plate images is labor-intensive and prone to inter-op

EndoGSim: Physics-Aware 4D Dynamic Endoscopic Scene Simulations via MLLM-Guided Gaussian Splatting

ResearchDGX agent

arXiv:2605.16022v1 Announce Type: new Abstract: In robot-assisted minimally invasive surgery, high-fidelity dynamic endoscopic scene reconstruction and simulation are crucial to enhancing downstream t

Enhancing Medical Image Segmentation via Heat Conduction Equation

ResearchDGX agent

arXiv:2511.03260v2 Announce Type: replace Abstract: Medical image segmentation models struggle to achieve efficient global context modeling and long-range dependency reasoning under practical computat

Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction

Model ReleasesDGX agent

arXiv:2605.15466v1 Announce Type: new Abstract: Learning predictive world models from unlabelled video is a foundational challenge in artificial intelligence. While Joint Embedding Predictive Architec

EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy

SafetyDGX agent

arXiv:2605.15711v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to backdoor attacks. Exi

Evaluation of Anatomical Shape Priors in Deep Learning-Based Cardiac Multi-Compartment Segmentation

ResearchDGX agent

arXiv:2605.15707v1 Announce Type: cross Abstract: Whole-heart multi-compartment CT segmentation is clinically important, but standard CNNs do not explicitly enforce anatomical plausibility. Based on s

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization

HardwareDGX agent

arXiv:2605.15824v1 Announce Type: new Abstract: Human-centric video customization, particularly at the garment level, has shown significant commercial value. However, existing approaches cannot suppor

FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar Reconstruction

Model ReleasesDGX agent

arXiv:2605.15320v1 Announce Type: cross Abstract: Avatar reconstruction has traditionally relied on per-subject optimization that requires hours of computation or on expensive preprocessing that limit

FLASH: Efficient Visuomotor Policy via Sparse Sampling

SafetyDGX agent

arXiv:2605.15492v1 Announce Type: cross Abstract: Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative d

Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization

Model ReleasesDGX agent

arXiv:2605.15980v1 Announce Type: new Abstract: Group Relative Policy Optimization has emerged as essential for aligning video diffusion models with human preferences, but faces a critical computation

Frequency-domain Event-based Imaging for Selective Surveillance

ResearchDGX agent

arXiv:2605.15392v1 Announce Type: cross Abstract: Event-based cameras (EBCs) are an attractive sensing modality for surveillance due to their reporting of pixel-level radiance changes with microsecond

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding

ResearchDGX agent

arXiv:2605.15951v1 Announce Type: new Abstract: Finetuning Large Vision-Language Models with reinforcement learning has emerged as a promising approach to enhance their capability in object-level grou

From Full and Partial Intraoral Scans to Crown Proposal: A Classification-Guided Restoration Assistance Pipeline

SafetyDGX agent

arXiv:2605.15241v1 Announce Type: cross Abstract: Single-unit crown restoration is among the most common procedures in clinical dentistry, with CAD/CAM workflows now designing crowns directly from int

From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection

SafetyDGX agent

arXiv:2602.20630v4 Announce Type: replace Abstract: Keypoint-based matching is a fundamental component of modern 3D vision systems, such as Structure-from-Motion (SfM) and SLAM. Most existing learning

GHOST: Geometry-Hierarchical Online Streaming Token Eviction for Efficient 3D Reconstruction

ResearchDGX agent

arXiv:2605.15852v1 Announce Type: new Abstract: Streaming 3D reconstruction from long monocular video sequences requires maintaining a key-value (KV) cache that grows linearly with sequence length, cr

GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective

SafetyDGX agent

arXiv:2605.15723v1 Announce Type: cross Abstract: Multimodal alignment is commonly learned from isolated image-text pairs via CLIP-style dual encoders, leaving the relational context among entities la

Grounded Reinforcement Learning for Visual Reasoning

Local AiDGX agent

arXiv:2505.23678v3 Announce Type: replace Abstract: While reinforcement learning (RL) over chains of thought has significantly advanced language models in tasks such as mathematics and coding, visual

Hestia: Voxel-Face-Aware Hierarchical Next-Best-View Acquisition for Efficient 3D Reconstruction

ApplicationsDGX agent

arXiv:2508.01014v4 Announce Type: replace-cross Abstract: Advances in 3D reconstruction and novel view synthesis have enabled efficient and photorealistic rendering. However, images for reconstruction

Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces

Model ReleasesDGX agent

arXiv:2605.15753v1 Announce Type: cross Abstract: Functional 3D scene graphs offer a versatile and flexible representation for 3D scene understanding and robotic manipulation, defined by object nodes,

Highly Detailed and Generalizable Broadleaf Tree Crown Instance Segmentation from UAV Imagery

ResearchDGX agent

arXiv:2605.15673v1 Announce Type: cross Abstract: We present a highly detailed instance segmentation model for delineating individual tree crowns in natural broadleaf forests using aerial imagery acqu

How to Choose Your Teacher for Fine Grained Image Recognition

TutorialsDGX agent

arXiv:2605.15689v1 Announce Type: new Abstract: Fine-grained image recognition classifies subcategories such as bird species or car models. While state-of-the-art (SOTA) models are accurate, they are

HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion

SafetyDGX agent

arXiv:2605.15741v1 Announce Type: new Abstract: Pixel-space diffusion models bypass the reconstruction bottleneck of Variational Autoencoders (VAEs) but face a fundamental 'granularity dilemma': captu

IHF-Harmony: Multi-Modality Magnetic Resonance Images Harmonization using Invertible Hierarchy Flow Model

ResearchDGX agent

arXiv:2602.21536v2 Announce Type: replace Abstract: Retrospective MRI harmonization is limited by poor scalability across modalities and reliance on traveling subject datasets. To address these challe

Inevitable Encounters: Backdoor Attacks Involving Lossy Compression

ApplicationsDGX agent

arXiv:2603.13864v2 Announce Type: replace-cross Abstract: Real-world backdoor attacks often require poisoned datasets to be stored and transmitted before being used to compromise deep learning systems

Integrating chemical structures as treatments improves representations of microscopy images for morphological profiling

TutorialsDGX agent

arXiv:2504.09544v3 Announce Type: replace-cross Abstract: Recent advances in self-supervised deep learning have improved our ability to quantify cellular morphological changes in high-throughput micro

Invaria: Learning Scale and Density Invariance in Point Clouds via Next-Resolution Prediction

TutorialsDGX agent

arXiv:2605.15923v1 Announce Type: new Abstract: Modern image encoders achieve high generalization by decoupling semantic meaning from resolution, an ability yet to be fully realized in the 3D domain.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching

Local AiDGX agent

arXiv:2506.23552v2 Announce Type: replace Abstract: The intrinsic link between facial motion and speech is often overlooked in generative modeling, where talking head synthesis and text-to-speech (TTS

LAPS: Improving Incremental LiDAR Mapping using Active Pooling and Sampling for Neural Distance Fields

ApplicationsDGX agent

arXiv:2605.15496v1 Announce Type: cross Abstract: Neural distance fields offer a compact and continuous representation of 3D geometry, making them attractive for incremental LiDAR mapping. However, th

Layer Selection in Feature-Based Losses Affects Image Quality and Microstructural Consistency in Deep Learning Super-Resolution of Brain Diffusion MRI

ResearchDGX agent

arXiv:2605.15895v1 Announce Type: cross Abstract: Clinical application of high-resolution diffusion MRI is hindered by hardware limitations and prohibitive scan times, motivating computational super-r

LDGuid: A Framework for Robust Change Detection via Latent Difference Guidance

TutorialsDGX agent

arXiv:2605.15582v1 Announce Type: new Abstract: Modern deep learning models for change detection (CD) often struggle to explicitly represent task-relevant semantic differences. This paper proposes the

Learn2Splat: Extending the Horizon of Learned 3DGS Optimization

Model ReleasesDGX agent

arXiv:2605.15760v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) optimization is most commonly performed using standard optimizers (Adam, SGD). While stable across diverse scenes, standard

Learning Disentangled Representations for Generalized Multi-view Clustering

Model ReleasesDGX agent

arXiv:2605.15640v1 Announce Type: new Abstract: Multi-View Clustering (MVC) has gained significant attention for its ability to leverage complementary information across diverse views. However, existi

Learning Dynamic Structural Specialization for Underwater Salient Object Detection

Model ReleasesDGX agent

arXiv:2605.15535v1 Announce Type: new Abstract: Underwater salient object detection (USOD) has attracted increasing attention for underwater visual scene understanding and vision-guided robotic applic

Learning Normalized Energy Models for Linear Inverse Problems

ResearchDGX agent

arXiv:2605.15487v1 Announce Type: cross Abstract: Generative diffusion models can provide powerful prior probability models for inverse problems in imaging, but existing implementations suffer from tw

LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs

SafetyDGX agent

arXiv:2605.15621v1 Announce Type: new Abstract: Large vision-language models (LVLMs) achieve strong multimodal understanding, but their inference cost grows rapidly with the number of visual tokens, e

LUIVITON: Learned Universal Interoperable VIrtual Try-ON

SafetyDGX agent

arXiv:2509.05030v2 Announce Type: replace Abstract: To enable large-scale reuse of real-world 3D assets, where garments and characters rarely share skeletons, templates, or dense correspondences, we p

MAgSeg: Segmentation of Agricultural Landscapes in High-Resolution Satellite Imagery using Multimodal Large Language Models

SafetyDGX agent

arXiv:2605.16179v1 Announce Type: new Abstract: Agricultural landscape segmentation in the Global South is challenging as it is characterized by fragmented plots, high intra-class variance, and a scar

Mask-Morph Graph U-Net: A Generalisable Mesh-Based Surrogate for Crashworthiness Field Prediction under Large Geometric Variation

Model ReleasesDGX agent

arXiv:2605.15231v1 Announce Type: cross Abstract: Nonlinear finite element crash simulations are accurate but computationally expensive, limiting their use in iterative design optimisation. Machine-le

MaTe: Images Are All You Need for Material Transfer via Diffusion Transformer

SafetyDGX agent

arXiv:2605.15660v1 Announce Type: new Abstract: Recent diffusion-based methods for material transfer rely on image fine-tuning or complex architectures with assistive networks, but face challenges inc

MI-CXR: A Benchmark for Longitudinal Reasoning over Multi-Interval Chest X-rays

Model ReleasesDGX agent

arXiv:2605.15574v1 Announce Type: new Abstract: Longitudinal chest X-ray (CXR) interpretation requires reasoning over disease evolution across multiple patient visits, yet most existing medical VQA be

MIND: Decoupling Model-Induced Label Noise via Latent Manifold Disentanglement

Local AiDGX agent

arXiv:2605.16081v1 Announce Type: cross Abstract: The paradigm of learning from automatic annotations driven by pre-trained experts and Foundation Models dominates data-hungry applications. However, i

← Previous
1…132133134135136…211
Next →