AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
23 Apr 2026

Online CS-based SAR Edge-Mapping

ResearchDGX agent

arXiv:2604.19989v1 Announce Type: new Abstract: With modern defense applications increasingly relying on inexpensive, small Unmanned Aerial Vehicles (UAVs), a major challenge lies in designing intelli

OnSiteVRU: A High-Resolution Trajectory Dataset for High-Density Vulnerable Road Users

SafetyDGX agent

arXiv:2503.23365v2 Announce Type: replace Abstract: With the acceleration of urbanization and the growth of transportation demands, the safety of vulnerable road users (VRUs, such as pedestrians and c

Opportunistic Bone-Loss Screening from Routine Knee Radiographs Using a Multi-Task Deep Learning Framework with Sensitivity-Constrained Threshold Optimization

ResearchDGX agent

arXiv:2604.20268v1 Announce Type: new Abstract: Background: Osteoporosis and osteopenia are often undiagnosed until fragility fractures occur. Dual-energy X-ray absorptiometry (DXA) is the reference s


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Optimizing Data Augmentation for Real-Time Small UAV Detection: A Lightweight Context-Aware Approach

ResearchDGX agent

arXiv:2604.19999v1 Announce Type: new Abstract: Visual detection of Unmanned Aerial Vehicles (UAVs) is a critical task in surveillance systems due to their small physical size and environmental challe

Pairing Regularization for Mitigating Many-to-One Collapse in GANs

ResearchDGX agent

arXiv:2604.20130v1 Announce Type: cross Abstract: Mode collapse remains a fundamental challenge in training generative adversarial networks (GANs). While existing works have primarily focused on inter

ParetoSlider: Diffusion Models Post-Training for Continuous Reward Control

ResearchDGX agent

arXiv:2604.20816v1 Announce Type: cross Abstract: Reinforcement Learning (RL) post-training has become the standard for aligning generative models with human preferences, yet most methods rely on a si

PASTA: A Patch-Agnostic Twofold-Stealthy Backdoor Attack on Vision Transformers

ResearchDGX agent

arXiv:2604.20047v1 Announce Type: new Abstract: Vision Transformers (ViTs) have achieved remarkable success across vision tasks, yet recent studies show they remain vulnerable to backdoor attacks. Exi

PFGNet: A Fully Convolutional Frequency-Guided Peripheral Gating Network for Efficient Spatiotemporal Predictive Learning

ResearchDGX agent

arXiv:2602.20537v3 Announce Type: replace Abstract: Spatiotemporal predictive learning (STPL) aims to forecast future frames from past observations and is essential across a wide range of applications

Physics-informed Active Polarimetric 3D Imaging for Specular Surfaces

ApplicationsDGX agent

arXiv:2602.19470v2 Announce Type: replace Abstract: 3D imaging of specular surfaces remains challenging in real-world scenarios, such as in-line inspection or hand-held scanning, requiring fast and ac

Physics-Informed Conditional Diffusion for Motion-Robust Retinal Temporal Laser Speckle Contrast Imaging

ResearchDGX agent

arXiv:2604.20594v1 Announce Type: new Abstract: Retinal laser speckle contrast imaging (LSCI) is a noninvasive optical modality for monitoring retinal blood flow dynamics. However, conventional tempor

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards

SafetyDGX agent

arXiv:2604.20486v1 Announce Type: new Abstract: Training multimodal agents via reinforcement learning for knowledge-intensive visual reasoning is fundamentally hindered by the extreme sparsity of outc

R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs

ResearchDGX agent

arXiv:2604.20696v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have demonstrated impressive performance in various multimodal understanding and reasoning tasks. However, they sti

Random Walk on Point Clouds for Feature Detection

ResearchDGX agent

arXiv:2604.20474v1 Announce Type: new Abstract: The points on the point clouds that can entirely outline the shape of the model are of critical importance, as they serve as the foundation for numerous

RareSpot+: A Benchmark, Model, and Active Learning Framework for Small and Rare Wildlife in Aerial Imagery

Model ReleasesDGX agent

arXiv:2604.20000v1 Announce Type: new Abstract: Automated wildlife monitoring from aerial imagery is vital for conservation but remains limited by two persistent challenges: the difficulty of detectin

RefAerial: A Benchmark and Approach for Referring Detection in Aerial Images

Model ReleasesDGX agent

arXiv:2604.20543v1 Announce Type: new Abstract: Referring detection refers to locate the target referred by natural languages, which has recently attracted growing research interests. However, existin

Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback

TutorialsDGX agent

arXiv:2604.20730v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown promising capabilities in generating Scalable Vector Graphics (SVG) via direct code synthesis. Howev

Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing

Model ReleasesDGX agent

arXiv:2604.20258v1 Announce Type: new Abstract: Instruction-based image editing (IIE) aims to modify images according to textual instructions while preserving irrelevant content. Despite recent advanc

retinalysis-vascx: An explainable software toolbox for the extraction of retinal vascular biomarkers

Model ReleasesDGX agent

arXiv:2602.08580v2 Announce Type: replace-cross Abstract: Automatic extraction of retinal vascular biomarkers from color fundus images (CFI) is crucial for large-scale studies of the retinal vasculatu

Retinex Meets Language: A Physics-Semantics-Guided Underwater Image Enhancement Network

ResearchDGX agent

arXiv:2603.07076v2 Announce Type: replace Abstract: Underwater images often suffer from severe degradation caused by light absorption and scattering, leading to color distortion, low contrast and redu

REVNET: Rotation-Equivariant Point Cloud Completion via Vector Neuron Anchor Transformer

Model ReleasesDGX agent

arXiv:2601.08558v2 Announce Type: replace Abstract: Incomplete point clouds captured by 3D sensors often result in the loss of both geometric and semantic information. Most existing point cloud comple

Robust Principal Component Completion

ResearchDGX agent

arXiv:2603.25132v2 Announce Type: replace Abstract: Robust principal component analysis (RPCA) seeks a low-rank component and a sparse component from their summation. Yet, in many applications of inte

Rodrigues Network for Learning Robot Actions

SafetyDGX agent

arXiv:2506.02618v2 Announce Type: replace-cross Abstract: Understanding and predicting articulated actions is important in robot learning. However, common architectures such as MLPs and Transformers l

Sampling-Aware Quantization for Diffusion Models

SafetyDGX agent

arXiv:2505.02242v2 Announce Type: replace Abstract: Diffusion models have recently emerged as the dominant approach in visual generation tasks. However, the lengthy denoising chains and the computatio

SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation

Model ReleasesDGX agent

arXiv:2604.19907v1 Announce Type: new Abstract: Recent agentic frameworks for 3D scene synthesis have advanced realism and diversity by integrating heterogeneous generation and editing tools. These to

Secure Rate-Distortion-Perception: A Randomized Distributed Function Computation Approach for Realism

ResearchDGX agent

arXiv:2604.20245v1 Announce Type: cross Abstract: Fundamental rate-distortion-perception (RDP) trade-offs arise in applications requiring maintained perceptual quality of reconstructed data, such as n

SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images

Model ReleasesDGX agent

arXiv:2512.08730v2 Announce Type: replace Abstract: Most existing methods for training-free open-vocabulary semantic segmentation are based on CLIP. While these approaches have made progress, they oft

Self-supervised pretraining for an iterative image size agnostic vision transformer

ResearchDGX agent

arXiv:2604.20392v1 Announce Type: new Abstract: Vision Transformers (ViTs) dominate self-supervised learning (SSL). While they have proven highly effective for large-scale pretraining, they are comput

Semantic-Fast-SAM: Efficient Semantic Segmenter

ResearchDGX agent

arXiv:2604.20169v1 Announce Type: new Abstract: We propose Semantic-Fast-SAM (SFS), a semantic segmentation framework that combines the Fast Segment Anything model with a semantic labeling pipeline to

Semantic-guided Gaussian Splatting for High-Fidelity Underwater Scene Reconstruction

ApplicationsDGX agent

arXiv:2509.00800v3 Announce Type: replace Abstract: Accurate 3D reconstruction in degraded imaging conditions remains a key challenge in photogrammetry and neural rendering. In underwater environments

Semi-Supervised Flow Matching for Mosaiced and Panchromatic Fusion Imaging

Model ReleasesDGX agent

arXiv:2604.20128v1 Announce Type: new Abstract: Fusing a low resolution (LR) mosaiced hyperspectral image (HSI) with a high resolution (HR) panchromatic (PAN) image offers a promising avenue for video

SGAP-Gaze: Scene Grid Attention Based Point-of-Gaze Estimation Network for Driver Gaze

Model ReleasesDGX agent

arXiv:2604.19888v1 Announce Type: new Abstract: Driver gaze estimation is essential for understanding the driver's situational awareness of surrounding traffic. Existing gaze estimation models use dri

SpaCeFormer: Fast Proposal-Free Open-Vocabulary 3D Instance Segmentation

ResearchDGX agent

arXiv:2604.20395v1 Announce Type: new Abstract: Open-vocabulary 3D instance segmentation is a core capability for robotics and AR/VR, but prior methods trade one bottleneck for another: multi-stage 2D

SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models

ResearchDGX agent

arXiv:2604.20705v1 Announce Type: new Abstract: Reinforcement learning (RL) with verifiable rewards (RLVR) has demonstrated the great potential of enhancing the reasoning abilities in multimodal large

Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation

TutorialsDGX agent

arXiv:2604.20336v1 Announce Type: new Abstract: Co-manipulation requires multiple humans to synchronize their motions with a shared object while ensuring reasonable interactions, maintaining natural p

Structure-Augmented Standard Plane Detection with Temporal Aggregation in Blind-Sweep Fetal Ultrasound

ResearchDGX agent

arXiv:2604.20591v1 Announce Type: new Abstract: In low-resource settings, blind-sweep ultrasound provides a practical and accessible method for identifying fetal growth restriction. However, unlike fr

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark

Model ReleasesDGX agent

arXiv:2604.20319v1 Announce Type: new Abstract: Fine-grained spatiotemporal reasoning on surgical videos is critical, yet the capabilities of Multi-modal Large Language Models (MLLMs) in this domain r

Survival of the Cheapest: Cost-Aware Hardware Adaptation for Adversarial Robustness

Model ReleasesDGX agent

arXiv:2409.07609v2 Announce Type: replace-cross Abstract: Deploying adversarially robust machine learning systems requires continuous trade-offs between robustness, cost, and latency. We present an au

TactileEval: A Step Towards Automated Fine-Grained Evaluation and Editing of Tactile Graphics

ResearchDGX agent

arXiv:2604.19829v1 Announce Type: new Abstract: Tactile graphics require careful expert validation before reaching blind and visually impaired (BVI) learners, yet existing datasets provide only coarse

The Role and Relationship of Initialization and Densification in 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2603.20714v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become the method of choice for photo-realistic 3D reconstruction of scenes, due to being able to efficiently and a

Topology-Aware Skeleton Detection via Lighthouse-Guided Structured Inference

TutorialsDGX agent

arXiv:2604.20123v1 Announce Type: new Abstract: In natural images, object skeletons are used to represent geometric shapes. However, even slight variations in pose or movement can cause noticeable cha

Towards reconstructing experimental sparse-view X-ray CT data with diffusion models

ApplicationsDGX agent

arXiv:2602.12755v2 Announce Type: replace Abstract: Diffusion-based image generators are promising priors for ill-posed inverse problems like sparse-view X-ray Computed Tomography (CT). As most studie

UniCon3R: Contact-aware 3D Human-Scene Reconstruction from Monocular Video

ResearchDGX agent

arXiv:2604.19923v1 Announce Type: new Abstract: We introduce UniCon3R (Unified Contact-aware 3D Reconstruction), a unified feed-forward framework for online human-scene 4D reconstruction from monocula

UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval

Model ReleasesDGX agent

arXiv:2604.20318v1 Announce Type: new Abstract: Composed image retrieval, multi-turn composed image retrieval, and composed video retrieval all share a common paradigm: composing the reference visual

Video-ToC: Video Tree-of-Cue Reasoning

Model ReleasesDGX agent

arXiv:2604.20473v1 Announce Type: new Abstract: Existing Video Large Language Models (Video LLMs) struggle with complex video understanding, exhibiting limited reasoning capabilities and potential hal

Visual Reasoning through Tool-supervised Reinforcement Learning

AgentsDGX agent

arXiv:2604.19945v1 Announce Type: new Abstract: In this paper, we investigate the problem of how to effectively master tool-use to solve complex visual reasoning tasks for Multimodal Large Language Mo

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

ApplicationsDGX agent

arXiv:2604.19858v1 Announce Type: new Abstract: We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into p

Weighted Knowledge Distillation for Semi-Supervised Segmentation of Maxillary Sinus in Panoramic X-ray Images

ResearchDGX agent

arXiv:2604.20213v1 Announce Type: new Abstract: Accurate segmentation of maxillary sinus in panoramic X-ray images is essential for dental diagnosis and surgical planning; however, this task remains r

Where are they looking in the operating room?

ResearchDGX agent

arXiv:2604.20574v1 Announce Type: new Abstract: Purpose: Gaze-following, the task of inferring where individuals are looking, has been widely studied in computer vision, advancing research in visual a

WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring

Model ReleasesDGX agent

arXiv:2604.20190v1 Announce Type: new Abstract: Wildfire monitoring requires timely, actionable situational awareness from airborne platforms, yet existing aerial visual question answering (VQA) bench

X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference

AgentsDGX agent

arXiv:2604.20289v1 Announce Type: new Abstract: Real-time world simulation is becoming a key infrastructure for scalable evaluation and online reinforcement learning of autonomous driving systems. Rec

X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis

Model ReleasesDGX agent

arXiv:2604.20350v1 Announce Type: new Abstract: Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely u

22 Apr 2026

A Controlled Benchmark of Visual State-Space Backbones with Domain-Shift and Boundary Analysis for Remote-Sensing Segmentation

Model ReleasesDGX agent

arXiv:2604.18721v1 Announce Type: cross Abstract: Visual state-space models (SSMs) are increasingly promoted as efficient alternatives to Vision Transformers, yet their practical advantages remain unc

A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation

AgentsDGX agent

arXiv:2604.18988v1 Announce Type: new Abstract: Multimodal empathetic response generation (MERG) aims to generate emotionally engaging and empathetic responses based on users' multimodal contexts. Exi

A Network-Aware Evaluation of Distributed Energy Resource Control in Smart Distribution Systems

ResearchDGX agent

arXiv:2604.19715v1 Announce Type: new Abstract: Distribution networks with high penetration of Distributed Energy Resources (DERs) increasingly rely on communication networks to coordinate grid-intera

AdaGScale: Viewpoint-Adaptive Gaussian Scaling in 3D Gaussian Splatting to Reduce Gaussian-Tile Pairs

HardwareDGX agent

arXiv:2604.18980v1 Announce Type: new Abstract: Reducing the number of Gaussian-tile pairs is one of the most promising approaches to improve 3D Gaussian Splatting (3D-GS) rendering speed on GPUs. How

Adapting Self-Supervised Representations as a Latent Space for Efficient Generation

ResearchDGX agent

arXiv:2510.14630v2 Announce Type: replace Abstract: We introduce Representation Tokenizer (RepTok), a generative modeling framework that represents an image using a single continuous latent token obta

Adaptive Slicing-Assisted Hyper Inference for Enhanced Small Object Detection in High-Resolution Imagery

ResearchDGX agent

arXiv:2604.19233v1 Announce Type: new Abstract: Deep learning-based object detectors have achieved remarkable success across numerous computer vision applications, yet they continue to struggle with s

AI-Enabled Image-Based Hybrid Vision/Force Control of Tendon-Driven Aerial Continuum Manipulators

AgentsDGX agent

arXiv:2604.18961v1 Announce Type: cross Abstract: This paper presents an AI-enabled cascaded hybrid vision/force control framework for tendon-driven aerial continuum manipulators based on constant-str

Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2604.19386v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) has attracted significant attention due to its flexible multimodal query method, yet its development is severely constrai

Align then Refine: Text-Guided 3D Prostate Lesion Segmentation

Local AiDGX agent

arXiv:2604.18713v1 Announce Type: new Abstract: Automated 3D segmentation of prostate lesions from biparametric MRI (bp-MRI) is essential for reliable algorithmic analysis, but achieving high precisio

← Previous
1…176177178179180…209
Next →