AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
14 Apr 2026

Delta Rectified Flow Sampling for Text-to-Image Editing

Model ReleasesDGX agent

arXiv:2509.05342v3 Announce Type: replace Abstract: We propose Delta Rectified Flow Sampling (DRFS), a novel inversion-free, path-aware editing framework within rectified flow models for text-to-image

Development and evaluation of CADe systems in low-prevalence setting: The RARE25 challenge for early detection of Barrett's neoplasia

Model ReleasesDGX agent

arXiv:2604.11171v1 Announce Type: new Abstract: Computer-aided detection (CADe) of early neoplasia in Barrett's esophagus is a low-prevalence surveillance problem in which clinically relevant findings

Device-Conditioned Neural Architecture Search for Efficient Robotic Manipulation

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.10170v1 Announce Type: cross Abstract: The growing complexity of visuomotor policies poses significant challenges for deployment with heterogeneous robotic hardware constraints. However, mo

DiffClean: Diffusion-based Makeup Removal for Accurate Age Estimation

ApplicationsDGX agent

arXiv:2507.13292v3 Announce Type: replace Abstract: Accurate age verification can protect underage users from unauthorized access to online platforms and e-commerce sites that provide age-restricted s

Differentiable Vector Quantization for Rate-Distortion Optimization of Generative Image Compression

ResearchDGX agent

arXiv:2604.10546v1 Announce Type: new Abstract: The rapid growth of visual data under stringent storage and bandwidth constraints makes extremely low-bitrate image compression increasingly important.

DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain

Model ReleasesDGX agent

arXiv:2604.10425v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have revolutionized general visual understanding. However, their application in the food domain rem

DisCo-FLoc: Using Dual-Level Visual-Geometric Contrasts to Disambiguate Depth-Aware Visual Floorplan Localization

ResearchDGX agent

arXiv:2601.01822v2 Announce Type: replace-cross Abstract: Since floorplan data is readily available, long-term persistent, and robust to changes in visual appearance, visual Floorplan Localization (FL

Do Instance Priors Help Weakly Supervised Semantic Segmentation?

ResearchDGX agent

arXiv:2604.11170v1 Announce Type: new Abstract: Semantic segmentation requires dense pixel-level annotations, which are costly and time-consuming to acquire. To address this, we present SeSAM, a frame

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding

Model ReleasesDGX agent

arXiv:2604.11177v1 Announce Type: new Abstract: We benchmark how internal reasoning traces, which we call thought streams, affect video scene understanding in vision-language models. Using four config

Do vision models perceive illusory motion in static images like humans?

Local AiDGX agent

arXiv:2604.09853v1 Announce Type: new Abstract: Understanding human motion processing is essential for building reliable, human-centered computer vision systems. Although deep neural networks (DNNs) a

DocRevive: A Unified Pipeline for Document Text Restoration

Model ReleasesDGX agent

arXiv:2604.10077v1 Announce Type: new Abstract: In Document Understanding, the challenge of reconstructing damaged, occluded, or incomplete text remains a critical yet unexplored problem. Subsequent d

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection

ApplicationsDGX agent

arXiv:2604.09920v1 Announce Type: new Abstract: Vision foundation models (VFMs) offer the promise of zero-shot object detection without task-specific training data, yet their performance in complex ag

DoReMi: Bridging 3D Domains via Topology-Aware Domain-Representation Mixture of Experts

Model ReleasesDGX agent

arXiv:2511.11232v2 Announce Type: replace Abstract: Constructing a unified 3D scene understanding model has long been hindered by the significant topological discrepancies across different sensor moda

DoSReMC: Domain Shift Resilient Mammography Classification using Batch Normalization Adaptation

ApplicationsDGX agent

arXiv:2508.15452v3 Announce Type: replace-cross Abstract: Numerous deep learning-based solutions have been developed for the automatic recognition of breast cancer using mammography images. However, t

DRIFT: Deep Restoration, ISP Fusion, and Tone-mapping

SafetyDGX agent

arXiv:2604.03402v2 Announce Type: replace-cross Abstract: Smartphone cameras have gained immense popularity with the adoption of high-resolution and high-dynamic range imaging. As a result, high-perfo

Dual-Branch Remote Sensing Infrared Image Super-Resolution

Local AiDGX agent

arXiv:2604.10112v1 Announce Type: new Abstract: Remote sensing infrared image super-resolution aims to recover sharper thermal observations from low-resolution inputs while preserving target contours,

Dual-Exposure Imaging with Events

SafetyDGX agent

arXiv:2604.10273v1 Announce Type: new Abstract: By combining complementary benefits of short- and long-exposure images, Dual-Exposure Imaging (DEI) enhances image quality in low-light scenarios. Howev

Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy

ApplicationsDGX agent

arXiv:2512.18994v2 Announce Type: replace Abstract: Taxonomic classification of ecological families, genera, and species underpins biodiversity monitoring and conservation. Existing computer vision me

Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders

ResearchDGX agent

arXiv:2512.13876v2 Announce Type: replace Abstract: Detection Transformers (DETR) formulate object detection as a set prediction problem and enable end-to-end training without post-processing. However

EagleVision: A Multi-Task Benchmark for Cross-Domain Perception in High-Speed Autonomous Racing

Model ReleasesDGX agent

arXiv:2604.11400v1 Announce Type: cross Abstract: High-speed autonomous racing presents extreme perception challenges, including large relative velocities and substantial domain shifts from convention

EDFNet: Early Fusion of Edge and Depth for Thin-Obstacle Segmentation in UAV Navigation

AgentsDGX agent

arXiv:2604.09694v1 Announce Type: new Abstract: Autonomous Unmanned Aerial Vehicles (UAVs) must reliably detect thin obstacles such as wires, poles, and branches to navigate safely in real-world envir

EdgeDAM: Real-time Object Tracking for Mobile Devices

Model ReleasesDGX agent

arXiv:2603.05463v2 Announce Type: replace Abstract: Single-object tracking (SOT) on edge devices is a critical computer vision task, requiring accurate and continuous target localization across video

EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model

ResearchDGX agent

arXiv:2604.10268v1 Announce Type: new Abstract: We propose EditCrafter, a high-resolution image editing method that operates without tuning, leveraging pretrained text-to-image (T2I) diffusion models

Eevee: Towards Close-up High-resolution Video-based Virtual Try-on

Model ReleasesDGX agent

arXiv:2511.18957v2 Announce Type: replace Abstract: Video virtual try-on technology provides a cost-effective solution for creating marketing videos in fashion e-commerce. However, its practical adopt

Efficient Matrix Implementation for Rotary Position Embedding

ResearchDGX agent

arXiv:2604.09742v1 Announce Type: cross Abstract: Rotary Position Embedding (RoPE) has become a core component of modern Transformer architectures across language, vision, and 3D domains. However, exi

Efficient Transceiver Design for Aerial Image Transmission and Large-scale Scene Reconstruction

ApplicationsDGX agent

arXiv:2604.11098v1 Announce Type: new Abstract: Large-scale three-dimensional (3D) scene reconstruction in low-altitude intelligent networks (LAIN) demands highly efficient wireless image transmission

EgoFun3D: Modeling Interactive Objects from Egocentric Videos using Function Templates

Model ReleasesDGX agent

arXiv:2604.11038v1 Announce Type: new Abstract: We present EgoFun3D, a coordinated task formulation, dataset, and benchmark for modeling interactive 3D objects from egocentric videos. Interactive obje

Empowering Video Translation using Multimodal Large Language Models

SafetyDGX agent

arXiv:2604.11283v1 Announce Type: new Abstract: Recent developments in video translation have further enhanced cross-lingual access to video content, with multimodal large language models (MLLMs) play

Energy-oriented Diffusion Bridge for Image Restoration with Foundational Diffusion Models

TutorialsDGX agent

arXiv:2604.10983v1 Announce Type: new Abstract: Diffusion bridge models have shown great promise in image restoration by explicitly connecting clean and degraded image distributions. However, they oft

Enhancing Fine-Grained Spatial Grounding in 3D CT Report Generation via Discriminative Guidance

SafetyDGX agent

arXiv:2604.10437v1 Announce Type: new Abstract: Vision--language models (VLMs) for radiology report generation (RRG) can produce long-form chest CT reports from volumetric scans and show strong potent

EviRCOD: Evidence-Guided Probabilistic Decoding for Referring Camouflaged Object Detection

Model ReleasesDGX agent

arXiv:2604.10894v1 Announce Type: new Abstract: Referring Camouflaged Object Detection (Ref-COD) focuses on segmenting specific camouflaged targets in a query image using category-aligned references.

ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos

ResearchDGX agent

arXiv:2604.10466v1 Announce Type: new Abstract: Visual feedback is critical for motor skill acquisition in sports and rehabilitation, and psychological studies show that observing near-perfect version

Exploring Cross-Modal Flows for Few-Shot Learning

Model ReleasesDGX agent

arXiv:2510.14543v4 Announce Type: replace Abstract: Aligning features from different modalities, is one of the most fundamental challenges for cross-modal tasks. Although pre-trained vision-language m

Exploring the best way for UAV visual localization under Low-altitude Multi-view Observation Condition: a Benchmark

Model ReleasesDGX agent

arXiv:2503.10692v2 Announce Type: replace Abstract: Absolute Visual Localization (AVL) enables an Unmanned Aerial Vehicle (UAV) to determine its position in GNSS-denied environments by establishing ge

FastSHADE: Fast Self-augmented Hierarchical Asymmetric Denoising for Efficient inference on mobile devices

Model ReleasesDGX agent

arXiv:2604.10275v1 Announce Type: new Abstract: Real-time image denoising is essential for modern mobile photography but remains challenging due to the strict latency and power constraints of edge dev

FF3R: Feedforward Feature 3D Reconstruction from Unconstrained views

ResearchDGX agent

arXiv:2604.09862v1 Announce Type: new Abstract: Recent advances in vision foundation models have revolutionized geometry reconstruction and semantic understanding. Yet, most of the existing approaches

FGML-DG: Feynman-Inspired Cognitive Science Paradigm for Cross-Domain Medical Image Segmentation

SafetyDGX agent

arXiv:2604.10524v1 Announce Type: new Abstract: In medical image segmentation across multiple modalities (e.g., MRI, CT, etc.) and heterogeneous data sources (e.g., different hospitals and devices), D

FineEdit: Fine-Grained Image Edit with Bounding Box Guidance

Model ReleasesDGX agent

arXiv:2604.10954v1 Announce Type: new Abstract: Diffusion-based image editing models have achieved significant progress in real world applications. However, conventional models typically rely on natur

Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models

Model ReleasesDGX agent

arXiv:2604.11576v1 Announce Type: new Abstract: Despite their impressive zero-shot abilities, vision-language models such as CLIP have been shown to be susceptible to adversarial attacks. To enhance i

Flow Gym: A framework for the development, benchmarking, training, and deployment of flow-field quantification methods

ResearchDGX agent

arXiv:2512.20642v2 Announce Type: replace-cross Abstract: Particle image velocimetry (PIV) and related optical-flow methods are widely used to quantify fluid motion, but their development and evaluati

FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action Models

ResearchDGX agent

arXiv:2604.09651v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are emerging as a cornerstone for robotics, with flow-matching policies like pi_0 showing great promise in generatin

FM-SIREN & FM-FINER: Implicit Neural Representation Using Nyquist-based Orthogonality

ResearchDGX agent

arXiv:2509.23438v3 Announce Type: replace Abstract: Existing periodic activation-based implicit neural representation (INR) networks, such as SIREN and FINER, suffer from hidden feature redundancy, wh

FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis

Model ReleasesDGX agent

arXiv:2512.18073v2 Announce Type: replace Abstract: Multimodal LLMs (MLLMs) are capable of performing complex data analysis, visual question answering, generation, and reasoning tasks. However, their

FRAMER: Frequency-Aligned Self-Distillation with Adaptive Modulation Leveraging Diffusion Priors for Real-World Image Super-Resolution

SafetyDGX agent

arXiv:2512.01390v3 Announce Type: replace Abstract: Real-image super-resolution (Real-ISR) seeks to recover HR images from LR inputs with mixed, unknown degradations. While diffusion models surpass GA

FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation

ApplicationsDGX agent

arXiv:2604.10512v1 Announce Type: new Abstract: The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse

Fusion Complexity Inversion: Why Simpler Cross View Modules Outperform SSMs and Cross View Attention Transformers for Pasture Biomass Regression

Model ReleasesDGX agent

arXiv:2603.07819v3 Announce Type: replace Abstract: Accurate estimation of pasture biomass from agricultural imagery is critical for sustainable livestock management, yet existing methods are limited

Gait Recognition with Temporal Kolmogorov-Arnold Networks

Local AiDGX agent

arXiv:2604.09990v1 Announce Type: new Abstract: Gait recognition is a biometric modality that identifies individuals from their characteristic walking patterns. Unlike conventional biometric traits, g

GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2601.03416v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have become widely deployed, yet their safety alignment remains fragile under adversarial inputs. Previous

GaNI: Global and Near Field Illumination Aware Neural Inverse Rendering

ResearchDGX agent

arXiv:2403.15651v5 Announce Type: replace Abstract: In this paper, we present GaNI, a Global and Near-field Illumination-aware neural inverse rendering technique that can reconstruct geometry, albedo,

GazeVaLM: A Multi-Observer Eye-Tracking Benchmark for Evaluating Clinical Realism in AI-Generated X-Rays

Model ReleasesDGX agent

arXiv:2604.11653v1 Announce Type: new Abstract: We introduce GazeVaLM, a public eye-tracking dataset for studying clinical perception during chest radiograph authenticity assessment. The dataset compr

Generalizable Deepfake Detection Based on Forgery-aware Layer Masking and Multi-artifact Subspace Decomposition

Model ReleasesDGX agent

arXiv:2601.01041v3 Announce Type: replace Abstract: Deepfake detection remains highly challenging, particularly in cross-dataset scenarios and complex real-world settings. This challenge mainly arises

GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2509.04334v4 Announce Type: replace Abstract: Geographic reasoning is a fundamental cognitive capability that requires models to infer plausible locations by synthesizing visual evidence with sp

GeoFormer: A Lightweight Swin Transformer for Joint Building Height and Footprint Estimation from Sentinel Imagery

Model ReleasesDGX agent

arXiv:2602.09932v2 Announce Type: replace Abstract: Building height (BH) and footprint (BF) are fundamental urban morphological parameters required by climate modelling, disaster-risk assessment, and

GeomPrompt: Geometric Prompt Learning for RGB-D Semantic Segmentation Under Missing and Degraded Depth

ResearchDGX agent

arXiv:2604.11585v1 Announce Type: new Abstract: Multimodal perception systems for robotics and embodied AI often assume reliable RGB-D sensing, but in practice, depth is frequently missing, noisy, or

Geoparsing: Diagram Parsing for Plane and Solid Geometry with a Unified Formal Language

ApplicationsDGX agent

arXiv:2604.11600v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress but continue to struggle with geometric reasoning, primarily due to the perce

GIF: A Conditional Multimodal Generative Framework for IR Drop Imaging in Chip Layouts

TutorialsDGX agent

arXiv:2604.09999v1 Announce Type: new Abstract: IR drop analysis is essential in physical chip design to ensure the power integrity of on-chip power delivery networks. Traditional Electronic Design Au

Global monitoring of methane point sources using deep learning on hyperspectral radiance measurements from EMIT

Local AiDGX agent

arXiv:2604.10094v1 Announce Type: new Abstract: Anthropogenic methane (CH4) point sources drive near-term climate forcing, safety hazards, and system inefficiencies. Space-based imaging spectroscopy i

GrOCE:Graph-Guided Online Concept Erasure for Text-to-Image Diffusion Models

ResearchDGX agent

arXiv:2511.12968v2 Announce Type: replace Abstract: Concept erasure aims to remove harmful, inappropriate, or copyrighted content from text-to-image diffusion models while preserving non-target semant

Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis

ResearchDGX agent

arXiv:2604.06939v2 Announce Type: replace Abstract: Autoregressive video synthesis offers a promising pathway for infinite-horizon generation but is fundamentally hindered by three intertwined challen

GS4City: Hierarchical Semantic Gaussian Splatting via City-Model Priors

ResearchDGX agent

arXiv:2604.11401v1 Announce Type: new Abstract: Recent semantic 3D Gaussian Splatting (3DGS) methods primarily rely on 2D foundation models, often yielding ambiguous boundaries and limited support for

← Previous
1…195196197198199…207
Next →