AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Model Releases

Delta Rectified Flow Sampling for Text-to-Image Editing

DGX agent

arXiv:2509.05342v3 Announce Type: replace Abstract: We propose Delta Rectified Flow Sampling (DRFS), a novel inversion-free, path-aware editing framework within rectified flow models for text-to-image

model-releasesarxiv-cs-cv
14 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Development and evaluation of CADe systems in low-prevalence setting: The RARE25 challenge for early detection of Barrett's neoplasia

DGX agent

arXiv:2604.11171v1 Announce Type: new Abstract: Computer-aided detection (CADe) of early neoplasia in Barrett's esophagus is a low-prevalence surveillance problem in which clinically relevant findings

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

Device-Conditioned Neural Architecture Search for Efficient Robotic Manipulation

DGX agent

arXiv:2604.10170v1 Announce Type: cross Abstract: The growing complexity of visuomotor policies poses significant challenges for deployment with heterogeneous robotic hardware constraints. However, mo

safetyarxiv-cs-cv
14 Apr 2026
Applications

DiffClean: Diffusion-based Makeup Removal for Accurate Age Estimation

DGX agent

arXiv:2507.13292v3 Announce Type: replace Abstract: Accurate age verification can protect underage users from unauthorized access to online platforms and e-commerce sites that provide age-restricted s

applicationsarxiv-cs-cv
14 Apr 2026
Research

Differentiable Vector Quantization for Rate-Distortion Optimization of Generative Image Compression

DGX agent

arXiv:2604.10546v1 Announce Type: new Abstract: The rapid growth of visual data under stringent storage and bandwidth constraints makes extremely low-bitrate image compression increasingly important.

researcharxiv-cs-cv
14 Apr 2026
Model Releases

DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain

DGX agent

arXiv:2604.10425v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have revolutionized general visual understanding. However, their application in the food domain rem

model-releasesarxiv-cs-cv
14 Apr 2026
Research

DisCo-FLoc: Using Dual-Level Visual-Geometric Contrasts to Disambiguate Depth-Aware Visual Floorplan Localization

DGX agent

arXiv:2601.01822v2 Announce Type: replace-cross Abstract: Since floorplan data is readily available, long-term persistent, and robust to changes in visual appearance, visual Floorplan Localization (FL

researcharxiv-cs-cv
14 Apr 2026
Research

Do Instance Priors Help Weakly Supervised Semantic Segmentation?

DGX agent

arXiv:2604.11170v1 Announce Type: new Abstract: Semantic segmentation requires dense pixel-level annotations, which are costly and time-consuming to acquire. To address this, we present SeSAM, a frame

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding

DGX agent

arXiv:2604.11177v1 Announce Type: new Abstract: We benchmark how internal reasoning traces, which we call thought streams, affect video scene understanding in vision-language models. Using four config

model-releasesarxiv-cs-cv
14 Apr 2026
Local Ai

Do vision models perceive illusory motion in static images like humans?

DGX agent

arXiv:2604.09853v1 Announce Type: new Abstract: Understanding human motion processing is essential for building reliable, human-centered computer vision systems. Although deep neural networks (DNNs) a

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

DocRevive: A Unified Pipeline for Document Text Restoration

DGX agent

arXiv:2604.10077v1 Announce Type: new Abstract: In Document Understanding, the challenge of reconstructing damaged, occluded, or incomplete text remains a critical yet unexplored problem. Subsequent d

model-releasesarxiv-cs-cv
14 Apr 2026
Applications

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection

DGX agent

arXiv:2604.09920v1 Announce Type: new Abstract: Vision foundation models (VFMs) offer the promise of zero-shot object detection without task-specific training data, yet their performance in complex ag

applicationsarxiv-cs-cv
14 Apr 2026
Model Releases

DoReMi: Bridging 3D Domains via Topology-Aware Domain-Representation Mixture of Experts

DGX agent

arXiv:2511.11232v2 Announce Type: replace Abstract: Constructing a unified 3D scene understanding model has long been hindered by the significant topological discrepancies across different sensor moda

model-releasesarxiv-cs-cv
14 Apr 2026
Applications

DoSReMC: Domain Shift Resilient Mammography Classification using Batch Normalization Adaptation

DGX agent

arXiv:2508.15452v3 Announce Type: replace-cross Abstract: Numerous deep learning-based solutions have been developed for the automatic recognition of breast cancer using mammography images. However, t

applicationsarxiv-cs-cv
14 Apr 2026
Safety

DRIFT: Deep Restoration, ISP Fusion, and Tone-mapping

DGX agent

arXiv:2604.03402v2 Announce Type: replace-cross Abstract: Smartphone cameras have gained immense popularity with the adoption of high-resolution and high-dynamic range imaging. As a result, high-perfo

safetyarxiv-cs-cv
14 Apr 2026
Local Ai

Dual-Branch Remote Sensing Infrared Image Super-Resolution

DGX agent

arXiv:2604.10112v1 Announce Type: new Abstract: Remote sensing infrared image super-resolution aims to recover sharper thermal observations from low-resolution inputs while preserving target contours,

local-aiarxiv-cs-cv
14 Apr 2026
Safety

Dual-Exposure Imaging with Events

DGX agent

arXiv:2604.10273v1 Announce Type: new Abstract: By combining complementary benefits of short- and long-exposure images, Dual-Exposure Imaging (DEI) enhances image quality in low-light scenarios. Howev

safetyarxiv-cs-cv
14 Apr 2026
Applications

Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy

DGX agent

arXiv:2512.18994v2 Announce Type: replace Abstract: Taxonomic classification of ecological families, genera, and species underpins biodiversity monitoring and conservation. Existing computer vision me

applicationsarxiv-cs-cv
14 Apr 2026
Research

Dual-R-DETR: Resolving Query Competition with Pairwise Routing in Transformer Decoders

DGX agent

arXiv:2512.13876v2 Announce Type: replace Abstract: Detection Transformers (DETR) formulate object detection as a set prediction problem and enable end-to-end training without post-processing. However

researcharxiv-cs-cv
14 Apr 2026
Model Releases

EagleVision: A Multi-Task Benchmark for Cross-Domain Perception in High-Speed Autonomous Racing

DGX agent

arXiv:2604.11400v1 Announce Type: cross Abstract: High-speed autonomous racing presents extreme perception challenges, including large relative velocities and substantial domain shifts from convention

model-releasesarxiv-cs-cv
14 Apr 2026
Agents

EDFNet: Early Fusion of Edge and Depth for Thin-Obstacle Segmentation in UAV Navigation

DGX agent

arXiv:2604.09694v1 Announce Type: new Abstract: Autonomous Unmanned Aerial Vehicles (UAVs) must reliably detect thin obstacles such as wires, poles, and branches to navigate safely in real-world envir

agentsarxiv-cs-cv
14 Apr 2026
Model Releases

EdgeDAM: Real-time Object Tracking for Mobile Devices

DGX agent

arXiv:2603.05463v2 Announce Type: replace Abstract: Single-object tracking (SOT) on edge devices is a critical computer vision task, requiring accurate and continuous target localization across video

model-releasesarxiv-cs-cv
14 Apr 2026
Research

EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model

DGX agent

arXiv:2604.10268v1 Announce Type: new Abstract: We propose EditCrafter, a high-resolution image editing method that operates without tuning, leveraging pretrained text-to-image (T2I) diffusion models

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Eevee: Towards Close-up High-resolution Video-based Virtual Try-on

DGX agent

arXiv:2511.18957v2 Announce Type: replace Abstract: Video virtual try-on technology provides a cost-effective solution for creating marketing videos in fashion e-commerce. However, its practical adopt

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Efficient Matrix Implementation for Rotary Position Embedding

DGX agent

arXiv:2604.09742v1 Announce Type: cross Abstract: Rotary Position Embedding (RoPE) has become a core component of modern Transformer architectures across language, vision, and 3D domains. However, exi

researcharxiv-cs-cv
14 Apr 2026
Applications

Efficient Transceiver Design for Aerial Image Transmission and Large-scale Scene Reconstruction

DGX agent

arXiv:2604.11098v1 Announce Type: new Abstract: Large-scale three-dimensional (3D) scene reconstruction in low-altitude intelligent networks (LAIN) demands highly efficient wireless image transmission

applicationsarxiv-cs-cv
14 Apr 2026
Model Releases

EgoFun3D: Modeling Interactive Objects from Egocentric Videos using Function Templates

DGX agent

arXiv:2604.11038v1 Announce Type: new Abstract: We present EgoFun3D, a coordinated task formulation, dataset, and benchmark for modeling interactive 3D objects from egocentric videos. Interactive obje

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

Empowering Video Translation using Multimodal Large Language Models

DGX agent

arXiv:2604.11283v1 Announce Type: new Abstract: Recent developments in video translation have further enhanced cross-lingual access to video content, with multimodal large language models (MLLMs) play

safetyarxiv-cs-cv
14 Apr 2026
Tutorials

Energy-oriented Diffusion Bridge for Image Restoration with Foundational Diffusion Models

DGX agent

arXiv:2604.10983v1 Announce Type: new Abstract: Diffusion bridge models have shown great promise in image restoration by explicitly connecting clean and degraded image distributions. However, they oft

tutorialsarxiv-cs-cv
14 Apr 2026
Safety

Enhancing Fine-Grained Spatial Grounding in 3D CT Report Generation via Discriminative Guidance

DGX agent

arXiv:2604.10437v1 Announce Type: new Abstract: Vision--language models (VLMs) for radiology report generation (RRG) can produce long-form chest CT reports from volumetric scans and show strong potent

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

EviRCOD: Evidence-Guided Probabilistic Decoding for Referring Camouflaged Object Detection

DGX agent

arXiv:2604.10894v1 Announce Type: new Abstract: Referring Camouflaged Object Detection (Ref-COD) focuses on segmenting specific camouflaged targets in a query image using category-aligned references.

model-releasesarxiv-cs-cv
14 Apr 2026
Research

ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos

DGX agent

arXiv:2604.10466v1 Announce Type: new Abstract: Visual feedback is critical for motor skill acquisition in sports and rehabilitation, and psychological studies show that observing near-perfect version

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Exploring Cross-Modal Flows for Few-Shot Learning

DGX agent

arXiv:2510.14543v4 Announce Type: replace Abstract: Aligning features from different modalities, is one of the most fundamental challenges for cross-modal tasks. Although pre-trained vision-language m

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

Exploring the best way for UAV visual localization under Low-altitude Multi-view Observation Condition: a Benchmark

DGX agent

arXiv:2503.10692v2 Announce Type: replace Abstract: Absolute Visual Localization (AVL) enables an Unmanned Aerial Vehicle (UAV) to determine its position in GNSS-denied environments by establishing ge

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

FastSHADE: Fast Self-augmented Hierarchical Asymmetric Denoising for Efficient inference on mobile devices

DGX agent

arXiv:2604.10275v1 Announce Type: new Abstract: Real-time image denoising is essential for modern mobile photography but remains challenging due to the strict latency and power constraints of edge dev

model-releasesarxiv-cs-cv
14 Apr 2026
Research

FF3R: Feedforward Feature 3D Reconstruction from Unconstrained views

DGX agent

arXiv:2604.09862v1 Announce Type: new Abstract: Recent advances in vision foundation models have revolutionized geometry reconstruction and semantic understanding. Yet, most of the existing approaches

researcharxiv-cs-cv
14 Apr 2026
Safety

FGML-DG: Feynman-Inspired Cognitive Science Paradigm for Cross-Domain Medical Image Segmentation

DGX agent

arXiv:2604.10524v1 Announce Type: new Abstract: In medical image segmentation across multiple modalities (e.g., MRI, CT, etc.) and heterogeneous data sources (e.g., different hospitals and devices), D

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

FineEdit: Fine-Grained Image Edit with Bounding Box Guidance

DGX agent

arXiv:2604.10954v1 Announce Type: new Abstract: Diffusion-based image editing models have achieved significant progress in real world applications. However, conventional models typically rely on natur

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models

DGX agent

arXiv:2604.11576v1 Announce Type: new Abstract: Despite their impressive zero-shot abilities, vision-language models such as CLIP have been shown to be susceptible to adversarial attacks. To enhance i

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Flow Gym: A framework for the development, benchmarking, training, and deployment of flow-field quantification methods

DGX agent

arXiv:2512.20642v2 Announce Type: replace-cross Abstract: Particle image velocimetry (PIV) and related optical-flow methods are widely used to quantify fluid motion, but their development and evaluati

researcharxiv-cs-cv
14 Apr 2026
Research

FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action Models

DGX agent

arXiv:2604.09651v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are emerging as a cornerstone for robotics, with flow-matching policies like pi_0 showing great promise in generatin

researcharxiv-cs-cv
14 Apr 2026
Research

FM-SIREN & FM-FINER: Implicit Neural Representation Using Nyquist-based Orthogonality

DGX agent

arXiv:2509.23438v3 Announce Type: replace Abstract: Existing periodic activation-based implicit neural representation (INR) networks, such as SIREN and FINER, suffer from hidden feature redundancy, wh

researcharxiv-cs-cv
14 Apr 2026
Model Releases

FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis

DGX agent

arXiv:2512.18073v2 Announce Type: replace Abstract: Multimodal LLMs (MLLMs) are capable of performing complex data analysis, visual question answering, generation, and reasoning tasks. However, their

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

FRAMER: Frequency-Aligned Self-Distillation with Adaptive Modulation Leveraging Diffusion Priors for Real-World Image Super-Resolution

DGX agent

arXiv:2512.01390v3 Announce Type: replace Abstract: Real-image super-resolution (Real-ISR) seeks to recover HR images from LR inputs with mixed, unknown degradations. While diffusion models surpass GA

safetyarxiv-cs-cv
14 Apr 2026
Applications

FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation

DGX agent

arXiv:2604.10512v1 Announce Type: new Abstract: The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse

applicationsarxiv-cs-cv
14 Apr 2026
Model Releases

Fusion Complexity Inversion: Why Simpler Cross View Modules Outperform SSMs and Cross View Attention Transformers for Pasture Biomass Regression

DGX agent

arXiv:2603.07819v3 Announce Type: replace Abstract: Accurate estimation of pasture biomass from agricultural imagery is critical for sustainable livestock management, yet existing methods are limited

model-releasesarxiv-cs-cv
14 Apr 2026
Local Ai

Gait Recognition with Temporal Kolmogorov-Arnold Networks

DGX agent

arXiv:2604.09990v1 Announce Type: new Abstract: Gait recognition is a biometric modality that identifies individuals from their characteristic walking patterns. Unlike conventional biometric traits, g

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models

DGX agent

arXiv:2601.03416v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have become widely deployed, yet their safety alignment remains fragile under adversarial inputs. Previous

model-releasesarxiv-cs-cv
14 Apr 2026
← Previous
1…244245246247248…259
Next →