AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Model Releases

Survival of the Cheapest: Cost-Aware Hardware Adaptation for Adversarial Robustness

DGX agent

arXiv:2409.07609v2 Announce Type: replace-cross Abstract: Deploying adversarially robust machine learning systems requires continuous trade-offs between robustness, cost, and latency. We present an au

model-releasesarxiv-cs-cv
23 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

TactileEval: A Step Towards Automated Fine-Grained Evaluation and Editing of Tactile Graphics

DGX agent

arXiv:2604.19829v1 Announce Type: new Abstract: Tactile graphics require careful expert validation before reaching blind and visually impaired (BVI) learners, yet existing datasets provide only coarse

researcharxiv-cs-cv
23 Apr 2026
Model Releases

The Role and Relationship of Initialization and Densification in 3D Gaussian Splatting

DGX agent

arXiv:2603.20714v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become the method of choice for photo-realistic 3D reconstruction of scenes, due to being able to efficiently and a

model-releasesarxiv-cs-cv
23 Apr 2026
Tutorials

Topology-Aware Skeleton Detection via Lighthouse-Guided Structured Inference

DGX agent

arXiv:2604.20123v1 Announce Type: new Abstract: In natural images, object skeletons are used to represent geometric shapes. However, even slight variations in pose or movement can cause noticeable cha

tutorialsarxiv-cs-cv
23 Apr 2026
Applications

Towards reconstructing experimental sparse-view X-ray CT data with diffusion models

DGX agent

arXiv:2602.12755v2 Announce Type: replace Abstract: Diffusion-based image generators are promising priors for ill-posed inverse problems like sparse-view X-ray Computed Tomography (CT). As most studie

applicationsarxiv-cs-cv
23 Apr 2026
Research

UniCon3R: Contact-aware 3D Human-Scene Reconstruction from Monocular Video

DGX agent

arXiv:2604.19923v1 Announce Type: new Abstract: We introduce UniCon3R (Unified Contact-aware 3D Reconstruction), a unified feed-forward framework for online human-scene 4D reconstruction from monocula

researcharxiv-cs-cv
23 Apr 2026
Model Releases

UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval

DGX agent

arXiv:2604.20318v1 Announce Type: new Abstract: Composed image retrieval, multi-turn composed image retrieval, and composed video retrieval all share a common paradigm: composing the reference visual

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

Video-ToC: Video Tree-of-Cue Reasoning

DGX agent

arXiv:2604.20473v1 Announce Type: new Abstract: Existing Video Large Language Models (Video LLMs) struggle with complex video understanding, exhibiting limited reasoning capabilities and potential hal

model-releasesarxiv-cs-cv
23 Apr 2026
Agents

Visual Reasoning through Tool-supervised Reinforcement Learning

DGX agent

arXiv:2604.19945v1 Announce Type: new Abstract: In this paper, we investigate the problem of how to effectively master tool-use to solve complex visual reasoning tasks for Multimodal Large Language Mo

agentsarxiv-cs-cv
23 Apr 2026
Applications

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

DGX agent

arXiv:2604.19858v1 Announce Type: new Abstract: We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into p

applicationsarxiv-cs-cv
23 Apr 2026
Research

Weighted Knowledge Distillation for Semi-Supervised Segmentation of Maxillary Sinus in Panoramic X-ray Images

DGX agent

arXiv:2604.20213v1 Announce Type: new Abstract: Accurate segmentation of maxillary sinus in panoramic X-ray images is essential for dental diagnosis and surgical planning; however, this task remains r

researcharxiv-cs-cv
23 Apr 2026
Research

Where are they looking in the operating room?

DGX agent

arXiv:2604.20574v1 Announce Type: new Abstract: Purpose: Gaze-following, the task of inferring where individuals are looking, has been widely studied in computer vision, advancing research in visual a

researcharxiv-cs-cv
23 Apr 2026
Model Releases

WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring

DGX agent

arXiv:2604.20190v1 Announce Type: new Abstract: Wildfire monitoring requires timely, actionable situational awareness from airborne platforms, yet existing aerial visual question answering (VQA) bench

model-releasesarxiv-cs-cv
23 Apr 2026
Agents

X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference

DGX agent

arXiv:2604.20289v1 Announce Type: new Abstract: Real-time world simulation is becoming a key infrastructure for scalable evaluation and online reinforcement learning of autonomous driving systems. Rec

agentsarxiv-cs-cv
23 Apr 2026
Model Releases

X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis

DGX agent

arXiv:2604.20350v1 Announce Type: new Abstract: Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely u

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

A Controlled Benchmark of Visual State-Space Backbones with Domain-Shift and Boundary Analysis for Remote-Sensing Segmentation

DGX agent

arXiv:2604.18721v1 Announce Type: cross Abstract: Visual state-space models (SSMs) are increasingly promoted as efficient alternatives to Vision Transformers, yet their practical advantages remain unc

model-releasesarxiv-cs-cv
22 Apr 2026
Agents

A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation

DGX agent

arXiv:2604.18988v1 Announce Type: new Abstract: Multimodal empathetic response generation (MERG) aims to generate emotionally engaging and empathetic responses based on users' multimodal contexts. Exi

agentsarxiv-cs-cv
22 Apr 2026
Research

A Network-Aware Evaluation of Distributed Energy Resource Control in Smart Distribution Systems

DGX agent

arXiv:2604.19715v1 Announce Type: new Abstract: Distribution networks with high penetration of Distributed Energy Resources (DERs) increasingly rely on communication networks to coordinate grid-intera

researcharxiv-cs-cv
22 Apr 2026
Hardware

AdaGScale: Viewpoint-Adaptive Gaussian Scaling in 3D Gaussian Splatting to Reduce Gaussian-Tile Pairs

DGX agent

arXiv:2604.18980v1 Announce Type: new Abstract: Reducing the number of Gaussian-tile pairs is one of the most promising approaches to improve 3D Gaussian Splatting (3D-GS) rendering speed on GPUs. How

hardwarearxiv-cs-cv
22 Apr 2026
Research

Adapting Self-Supervised Representations as a Latent Space for Efficient Generation

DGX agent

arXiv:2510.14630v2 Announce Type: replace Abstract: We introduce Representation Tokenizer (RepTok), a generative modeling framework that represents an image using a single continuous latent token obta

researcharxiv-cs-cv
22 Apr 2026
Research

Adaptive Slicing-Assisted Hyper Inference for Enhanced Small Object Detection in High-Resolution Imagery

DGX agent

arXiv:2604.19233v1 Announce Type: new Abstract: Deep learning-based object detectors have achieved remarkable success across numerous computer vision applications, yet they continue to struggle with s

researcharxiv-cs-cv
22 Apr 2026
Agents

AI-Enabled Image-Based Hybrid Vision/Force Control of Tendon-Driven Aerial Continuum Manipulators

DGX agent

arXiv:2604.18961v1 Announce Type: cross Abstract: This paper presents an AI-enabled cascaded hybrid vision/force control framework for tendon-driven aerial continuum manipulators based on constant-str

agentsarxiv-cs-cv
22 Apr 2026
Model Releases

Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval

DGX agent

arXiv:2604.19386v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) has attracted significant attention due to its flexible multimodal query method, yet its development is severely constrai

model-releasesarxiv-cs-cv
22 Apr 2026
Local Ai

Align then Refine: Text-Guided 3D Prostate Lesion Segmentation

DGX agent

arXiv:2604.18713v1 Announce Type: new Abstract: Automated 3D segmentation of prostate lesions from biparametric MRI (bp-MRI) is essential for reliable algorithmic analysis, but achieving high precisio

local-aiarxiv-cs-cv
22 Apr 2026
Safety

AlignedCut: Visual Concepts Discovery on Brain-Guided Universal Feature Space

DGX agent

arXiv:2406.18344v2 Announce Type: replace Abstract: We study the intriguing connection between visual data, deep networks, and the brain. Our method creates a universal channel alignment by using brai

safetyarxiv-cs-cv
22 Apr 2026
Safety

Allo{SR}^2: Rectifying One-Step Super-Resolution to Stay Real via Allomorphic Generative Flows

DGX agent

arXiv:2604.19238v1 Announce Type: new Abstract: Real-world image super-resolution (Real-SR) has been revolutionized by leveraging the powerful generative priors of large-scale diffusion and flow-based

safetyarxiv-cs-cv
22 Apr 2026
Local Ai

An Object-Centered Data Acquisition Method for 3D Gaussian Splatting using Mobile Phones

DGX agent

arXiv:2604.19216v1 Announce Type: new Abstract: Data acquisition through mobile phones remains a challenge for 3D Gaussian Splatting (3DGS). In this work we target the object-centered scenario and ena

local-aiarxiv-cs-cv
22 Apr 2026
Model Releases

AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model

DGX agent

arXiv:2604.19747v1 Announce Type: new Abstract: Sparse-view 3D reconstruction is essential for modeling scenes from casual captures, but remain challenging for non-generative reconstruction. Existing

model-releasesarxiv-cs-cv
22 Apr 2026
Local Ai

Attend what matters: Leveraging vision foundational models for breast cancer classification using mammograms

DGX agent

arXiv:2604.19350v1 Announce Type: new Abstract: Vision Transformers (exttt{ViT}) have become the architecture of choice for many computer vision tasks, yet their performance in computer-aided diagnost

local-aiarxiv-cs-cv
22 Apr 2026
Local Ai

Autonomous Skeletal Landmark Localization towards Agentic C-Arm Control

DGX agent

arXiv:2604.18740v1 Announce Type: new Abstract: Purpose: Automated C-arm positioning ensures timely treatment in patients requiring emergent interventions. When a conventional Deep Learning (DL) appro

local-aiarxiv-cs-cv
22 Apr 2026
Model Releases

BALTIC: A Benchmark and Cross-Domain Strategy for 3D Reconstruction Across Air and Underwater Domains Under Varying Illumination

DGX agent

arXiv:2604.19133v1 Announce Type: new Abstract: Robust 3D reconstruction across varying environmental conditions remains a critical challenge for robotic perception, particularly when transitioning be

model-releasesarxiv-cs-cv
22 Apr 2026
Research

Benchmarking Vision Foundation Models for Domain-Generalizable Face Anti-Spoofing

DGX agent

arXiv:2604.19196v1 Announce Type: new Abstract: Face Anti-Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends lev

researcharxiv-cs-cv
22 Apr 2026
Model Releases

Bridging Foundation Models and ASTM Metallurgical Standards for Automated Grain Size Estimation from Microscopy Images

DGX agent

arXiv:2604.18957v1 Announce Type: new Abstract: Extracting standardized metallurgical metrics from microscopy images remains challenging due to complex grain morphology and the data demands of supervi

model-releasesarxiv-cs-cv
22 Apr 2026
Safety

Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing

DGX agent

arXiv:2512.19302v2 Announce Type: replace Abstract: Large Vision--Language Models (LVLMs) hold great promise for advancing optical remote sensing (RS) analysis, yet existing reasoning segmentation fra

safetyarxiv-cs-cv
22 Apr 2026
Safety

CAHAL: Clinically Applicable resolution enHAncement for Low-resolution MRI scans

DGX agent

arXiv:2604.18781v1 Announce Type: new Abstract: Large-scale automated morphometric analysis of brain MRI is limited by the thick-slice, anisotropic acquisitions prevalent in routine clinical practice.

safetyarxiv-cs-cv
22 Apr 2026
Local Ai

Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow Matching

DGX agent

arXiv:2604.18623v1 Announce Type: new Abstract: Scene Graph Generation (SGG) unifies object localization and visual relationship reasoning by predicting boxes and subject-predicate-object triples. Yet

local-aiarxiv-cs-cv
22 Apr 2026
Research

Centralized Copy-Paste: Enhanced Data Augmentation Strategy for Wildland Fire Semantic Segmentation

DGX agent

arXiv:2507.06321v2 Announce Type: replace Abstract: Collecting and annotating images for the purpose of training segmentation models is often cost prohibitive. In the domain of wildland fire science,

researcharxiv-cs-cv
22 Apr 2026
Agents

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation

DGX agent

arXiv:2604.19741v1 Announce Type: new Abstract: We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video

agentsarxiv-cs-cv
22 Apr 2026
Model Releases

CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation

DGX agent

arXiv:2602.20409v2 Announce Type: replace Abstract: Recent vision-language models (VLMs) such as CLIP demonstrate impressive cross-modal reasoning, extending beyond images to 3D perception. Yet, these

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation

DGX agent

arXiv:2604.19636v1 Announce Type: new Abstract: Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, curren

model-releasesarxiv-cs-cv
22 Apr 2026
Research

Colour Extraction Pipeline for Odonates using Computer Vision

DGX agent

arXiv:2604.18725v1 Announce Type: new Abstract: The correlation between insect morphological traits and climate has been documented in physiological studies, but such studies remain limited by the tim

researcharxiv-cs-cv
22 Apr 2026
Model Releases

Concept Inconsistency in Dermoscopic Concept Bottleneck Models: A Rough-Set Analysis of the Derm7pt Dataset

DGX agent

arXiv:2604.19323v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) route predictions exclusively through a clinically grounded concept layer, binding interpretability to concept-label

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

ConvVitMamba: Efficient Multiscale Convolution, Transformer, and Mamba-Based Sequence modelling for Hyperspectral Image Classification

DGX agent

arXiv:2604.18856v1 Announce Type: new Abstract: Hyperspectral image (HSI) classification remains challenging due to high spectral dimensionality, redundancy, and limited labeled data. Although convolu

model-releasesarxiv-cs-cv
22 Apr 2026
Safety

CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers

DGX agent

arXiv:2604.19632v1 Announce Type: new Abstract: Graphic design images consist of multiple editable layers, such as text, background, and decorative elements, while most generative models produce raste

safetyarxiv-cs-cv
22 Apr 2026
Model Releases

CrossPan: A Comprehensive Benchmark for Cross-Sequence Pancreas MRI Segmentation and Generalization

DGX agent

arXiv:2604.18797v1 Announce Type: new Abstract: Automatic pancreas segmentation is fundamental to abdominal MRI analysis, yet deep learning models trained on one MRI sequence often fail catastrophical

model-releasesarxiv-cs-cv
22 Apr 2026
Research

Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasets

DGX agent

arXiv:2304.02296v2 Announce Type: replace Abstract: In our study, we conducted a comprehensive analysis of three widely used datasets in the domain of building footprint extraction using deep neural n

researcharxiv-cs-cv
22 Apr 2026
Model Releases

DDF2Pol: A Dual-Domain Feature Fusion Network for PolSAR Image Classification

DGX agent

arXiv:2604.18853v1 Announce Type: new Abstract: This paper presents DDF2Pol, a lightweight dual-domain convolutional neural network for PolSAR image classification. The proposed architecture integrate

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

Deep sprite-based image models: An analysis

DGX agent

arXiv:2604.19480v1 Announce Type: new Abstract: While foundation models drive steady progress in image segmentation and diffusion algorithms compose always more realistic images, the seemingly simple

model-releasesarxiv-cs-cv
22 Apr 2026
← Previous
1…221222223224225…261
Next →