AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Local Ai

2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA

DGX agent

arXiv:2604.23935v1 Announce Type: new Abstract: Audio-based video object segmentation aims to locate and segment objects in videos conditioned on audio cues, requiring precise understanding of both ap

local-aiarxiv-cs-cv
28 Apr 2026
Model Releases

6thGrid-Net: Unified Remote Sensing Image Dehazing Based on Color Restoration and Edge-Preserving

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2604.24149v1 Announce Type: new Abstract: Remote sensing images are frequently degraded by adverse weather conditions, particularly clouds and haze, which severely impair downstream applications

model-releasesarxiv-cs-cv
28 Apr 2026
Local Ai

A Digital Pathology Resource for Liver Cancer Quantification with Datasets, Benchmarks, and Tools

DGX agent

arXiv:2604.22858v1 Announce Type: new Abstract: Liver cancer, especially hepatocellular carcinoma (HCC), imposes a substantial global disease burden. Accurate diagnosis and prognostic assessment direc

local-aiarxiv-cs-cv
28 Apr 2026
Local Ai

A Graph-Augmented knowledge Distillation based Dual-Stream Vision Transformer with Region-Aware Attention for Gastrointestinal Disease Classification with Explainable AI

DGX agent

arXiv:2512.21372v2 Announce Type: replace-cross Abstract: The accurate classification of gastrointestinal diseases from endoscopic and histopathological imagery remains a significant challenge in medi

local-aiarxiv-cs-cv
28 Apr 2026
Research

A Heterogeneous Two-Stream Framework for Video Action Recognition with Comparative Fusion Analysis

DGX agent

arXiv:2604.23415v1 Announce Type: new Abstract: Most two-stream action recognition networks apply the same convolutional backbone to both RGB and optical flow streams, ignoring the fact that the two m

researcharxiv-cs-cv
28 Apr 2026
Applications

A Hierarchical Ensemble Inference Pipeline for Robust White Blood Cell Classification Under Domain Shifts

DGX agent

arXiv:2604.23271v1 Announce Type: new Abstract: Automated white blood cell (WBC) classification is essential for scalable leukaemia screening. However, real-world deployment is challenged by domain sh

applicationsarxiv-cs-cv
28 Apr 2026
Model Releases

A Hierarchical Self-Consistent Regularization Approach to Satellite Image Time Series Classification

DGX agent

arXiv:2510.04916v2 Announce Type: replace Abstract: Deep learning has become increasingly important in remote sensing image classification due to its ability to extract semantic information from compl

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

A Pose-only Geometric Constraint for Multi-Camera Pose Adjustment

DGX agent

arXiv:2604.23704v1 Announce Type: new Abstract: Multi-camera systems offer rich observation capabilities for visual navigation and 3D scene reconstruction; however, the resulting feature redundancy of

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

A satellite foundation model for improved wealth monitoring

DGX agent

arXiv:2604.23166v1 Announce Type: cross Abstract: Poverty statistics guide social policy, but in many low- and middle-income countries, censuses and household surveys that collect these data are costl

model-releasesarxiv-cs-cv
28 Apr 2026
Research

A Synergistic CNN-Transformer Network with Pooling Attention Fusion for Hyperspectral Image Classification

DGX agent

arXiv:2604.23622v1 Announce Type: new Abstract: In the hyperspectral image (HSI) classification task, each pixel is categorized into a specific land-cover category or material. Convolutional neural ne

researcharxiv-cs-cv
28 Apr 2026
Research

A Topology fixated Shape Gradient Framework for Non Simple Boundary Extraction for CIE Lab color images with Repulsive Energy

DGX agent

arXiv:2604.23167v1 Announce Type: new Abstract: A levelset free but a hybrid image segmentation approach based on a modified version of the piece wise constant shape gradient of an Mumford Shah shape

researcharxiv-cs-cv
28 Apr 2026
Research

Accelerating New Product Introduction for Visual Quality Inspection via Few-Shot Diffusion-Based Defect Synthesis

DGX agent

arXiv:2604.22850v1 Announce Type: new Abstract: Industrial visual inspection systems often suffer from a severe scarcity of labeled defect data, particularly during the early stages of New Product Int

researcharxiv-cs-cv
28 Apr 2026
Research

AD-Relight: Training-Free Banner Relighting via Illumination Translation with Diffusion Priors

DGX agent

arXiv:2604.24407v1 Announce Type: new Abstract: The recent surge in content consumption through streaming services has driven a growing demand for personalized content. Personalized advertisements (ad

researcharxiv-cs-cv
28 Apr 2026
Model Releases

Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model

DGX agent

arXiv:2508.06206v4 Announce Type: replace-cross Abstract: Affordance grounding focuses on predicting the specific regions of objects that are associated with the actions to be performed by robots. It

model-releasesarxiv-cs-cv
28 Apr 2026
Agents

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method

DGX agent

arXiv:2604.22836v1 Announce Type: new Abstract: This report describes a Ref-VOS pipeline centered on Sa2VA and organized with explicit agent roles. The key idea is that Sa2VA should provide the first

agentsarxiv-cs-cv
28 Apr 2026
Safety

AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance

DGX agent

arXiv:2604.23909v1 Announce Type: new Abstract: Navigational aids for blind and low vision individuals struggle conveying dynamic real-world environments, leading to cognitive overload from continuous

safetyarxiv-cs-cv
28 Apr 2026
Research

An Affordable,Wearable Stereo-Eye-Tracking Platform

DGX agent

arXiv:2604.24331v1 Announce Type: new Abstract: Research on video-based eye-tracking has long explored stereo and glint-based methods, yet existing wearable eye trackers - both commercial and open-sou

researcharxiv-cs-cv
28 Apr 2026
Safety

AnemiaVision: Non-Invasive Anemia Detection via Smartphone Imagery Using EfficientNet-B3 with TrivialAugmentWide, Mixup Augmentation, and Persistent Patient History Management

DGX agent

arXiv:2604.22964v1 Announce Type: new Abstract: Anemia affects over one billion people globally and remains severely under-diagnosed in low-resource regions where laboratory blood tests are inaccessib

safetyarxiv-cs-cv
28 Apr 2026
Safety

Animalbooth: multimodal feature enhancement for animal subject personalization

DGX agent

arXiv:2509.16702v2 Announce Type: replace Abstract: Personalized animal image generation is challenging due to rich appearance cues and large morphological variability. Existing approaches often exhib

safetyarxiv-cs-cv
28 Apr 2026
Agents

Attention-Augmented YOLOv8 with Ghost Convolution for Real-Time Vehicle Detection in Intelligent Transportation Systems

DGX agent

arXiv:2604.22856v1 Announce Type: new Abstract: Accurate vehicle detection is a critical component of autonomous driving, traffic surveillance, and intelligent transportation systems. This paper prese

agentsarxiv-cs-cv
28 Apr 2026
Model Releases

ATTN-FIQA: Interpretable Attention-based Face Image Quality Assessment with Vision Transformers

DGX agent

arXiv:2604.22841v1 Announce Type: new Abstract: Face Image Quality Assessment (FIQA) aims to assess the recognition utility of face samples and is essential for reliable face recognition (FR) systems.

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

AusSmoke meets MultiNatSmoke: a fully-labelled diverse smoke segmentation dataset

DGX agent

arXiv:2604.23542v1 Announce Type: new Abstract: Wildfires are an escalating global concern due to the devastating impacts on the environment, economy, and human health, with notable incidents such as

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark

DGX agent

arXiv:2604.24441v1 Announce Type: new Abstract: Autonomous agents capable of navigating Graphical User Interfaces (GUIs) hold the potential to revolutionize digital productivity. However, achieving tr

model-releasesarxiv-cs-cv
28 Apr 2026
Tutorials

AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering

DGX agent

arXiv:2510.18346v2 Announce Type: replace Abstract: Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse ques

tutorialsarxiv-cs-cv
28 Apr 2026
Research

Aycromo: An Open-Source Platform for Automatic Chromosome Detection in Metaphase Images Based on Deep Learning

DGX agent

arXiv:2604.24685v1 Announce Type: new Abstract: Chromosome analysis is a fundamental step in the diagnosis of genetic diseases, but the manual karyotyping workflow is time-consuming and heavily depend

researcharxiv-cs-cv
28 Apr 2026
Research

B-FIRE: Binning-Free Diffusion Implicit Neural Representation for Hyper-Accelerated Motion-Resolved MRI

DGX agent

arXiv:2601.06166v2 Announce Type: replace Abstract: Accelerated dynamic volumetric magnetic resonance imaging (4DMRI) is essential for applications relying on motion resolution. Existing 4DMRI produce

researcharxiv-cs-cv
28 Apr 2026
Model Releases

Benchmarking Pathology Foundation Models for Breast Cancer Survival Prediction

DGX agent

arXiv:2604.24679v1 Announce Type: new Abstract: Pathology foundation models (PFMs) have recently emerged as powerful pretrained encoders for computational pathology, enabling transfer learning across

model-releasesarxiv-cs-cv
28 Apr 2026
Research

BIMStruct3D: A Fully Automated Hybrid Learning Scan-to-BIM Pipeline with Integrated Topology Refinement

DGX agent

arXiv:2604.24311v1 Announce Type: new Abstract: Automatic generation of Building Information Models (BIM) from building scans is a key challenge in architecture and construction. We present a modular

researcharxiv-cs-cv
28 Apr 2026
Model Releases

BIR-Adapter: A parameter-efficient diffusion adapter for blind image restoration

DGX agent

arXiv:2509.06904v3 Announce Type: replace Abstract: We introduce the BIR-Adapter, a parameter-efficient diffusion adapter for blind image restoration. Diffusion-based restoration methods have demonstr

model-releasesarxiv-cs-cv
28 Apr 2026
Safety

BMD-45: A Large-Scale CCTV Vehicle Detection Dataset for Urban Traffic in Developing Cities

DGX agent

arXiv:2604.24419v1 Announce Type: new Abstract: Robust vehicle detection from fixed CCTV cameras is critical for Intelligent Transportation Systems. Yet existing benchmarks predominantly feature relat

safetyarxiv-cs-cv
28 Apr 2026
Model Releases

Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations

DGX agent

arXiv:2603.08592v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) have achieved remarkable success in 2D visual understanding, their ability to reason about 3D space r

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

Breaking Degradation Coupling: A Structural Entropy Guided Decoupled Framework and Benchmark for Infrared Enhancement

DGX agent

arXiv:2604.22886v1 Announce Type: new Abstract: Thermal infrared image enhancement aims to restore high-quality images from complex compound degradations. Existing all-in-one approaches typically empl

model-releasesarxiv-cs-cv
28 Apr 2026
Safety

Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training

DGX agent

arXiv:2604.23121v1 Announce Type: cross Abstract: Have you ever post-trained a generalist vision-language-action (VLA) policy on a small demonstration dataset, only to find that it stops responding to

safetyarxiv-cs-cv
28 Apr 2026
Local Ai

Breaking the Resource Wall: Geometry-Guided Sequence Modeling for Efficient Semantic Segmentation

DGX agent

arXiv:2604.23399v1 Announce Type: new Abstract: High-performance semantic segmentation has achieved significant progress in recent years, often driven by increasingly large backbones and higher comput

local-aiarxiv-cs-cv
28 Apr 2026
Research

Breaking the Scalability Limit of Multi-Projector Calibration with Embedded Cameras

DGX agent

arXiv:2604.24024v1 Announce Type: new Abstract: Conventional multi-projector calibration requires projecting and capturing structured light patterns for each projector sequentially, causing calibratio

researcharxiv-cs-cv
28 Apr 2026
Research

BrickNet: Graph-Backed Generative Brick Assembly

DGX agent

arXiv:2604.22984v1 Announce Type: new Abstract: We train a language model to generate LEGO-brick build sequences. While prior work has been restricted to discrete, voxel-like towers, we consider a muc

researcharxiv-cs-cv
28 Apr 2026
Applications

Bridging Restoration and Generation Manifolds in One-Step Diffusion for Real-World Super-Resolution

DGX agent

arXiv:2604.24136v1 Announce Type: new Abstract: Pretrained diffusion models have revolutionized real-world image super-resolution (Real-ISR) but suffer from computational bottlenecks due to iterative

applicationsarxiv-cs-cv
28 Apr 2026
Model Releases

Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search

DGX agent

arXiv:2604.23282v1 Announce Type: new Abstract: Text-based person anomaly search retrieves specific behavioral events from surveillance archives using natural-language queries. Although recent pose-aw

model-releasesarxiv-cs-cv
28 Apr 2026
Research

Bringing a Personal Point of View: Evaluating Dynamic 3D Gaussian Splatting for Egocentric Scene Reconstruction

DGX agent

arXiv:2604.23803v1 Announce Type: new Abstract: Egocentric video provides a unique view into human perception and interaction, with growing relevance for augmented reality, robotics, and assistive tec

researcharxiv-cs-cv
28 Apr 2026
Research

BSViT: A Burst Spiking Vision Transformer for Expressive and Efficient Visual Representation Learning

DGX agent

arXiv:2604.23165v1 Announce Type: new Abstract: Spiking Vision Transformers (S-ViTs) offer a promising framework for energy-efficient visual learning. However, existing designs remain limited by two f

researcharxiv-cs-cv
28 Apr 2026
Research

BurstGP: Enhancing Raw Burst Image Super Resolution with Generative Priors

DGX agent

arXiv:2604.23508v1 Announce Type: new Abstract: Burst image super resolution (BISR) aims to construct a single high-resolution (HR) image by aggregating information from multiple low-resolution (LR) f

researcharxiv-cs-cv
28 Apr 2026
Safety

BVI-Mamba: Video Enhancement Using a Visual State-Space Model for Low-Light and Underwater Environments

DGX agent

arXiv:2604.23655v1 Announce Type: new Abstract: Videos captured in low-light and underwater conditions often suffer from distortions such as noise, low contrast, color imbalance, and blur. These issue

safetyarxiv-cs-cv
28 Apr 2026
Safety

CA-IDD: Cross-Attention Guided Identity-Conditional Diffusion for Identity-Consistent Face Swapping

DGX agent

arXiv:2604.24493v1 Announce Type: new Abstract: Face swapping aims to optimize realistic facial image generation by leveraging the identity of a source face onto a target face while preserving pose, e

safetyarxiv-cs-cv
28 Apr 2026
Research

Caries DETR: Tooth Structure-aware Prior and Lesion-aware Dynamic Loss Refinement for DETR Based Caries Detection

DGX agent

arXiv:2604.23718v1 Announce Type: new Abstract: As dental caries appear as subtle, low-contrast lesions in intraoral imaging, existing deep learning models face significant challenges in the early det

researcharxiv-cs-cv
28 Apr 2026
Applications

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM

DGX agent

arXiv:2603.27507v2 Announce Type: replace Abstract: Recent advancements in multi-modal large language models (MLLMs) have shown strong potential for 3D scene understanding. However, existing methods s

applicationsarxiv-cs-cv
28 Apr 2026
Model Releases

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

DGX agent

arXiv:2604.23781v1 Announce Type: new Abstract: Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surroundi

model-releasesarxiv-cs-cv
28 Apr 2026
Research

CLIP-Guided Data Augmentation for Night-Time Image Dehazing

DGX agent

arXiv:2604.05500v2 Announce Type: replace Abstract: Nighttime image dehazing faces a more complex degradation pattern than its daytime counterpart, as haze scattering couples with low illumination, no

researcharxiv-cs-cv
28 Apr 2026
Agents

CLLAP: Contrastive Learning-based LiDAR-Augmented Pretraining for Enhanced Radar-Camera Fusion

DGX agent

arXiv:2604.24044v1 Announce Type: new Abstract: Accurate 3D object detection is critical for autonomous driving, necessitating reliable, cost-effective sensors capable of operating in adverse weather

agentsarxiv-cs-cv
28 Apr 2026
← Previous
1…210211212213214…261
Next →