AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

PLAS-Net: Pixel-Level Area Segmentation for UAV-Based Beach Litter Monitoring

DGX agent

arXiv:2604.21313v1 Announce Type: new Abstract: Accurate quantification of the physical exposure area of beach litter, rather than simple item counts, is essential for credible ecological risk assessm

researcharxiv-cs-cv
24 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Pre-process for segmentation task with nonlinear diffusion filters

DGX agent

arXiv:2604.21422v1 Announce Type: new Abstract: This paper deals with the case of using nonlinear diffusion filters to obtain piecewise constant images as a previous process for segmentation technique

researcharxiv-cs-cv
24 Apr 2026
Model Releases

Pretrain Where? Investigating How Pretraining Data Diversity Impacts Geospatial Foundation Model Performance

DGX agent

arXiv:2604.21104v1 Announce Type: new Abstract: New geospatial foundation models introduce a new model architecture and pretraining dataset, often sampled using different notions of data diversity. Pe

model-releasesarxiv-cs-cv
24 Apr 2026
Research

Projected Gradient Unlearning for Text-to-Image Diffusion Models: Defending Against Concept Revival Attacks

DGX agent

arXiv:2604.21041v1 Announce Type: new Abstract: Machine unlearning for text-to-image diffusion models aims to selectively remove undesirable concepts from pre-trained models without costly retraining.

researcharxiv-cs-cv
24 Apr 2026
Research

Prototype-Based Test-Time Adaptation of Vision-Language Models

DGX agent

arXiv:2604.21360v1 Announce Type: new Abstract: Test-time adaptation (TTA) has emerged as a promising paradigm for vision-language models (VLMs) to bridge the distribution gap between pre-training and

researcharxiv-cs-cv
24 Apr 2026
Model Releases

RailVQA: A Benchmark and Framework for Efficient Interpretable Visual Cognition in Automatic Train Operation

DGX agent

arXiv:2603.27112v2 Announce Type: replace Abstract: As Automatic Train Operation (ATO) advances toward GoA4 and beyond, it increasingly depends on efficient, reliable cab-view visual perception and de

model-releasesarxiv-cs-cv
24 Apr 2026
Safety

Ramen: Robust Test-Time Adaptation of Vision-Language Models with Active Sample Selection

DGX agent

arXiv:2604.21728v1 Announce Type: new Abstract: Pretrained vision-language models such as CLIP exhibit strong zero-shot generalization but remain sensitive to distribution shifts. Test-time adaptation

safetyarxiv-cs-cv
24 Apr 2026
Model Releases

Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment

DGX agent

arXiv:2604.21160v1 Announce Type: new Abstract: Point-Vision-Language Models promise to empower embodied agents with executable spatial reasoning, yet they frequently succumb to geometric hallucinatio

model-releasesarxiv-cs-cv
24 Apr 2026
Tutorials

Reshoot-Anything: A Self-Supervised Model for In-the-Wild Video Reshooting

DGX agent

arXiv:2604.21776v1 Announce Type: new Abstract: Precise camera control for reshooting dynamic videos is bottlenecked by the severe scarcity of paired multi-view data for non-rigid scenes. We overcome

tutorialsarxiv-cs-cv
24 Apr 2026
Safety

Rethinking Cross-Domain Evaluation for Face Forgery Detection with Semantic Fine-grained Alignment and Mixture-of-Experts

DGX agent

arXiv:2604.21478v1 Announce Type: new Abstract: Nowadays, visual data forgery detection plays an increasingly important role in social and economic security with the rapid development of generative mo

safetyarxiv-cs-cv
24 Apr 2026
Tutorials

S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images

DGX agent

arXiv:2604.21409v1 Announce Type: new Abstract: We present S1-VL, a multimodal reasoning model for scientific domains that natively supports two complementary reasoning paradigms: Scientific Reasoning

tutorialsarxiv-cs-cv
24 Apr 2026
Tutorials

Sapiens2

DGX agent

arXiv:2604.21681v1 Announce Type: new Abstract: We present Sapiens2, a model family of high-resolution transformers for human-centric vision focused on generalization, versatility, and high-fidelity o

tutorialsarxiv-cs-cv
24 Apr 2026
Model Releases

SatSAM2: Motion-Constrained Video Object Tracking in Satellite Imagery using Promptable SAM2 and Kalman Priors

DGX agent

arXiv:2511.18264v3 Announce Type: replace Abstract: Existing satellite video tracking methods often struggle with generalization, requiring scenario-specific training to achieve satisfactory performan

model-releasesarxiv-cs-cv
24 Apr 2026
Model Releases

SCASeg: Strip Cross-Attention for Efficient Semantic Segmentation

DGX agent

arXiv:2411.17061v2 Announce Type: replace Abstract: The Vision Transformer (ViT) has achieved notable success in computer vision, with its variants widely validated across various downstream tasks, in

model-releasesarxiv-cs-cv
24 Apr 2026
Research

Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers

DGX agent

arXiv:2604.21592v1 Announce Type: new Abstract: Recent breakthroughs in 3D generative modeling have yielded remarkable progress in static shape synthesis, yet high-fidelity dynamic 4D generation remai

researcharxiv-cs-cv
24 Apr 2026
Safety

Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs

DGX agent

arXiv:2604.21926v1 Announce Type: new Abstract: Understanding human activities and their surrounding environments typically relies on visual perception, yet cameras pose persistent challenges in priva

safetyarxiv-cs-cv
24 Apr 2026
Safety

SGG-R^{rm 3}: From Next-Token Prediction to End-to-End Unbiased Scene Graph Generation

DGX agent

arXiv:2603.07961v3 Announce Type: replace Abstract: Scene Graph Generation (SGG) structures visual scenes as graphs of objects and their relations. While Multimodal Large Language Models (MLLMs) have

safetyarxiv-cs-cv
24 Apr 2026
Local Ai

Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation

DGX agent

arXiv:2604.21221v1 Announce Type: new Abstract: We introduce Sparse Forcing, a training-and-inference paradigm for autoregressive video diffusion models that improves long-horizon generation quality w

local-aiarxiv-cs-cv
24 Apr 2026
Model Releases

SparseGF: A Height-Aware Sparse Segmentation Framework with Context Compression for Robust Ground Filtering Across Urban to Natural Scenes

DGX agent

arXiv:2604.21356v1 Announce Type: new Abstract: High-quality digital terrain models derived from airborne laser scanning (ALS) data are essential for a wide range of geospatial analyses, and their gen

model-releasesarxiv-cs-cv
24 Apr 2026
Agents

SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning

DGX agent

arXiv:2604.21190v1 Announce Type: new Abstract: Understanding visual scenes requires not only recognizing objects but also reasoning about their spatial relationships. Unlike general vision-language t

agentsarxiv-cs-cv
24 Apr 2026
Model Releases

StyleID: A Perception-Aware Dataset and Metric for Stylization-Agnostic Facial Identity Recognition

DGX agent

arXiv:2604.21689v1 Announce Type: cross Abstract: Creative face stylization aims to render portraits in diverse visual idioms such as cartoons, sketches, and paintings while retaining recognizable ide

model-releasesarxiv-cs-cv
24 Apr 2026
Model Releases

SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding

DGX agent

arXiv:2511.03325v3 Announce Type: replace Abstract: Video Question Answering (VideoQA) in the surgical domain aims to enhance intraoperative understanding by enabling AI models to reason over temporal

model-releasesarxiv-cs-cv
24 Apr 2026
Local Ai

Teacher-Guided Routing for Sparse Vision Mixture-of-Experts

DGX agent

arXiv:2604.21330v1 Announce Type: new Abstract: Recent progress in deep learning has been driven by increasingly large-scale models, but the resulting computational cost has become a critical bottlene

local-aiarxiv-cs-cv
24 Apr 2026
Model Releases

TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval

DGX agent

arXiv:2604.21806v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) is an important image retrieval paradigm that enables users to retrieve a target image using a multimodal query that cons

model-releasesarxiv-cs-cv
24 Apr 2026
Safety

Temporal Prototyping and Hierarchical Alignment for Unsupervised Video-based Visible-Infrared Person Re-Identification

DGX agent

arXiv:2604.21324v1 Announce Type: new Abstract: Visible-infrared person re-identification (VI-ReID) enables cross-modality identity matching for all-day surveillance, yet existing methods predominantl

safetyarxiv-cs-cv
24 Apr 2026
Model Releases

TimePre: Bridging Accuracy, Efficiency, and Stability in Probabilistic Time-Series Forecasting

DGX agent

arXiv:2511.18539v2 Announce Type: replace-cross Abstract: We propose TimePre, a simple framework that unifies the efficiency of Multilayer Perceptron (MLP)-based models with the distributional flexibi

model-releasesarxiv-cs-cv
24 Apr 2026
Research

Transformer-Progressive Mamba Network for Lightweight Image Super-Resolution

DGX agent

arXiv:2511.03232v2 Announce Type: replace Abstract: Recently, Mamba-based super-resolution (SR) methods have demonstrated the ability to capture global receptive fields with linear complexity, address

researcharxiv-cs-cv
24 Apr 2026
Safety

Tumor-anchored deep feature random forests for out-of-distribution detection in lung cancer segmentation

DGX agent

arXiv:2512.08216v3 Announce Type: replace-cross Abstract: Accurate segmentation of lung tumors from 3D computed tomography (CT) scans is essential for automated treatment planning and response assessm

safetyarxiv-cs-cv
24 Apr 2026
Applications

TV Subgradient-Guided Multi-Source Fusion for Spectral Imaging in Dual-Camera CASSI Systems

DGX agent

arXiv:2509.10897v2 Announce Type: replace Abstract: Balancing spectral, spatial, and temporal resolutions is a key challenge in spectral imaging. The Dual-Camera Coded Aperture Snapshot Spectral Imagi

applicationsarxiv-cs-cv
24 Apr 2026
Research

UAU-Net: Uncertainty-aware Representation Learning and Evidential Classification for Facial Action Unit Detection

DGX agent

arXiv:2604.21227v1 Announce Type: new Abstract: Facial action unit (AU) detection remains challenging because it involves heterogeneous, AU-specific uncertainties arising at both the representation an

researcharxiv-cs-cv
24 Apr 2026
Local Ai

UHR-DETR: Efficient End-to-End Small Object Detection for Ultra-High-Resolution Remote Sensing Imagery

DGX agent

arXiv:2604.21435v1 Announce Type: new Abstract: Ultra-High-Resolution (UHR) imagery has become essential for modern remote sensing, offering unprecedented spatial coverage. However, detecting small ob

local-aiarxiv-cs-cv
24 Apr 2026
Safety

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection

DGX agent

arXiv:2604.21904v1 Announce Type: new Abstract: In recent years, significant progress has been made in both image generation and generated image detection. Despite their rapid, yet largely independent

safetyarxiv-cs-cv
24 Apr 2026
Model Releases

Unlocking Multi-Spectral Data for Multi-Modal Models with Guided Inputs and Chain-of-Thought Reasoning

DGX agent

arXiv:2604.21032v1 Announce Type: new Abstract: Multi-spectral imagery is a valuable input signal for Remote Sensing applications, such as land-use and land-cover classification and environmental moni

model-releasesarxiv-cs-cv
24 Apr 2026
Local Ai

Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation

DGX agent

arXiv:2604.21713v1 Announce Type: new Abstract: Feed-forward visual geometry estimation has recently made rapid progress. However, an important gap remains: multi-frame models usually produce better c

local-aiarxiv-cs-cv
24 Apr 2026
Model Releases

Unsharp Measurement with Adaptive Gaussian POVMs for Quantum-Inspired Image Processing

DGX agent

arXiv:2604.04685v2 Announce Type: replace-cross Abstract: We propose a data-adaptive probabilistic intensity remapping framework for structure-preserving transformation of grayscale images. The sugges

model-releasesarxiv-cs-cv
24 Apr 2026
Safety

VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models

DGX agent

arXiv:2510.18457v3 Announce Type: replace Abstract: The performance of Latent Diffusion Models (LDMs) is critically dependent on the quality of their visual tokenizers. While recent works have explore

safetyarxiv-cs-cv
24 Apr 2026
Safety

VFM^{4}SDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection

DGX agent

arXiv:2604.21502v1 Announce Type: new Abstract: In real-world scenarios, continual changes in weather, illumination, and imaging conditions cause significant domain shifts, leading detectors trained o

safetyarxiv-cs-cv
24 Apr 2026
Model Releases

VidHal: Benchmarking Temporal Hallucinations in Vision LLMs

DGX agent

arXiv:2411.16771v3 Announce Type: replace Abstract: Vision Large Language Models (VLLMs) are widely acknowledged to be prone to hallucinations. Existing research addressing this problem has primarily

model-releasesarxiv-cs-cv
24 Apr 2026
Applications

Vista4D: Video Reshooting with 4D Point Clouds

DGX agent

arXiv:2604.21915v1 Announce Type: new Abstract: We present Vista4D, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically,

applicationsarxiv-cs-cv
24 Apr 2026
Model Releases

WFM: 3D Wavelet Flow Matching for Ultrafast Multi-Modal MRI Synthesis

DGX agent

arXiv:2604.21146v1 Announce Type: new Abstract: Diffusion models have achieved remarkable quality in multi-modal MRI synthesis, but their computational cost (hundreds of sampling steps and separate mo

model-releasesarxiv-cs-cv
24 Apr 2026
Applications

WildSplatter: Feed-forward 3D Gaussian Splatting with Appearance Control from Unconstrained Images

DGX agent

arXiv:2604.21182v1 Announce Type: new Abstract: We propose WildSplatter, a feed-forward 3D Gaussian Splatting (3DGS) model for unconstrained images with unknown camera parameters and varying lighting

applicationsarxiv-cs-cv
24 Apr 2026
Model Releases

WorldMark: A Unified Benchmark Suite for Interactive Video World Models

DGX agent

arXiv:2604.21686v1 Announce Type: new Abstract: Interactive video generation models such as Genie, YUME, HY-World, and Matrix-Game are advancing rapidly, yet every model is evaluated on its own benchm

model-releasesarxiv-cs-cv
24 Apr 2026
Model Releases

You Only Gaussian Once: Controllable 3D Gaussian Splatting for Ultra-Densely Sampled Scenes

DGX agent

arXiv:2604.21400v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has revolutionized neural rendering, yet existing methods remain predominantly research prototypes ill-suited for productio

model-releasesarxiv-cs-cv
24 Apr 2026
Research

3D Smoke Scene Reconstruction Guided by Vision Priors from Multimodal Large Language Models

DGX agent

arXiv:2604.05687v2 Announce Type: replace Abstract: Reconstructing 3D scenes from smoke-degraded multi-view images is particularly difficult because smoke introduces strong scattering effects, view-de

researcharxiv-cs-cv
23 Apr 2026
Research

A Computational Model of Message Sensation Value in Short Video Multimodal Features that Predicts Sensory and Behavioral Engagement

DGX agent

arXiv:2604.19995v1 Announce Type: new Abstract: The contemporary media landscape is characterized by sensational short videos. While prior research examines the effects of individual multimodal featur

researcharxiv-cs-cv
23 Apr 2026
Research

A novel attention mechanism for noise-adaptive and robust segmentation of microtubules in microscopy images

DGX agent

arXiv:2507.07800v3 Announce Type: replace-cross Abstract: Segmenting cytoskeletal filaments in microscopy images is essential for studying their roles in cellular processes. However, this task is high

researcharxiv-cs-cv
23 Apr 2026
Safety

A Synchronized Audio-Visual Multi-View Capture System

DGX agent

arXiv:2603.23089v2 Announce Type: replace Abstract: Multi-view capture systems have been an important tool in research for recording human motion under controlling conditions. Most existing systems ar

safetyarxiv-cs-cv
23 Apr 2026
Model Releases

Adapting TrOCR for Printed Tigrinya Text Recognition: Word-Aware Loss Weighting for Cross-Script Transfer Learning

DGX agent

arXiv:2604.20813v1 Announce Type: new Abstract: Transformer-based OCR models have shown strong performance on Latin and CJK scripts, but their application to African syllabic writing systems remains l

model-releasesarxiv-cs-cv
23 Apr 2026
← Previous
1…218219220221222…261
Next →