AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

See Through the Noise: Improving Domain Generalization in Gaze Estimation

DGX agent

arXiv:2604.16562v1 Announce Type: new Abstract: Generalizable gaze estimation methods have garnered increasing attention due to their critical importance in real-world applications and have achieved s

safetyarxiv-cs-cv
21 Apr 2026
Research
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

SegTTA: Training-Free Test-Time Augmentation for Zero-Shot Medical Imaging Segmentation

DGX agent

arXiv:2604.17451v1 Announce Type: new Abstract: Increasingly advanced data augmentation techniques have greatly aided clinical medical research, increasing data diversity and improving model generaliz

researcharxiv-cs-cv
21 Apr 2026
Agents

Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation

DGX agent

arXiv:2604.16958v1 Announce Type: new Abstract: Narrative-driven product photography has become a prevalent paradigm in modern marketing, as coherent visual storytelling helps convey product value and

agentsarxiv-cs-cv
21 Apr 2026
Research

Self-Supervised Super-Resolution for Sentinel-5P Hyperspectral Images

DGX agent

arXiv:2604.17652v1 Announce Type: new Abstract: Sentinel-5P (S5P) plays a critical role in atmospheric monitoring; however, its spatial resolution limits fine-scale analysis. Existing super-resolution

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Semantically Stable Image Composition Analysisvia Saliency and Gradient Vector Flow Fusion

DGX agent

arXiv:2604.16500v1 Announce Type: new Abstract: The reliable computational assessment of photographic composition requires features that are discriminative of spatial layout yet robust to semantic con

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

DGX agent

arXiv:2604.18476v1 Announce Type: new Abstract: Camera-only 3D object detection has emerged as a cost-effective and scalable alternative to LiDAR for autonomous driving, yet existing methods primarily

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

SemMorph3D: Unsupervised Semantic-Aware 3D Morphing via Mesh-Guided Gaussians

DGX agent

arXiv:2510.02034v2 Announce Type: replace Abstract: We introduce METHODNAME, a novel framework for semantic-aware 3D shape and texture morphing directly from multi-view images. While 3D Gaussian Splat

model-releasesarxiv-cs-cv
21 Apr 2026
Research

SentiAvatar: Towards Expressive and Interactive Digital Humans

DGX agent

arXiv:2604.02908v2 Announce Type: replace Abstract: We present SentiAvatar, a framework for building expressive interactive 3D digital humans, and use it to create SuSu, a virtual character that speak

researcharxiv-cs-cv
21 Apr 2026
Model Releases

SetFlow: Generating Structured Sets of Representations for Multiple Instance Learning

DGX agent

arXiv:2604.16362v1 Announce Type: cross Abstract: Data scarcity and weak supervision continue to limit the performance of machine learning models in many real-world applications, such as mammography,

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Sharpening Lightweight Models for Generalized Polyp Segmentation: A Boundary Guided Distillation from Foundation Models

DGX agent

arXiv:2604.17865v1 Announce Type: new Abstract: Automated polyp segmentation is critical for early colorectal cancer detection and its prevention, yet remains challenging due to weak boundaries, large

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation

DGX agent

arXiv:2511.10370v2 Announce Type: replace Abstract: Geospatial foundation models (GFMs) for Earth observation often fail to perform reliably in environments underrepresented during pretraining. We int

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

SIF: Semantically In-Distribution Fingerprints for Large Vision-Language Models

DGX agent

arXiv:2604.17041v1 Announce Type: new Abstract: The public accessibility of large vision-language models (LVLMs) raises serious concerns about unauthorized model reuse and intellectual property infrin

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Sky2Ground: A Benchmark for Site Modeling under Varying Altitude

DGX agent

arXiv:2603.13740v3 Announce Type: replace Abstract: We introduce Sky2Ground, a three-view dataset designed for varying altitude camera localization, correspondence learning, and reconstruction. The da

model-releasesarxiv-cs-cv
21 Apr 2026
Research

SMILE-UHURA Challenge -- Small Vessel Segmentation at Mesoscopic Scale from Ultra-High Resolution 7T Magnetic Resonance Angiograms

DGX agent

arXiv:2411.09593v2 Announce Type: replace-cross Abstract: The human brain receives nutrients and oxygen through an intricate network of blood vessels. Pathology affecting small vessels, at the mesosco

researcharxiv-cs-cv
21 Apr 2026
Safety

Soft Label Pruning and Quantization for Large-Scale Dataset Distillation

DGX agent

arXiv:2604.18135v1 Announce Type: new Abstract: Large-scale dataset distillation requires storing auxiliary soft labels that can be 30-40x larger on ImageNet-1K and 200x larger on ImageNet-21K than th

safetyarxiv-cs-cv
21 Apr 2026
Safety

Source-Free Domain Adaptation with Vision-Language Prior

DGX agent

arXiv:2604.17748v1 Announce Type: new Abstract: Source-Free Domain Adaptation (SFDA) seeks to adapt a source model, which is pre-trained on a supervised source domain, for a target domain, with only a

safetyarxiv-cs-cv
21 Apr 2026
Safety

(Sparse) Attention to the Details: Preserving Spectral Fidelity in ML-based Weather Forecasting Models

DGX agent

arXiv:2604.16429v1 Announce Type: cross Abstract: We introduce Mosaic, a probabilistic weather forecasting model that addresses two principal sources of spectral degradation in ML-based weather predic

safetyarxiv-cs-cv
21 Apr 2026
Local Ai

Spatial-Regularization-Aware Dual-Branch Collaborative Inference for Training-Free OVSS in Remote Sensing Imagery

DGX agent

arXiv:2601.21159v2 Announce Type: replace Abstract: High-resolution remote sensing images contain densely distributed objects with pronounced scale variations and complex boundaries, which impose high

local-aiarxiv-cs-cv
21 Apr 2026
Research

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning

DGX agent

arXiv:2604.17385v1 Announce Type: new Abstract: Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge fo

researcharxiv-cs-cv
21 Apr 2026
Local Ai

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

DGX agent

arXiv:2603.27437v2 Announce Type: replace Abstract: Large vision-language models (VLMs) still struggle with reliable 3D spatial reasoning, a core capability for embodied and physical AI systems. This

local-aiarxiv-cs-cv
21 Apr 2026
Model Releases

Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models

DGX agent

arXiv:2604.17873v1 Announce Type: new Abstract: Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Spectral Forensics of Diffusion Attention Graphs for Copy-Move Forgery Detection

DGX agent

arXiv:2604.17287v1 Announce Type: new Abstract: Copy-move forgery, where a region within an image is duplicated to hide or fabricate content, remains a persistent threat to visual media integrity. We

researcharxiv-cs-cv
21 Apr 2026
Research

Speculative Decoding for Autoregressive Video Generation

DGX agent

arXiv:2604.17397v1 Announce Type: new Abstract: Autoregressive video diffusion is emerging as a promising paradigm for streaming video synthesis, with step distillation serving as the primary means of

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Spike-NVPT: Learning Robust Visual Prompts via Bio-Inspired Temporal Filtering and Discretization

DGX agent

arXiv:2604.18284v1 Announce Type: new Abstract: Pre-trained vision models have found widespread application across diverse domains. Prompt tuning-based methods have emerged as a parameter-efficient pa

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Splatography: Sparse multi-view dynamic Gaussian Splatting for filmmaking challenges

DGX agent

arXiv:2511.05152v2 Announce Type: replace Abstract: Deformable Gaussian Splatting (GS) accomplishes photorealistic dynamic 3-D reconstruction from dense multi-view video (MVV) by learning to deform a

researcharxiv-cs-cv
21 Apr 2026
Local Ai

ST-pi: Structured SpatioTemporal VLA for Robotic Manipulation

DGX agent

arXiv:2604.17880v1 Announce Type: cross Abstract: Vision-language-action (VLA) models have achieved great success on general robotic tasks, but still face challenges in fine-grained spatiotemporal man

local-aiarxiv-cs-cv
21 Apr 2026
Research

StableMTL: Repurposing Latent Diffusion Models for Multi-Task Learning from Partially Annotated Synthetic Datasets

DGX agent

arXiv:2506.08013v2 Announce Type: replace Abstract: Multi-task learning for dense prediction is limited by the need for extensive annotation for every task, though recent works have explored training

researcharxiv-cs-cv
21 Apr 2026
Local Ai

Structure-Adaptive Sparse Diffusion in Voxel Space for 3D Medical Image Enhancement

DGX agent

arXiv:2604.17773v1 Announce Type: new Abstract: Three-dimensional (3D) medical image enhancement, including denoising and super-resolution, is critical for clinical diagnosis in CT, PET, and MRI. Alth

local-aiarxiv-cs-cv
21 Apr 2026
Research

Structured 3D-SVD: A Practical Framework for the Compression and Reconstruction of Biological Volumetric Images

DGX agent

arXiv:2604.16947v1 Announce Type: cross Abstract: This work introduces Structured 3D-SVD as a practical framework for the reconstruction, compression, and analysis of biological volumetric data. Inspi

researcharxiv-cs-cv
21 Apr 2026
Research

Style-Based Neural Architectures for Real-Time Weather Classification

DGX agent

arXiv:2604.18251v1 Announce Type: new Abstract: In this paper, we present three neural network architectures designed for real-time classification of weather conditions (sunny, rain, snow, fog) from i

researcharxiv-cs-cv
21 Apr 2026
Safety

Sub-metre Lunar DEM Generation and Validation from Chandrayaan-2 OHRC Multi-View Imagery Using an Open-Source Pipeline

DGX agent

arXiv:2604.01032v3 Announce Type: replace Abstract: High-resolution digital elevation models (DEMs) of the lunar surface are essential for surface mobility planning, landing site characterization, and

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

Subject-Aware Multi-Granularity Alignment for Zero-Shot EEG-to-Image Retrieval

DGX agent

arXiv:2604.17782v1 Announce Type: new Abstract: Zero-shot EEG-to-image retrieval aims to decode perceived visual content from electroencephalography (EEG) by aligning neural responses with pretrained

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

SynAgent: Generalizable Cooperative Humanoid Manipulation via Solo-to-Cooperative Agent Synergy

DGX agent

arXiv:2604.18557v1 Announce Type: new Abstract: Controllable cooperative humanoid manipulation is a fundamental yet challenging problem for embodied intelligence, due to severe data scarcity, complexi

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

SynthPID: P&ID digitization from Topology-Preserving Synthetic Data

DGX agent

arXiv:2604.16513v1 Announce Type: new Abstract: Automating the digitization of Piping and Instrumentation Diagrams (P&IDs) into structured process graphs would unlock significant value in plant operat

model-releasesarxiv-cs-cv
21 Apr 2026
Local Ai

T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability

DGX agent

arXiv:2604.18573v1 Announce Type: new Abstract: Despite recent progress, vision-language encoders struggle with two core limitations: (1) weak alignment between language and dense vision features, whi

local-aiarxiv-cs-cv
21 Apr 2026
Model Releases

TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation

DGX agent

arXiv:2603.02972v2 Announce Type: replace Abstract: Vision-Language Navigation (VLN) presents a unique challenge for Large Vision-Language Models (VLMs) due to their inherent architectural mismatch: V

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation

DGX agent

arXiv:2604.17005v1 Announce Type: new Abstract: Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semant

safetyarxiv-cs-cv
21 Apr 2026
Research

Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models

DGX agent

arXiv:2604.18107v1 Announce Type: new Abstract: Vision-Language-Action models (VLAs) achieve remarkable performance in sequential decision-making but remain fragile to subtle environmental shifts, suc

researcharxiv-cs-cv
21 Apr 2026
Model Releases

The First Challenge on Mobile Real-World Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

DGX agent

arXiv:2604.17306v1 Announce Type: new Abstract: This paper provides a review of the NTIRE 2026 challenge on mobile real-world image super-resolution, highlighting the proposed solutions and the result

model-releasesarxiv-cs-cv
21 Apr 2026
Research

The Gait Signature of Frailty: Transfer Learning based Deep Gait Models for Scalable Frailty Assessment

DGX agent

arXiv:2603.24434v2 Announce Type: replace Abstract: Frailty is a condition in aging medicine characterized by diminished physiological reserve and increased vulnerability to stressors. However, frailt

researcharxiv-cs-cv
21 Apr 2026
Safety

The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images with Minimal 3D Knowledge

DGX agent

arXiv:2506.09885v2 Announce Type: replace Abstract: Recent advances in feed-forward Novel View Synthesis (NVS) have led to a divergence between two design philosophies: bias-driven methods, which rely

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

TimeColor: Flexible Reference Colorization via Temporal Concatenation

DGX agent

arXiv:2601.00296v2 Announce Type: replace Abstract: Most colorization models condition only on a single reference, typically the first frame of the scene. However, this approach ignores other sources

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

TinySR: Pruning Diffusion for Real-World Image Super-Resolution

DGX agent

arXiv:2508.17434v2 Announce Type: replace Abstract: Real-world image super-resolution (Real-ISR) focuses on recovering high-quality images from low-resolution inputs that suffer from complex degradati

model-releasesarxiv-cs-cv
21 Apr 2026
Research

ToLL: Topological Layout Learning with Asymmetric Cross-View Structural Distillation for 3D Scene Graph Generation Pretraining

DGX agent

arXiv:2603.28178v2 Announce Type: replace Abstract: 3D Scene Graph (3DSG) generation plays a pivotal role in spatial understanding and affordance perception. To mitigate generalization issues from dat

researcharxiv-cs-cv
21 Apr 2026
Local Ai

Topology-Aware Layer Pruning for Large Vision-Language Models

DGX agent

arXiv:2604.16502v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in natural language understanding and reasoning, while recent extensions that incorpo

local-aiarxiv-cs-cv
21 Apr 2026
Model Releases

Towards Generalizable Deepfake Image Detection with Vision Transformers

DGX agent

arXiv:2604.17376v1 Announce Type: new Abstract: In today's day and age, we face a challenge in detecting deepfake images because of the fast evolution of modern generative models and the poor generali

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Towards Joint Quantization and Token Pruning of Vision-Language Models

DGX agent

arXiv:2604.17320v1 Announce Type: new Abstract: Deploying Vision-Language Models (VLMs) under aggressive low-bit inference remains challenging because inference cost is dominated by the long visual-to

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training

DGX agent

arXiv:2603.23885v3 Announce Type: replace Abstract: Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Tradit

model-releasesarxiv-cs-cv
21 Apr 2026
← Previous
1…231232233234235…261
Next →