AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
28 Apr 2026

AD-Relight: Training-Free Banner Relighting via Illumination Translation with Diffusion Priors

ResearchDGX agent

arXiv:2604.24407v1 Announce Type: new Abstract: The recent surge in content consumption through streaming services has driven a growing demand for personalized content. Personalized advertisements (ad

Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model

Model ReleasesDGX agent

arXiv:2508.06206v4 Announce Type: replace-cross Abstract: Affordance grounding focuses on predicting the specific regions of objects that are associated with the actions to be performed by robots. It

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method

AgentsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.22836v1 Announce Type: new Abstract: This report describes a Ref-VOS pipeline centered on Sa2VA and organized with explicit agent roles. The key idea is that Sa2VA should provide the first

AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance

SafetyDGX agent

arXiv:2604.23909v1 Announce Type: new Abstract: Navigational aids for blind and low vision individuals struggle conveying dynamic real-world environments, leading to cognitive overload from continuous

An Affordable,Wearable Stereo-Eye-Tracking Platform

ResearchDGX agent

arXiv:2604.24331v1 Announce Type: new Abstract: Research on video-based eye-tracking has long explored stereo and glint-based methods, yet existing wearable eye trackers - both commercial and open-sou

AnemiaVision: Non-Invasive Anemia Detection via Smartphone Imagery Using EfficientNet-B3 with TrivialAugmentWide, Mixup Augmentation, and Persistent Patient History Management

SafetyDGX agent

arXiv:2604.22964v1 Announce Type: new Abstract: Anemia affects over one billion people globally and remains severely under-diagnosed in low-resource regions where laboratory blood tests are inaccessib

Animalbooth: multimodal feature enhancement for animal subject personalization

SafetyDGX agent

arXiv:2509.16702v2 Announce Type: replace Abstract: Personalized animal image generation is challenging due to rich appearance cues and large morphological variability. Existing approaches often exhib

Attention-Augmented YOLOv8 with Ghost Convolution for Real-Time Vehicle Detection in Intelligent Transportation Systems

AgentsDGX agent

arXiv:2604.22856v1 Announce Type: new Abstract: Accurate vehicle detection is a critical component of autonomous driving, traffic surveillance, and intelligent transportation systems. This paper prese

ATTN-FIQA: Interpretable Attention-based Face Image Quality Assessment with Vision Transformers

Model ReleasesDGX agent

arXiv:2604.22841v1 Announce Type: new Abstract: Face Image Quality Assessment (FIQA) aims to assess the recognition utility of face samples and is essential for reliable face recognition (FR) systems.

AusSmoke meets MultiNatSmoke: a fully-labelled diverse smoke segmentation dataset

Model ReleasesDGX agent

arXiv:2604.23542v1 Announce Type: new Abstract: Wildfires are an escalating global concern due to the devastating impacts on the environment, economy, and human health, with notable incidents such as

AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark

Model ReleasesDGX agent

arXiv:2604.24441v1 Announce Type: new Abstract: Autonomous agents capable of navigating Graphical User Interfaces (GUIs) hold the potential to revolutionize digital productivity. However, achieving tr

AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering

TutorialsDGX agent

arXiv:2510.18346v2 Announce Type: replace Abstract: Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse ques

Aycromo: An Open-Source Platform for Automatic Chromosome Detection in Metaphase Images Based on Deep Learning

ResearchDGX agent

arXiv:2604.24685v1 Announce Type: new Abstract: Chromosome analysis is a fundamental step in the diagnosis of genetic diseases, but the manual karyotyping workflow is time-consuming and heavily depend

B-FIRE: Binning-Free Diffusion Implicit Neural Representation for Hyper-Accelerated Motion-Resolved MRI

ResearchDGX agent

arXiv:2601.06166v2 Announce Type: replace Abstract: Accelerated dynamic volumetric magnetic resonance imaging (4DMRI) is essential for applications relying on motion resolution. Existing 4DMRI produce

Benchmarking Pathology Foundation Models for Breast Cancer Survival Prediction

Model ReleasesDGX agent

arXiv:2604.24679v1 Announce Type: new Abstract: Pathology foundation models (PFMs) have recently emerged as powerful pretrained encoders for computational pathology, enabling transfer learning across

BIMStruct3D: A Fully Automated Hybrid Learning Scan-to-BIM Pipeline with Integrated Topology Refinement

ResearchDGX agent

arXiv:2604.24311v1 Announce Type: new Abstract: Automatic generation of Building Information Models (BIM) from building scans is a key challenge in architecture and construction. We present a modular

BIR-Adapter: A parameter-efficient diffusion adapter for blind image restoration

Model ReleasesDGX agent

arXiv:2509.06904v3 Announce Type: replace Abstract: We introduce the BIR-Adapter, a parameter-efficient diffusion adapter for blind image restoration. Diffusion-based restoration methods have demonstr

BMD-45: A Large-Scale CCTV Vehicle Detection Dataset for Urban Traffic in Developing Cities

SafetyDGX agent

arXiv:2604.24419v1 Announce Type: new Abstract: Robust vehicle detection from fixed CCTV cameras is critical for Intelligent Transportation Systems. Yet existing benchmarks predominantly feature relat

Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations

Model ReleasesDGX agent

arXiv:2603.08592v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) have achieved remarkable success in 2D visual understanding, their ability to reason about 3D space r

Breaking Degradation Coupling: A Structural Entropy Guided Decoupled Framework and Benchmark for Infrared Enhancement

Model ReleasesDGX agent

arXiv:2604.22886v1 Announce Type: new Abstract: Thermal infrared image enhancement aims to restore high-quality images from complex compound degradations. Existing all-in-one approaches typically empl

Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training

SafetyDGX agent

arXiv:2604.23121v1 Announce Type: cross Abstract: Have you ever post-trained a generalist vision-language-action (VLA) policy on a small demonstration dataset, only to find that it stops responding to

Breaking the Resource Wall: Geometry-Guided Sequence Modeling for Efficient Semantic Segmentation

Local AiDGX agent

arXiv:2604.23399v1 Announce Type: new Abstract: High-performance semantic segmentation has achieved significant progress in recent years, often driven by increasingly large backbones and higher comput

Breaking the Scalability Limit of Multi-Projector Calibration with Embedded Cameras

ResearchDGX agent

arXiv:2604.24024v1 Announce Type: new Abstract: Conventional multi-projector calibration requires projecting and capturing structured light patterns for each projector sequentially, causing calibratio

BrickNet: Graph-Backed Generative Brick Assembly

ResearchDGX agent

arXiv:2604.22984v1 Announce Type: new Abstract: We train a language model to generate LEGO-brick build sequences. While prior work has been restricted to discrete, voxel-like towers, we consider a muc

Bridging Restoration and Generation Manifolds in One-Step Diffusion for Real-World Super-Resolution

ApplicationsDGX agent

arXiv:2604.24136v1 Announce Type: new Abstract: Pretrained diffusion models have revolutionized real-world image super-resolution (Real-ISR) but suffer from computational bottlenecks due to iterative

Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search

Model ReleasesDGX agent

arXiv:2604.23282v1 Announce Type: new Abstract: Text-based person anomaly search retrieves specific behavioral events from surveillance archives using natural-language queries. Although recent pose-aw

Bringing a Personal Point of View: Evaluating Dynamic 3D Gaussian Splatting for Egocentric Scene Reconstruction

ResearchDGX agent

arXiv:2604.23803v1 Announce Type: new Abstract: Egocentric video provides a unique view into human perception and interaction, with growing relevance for augmented reality, robotics, and assistive tec

BSViT: A Burst Spiking Vision Transformer for Expressive and Efficient Visual Representation Learning

ResearchDGX agent

arXiv:2604.23165v1 Announce Type: new Abstract: Spiking Vision Transformers (S-ViTs) offer a promising framework for energy-efficient visual learning. However, existing designs remain limited by two f

BurstGP: Enhancing Raw Burst Image Super Resolution with Generative Priors

ResearchDGX agent

arXiv:2604.23508v1 Announce Type: new Abstract: Burst image super resolution (BISR) aims to construct a single high-resolution (HR) image by aggregating information from multiple low-resolution (LR) f

BVI-Mamba: Video Enhancement Using a Visual State-Space Model for Low-Light and Underwater Environments

SafetyDGX agent

arXiv:2604.23655v1 Announce Type: new Abstract: Videos captured in low-light and underwater conditions often suffer from distortions such as noise, low contrast, color imbalance, and blur. These issue

CA-IDD: Cross-Attention Guided Identity-Conditional Diffusion for Identity-Consistent Face Swapping

SafetyDGX agent

arXiv:2604.24493v1 Announce Type: new Abstract: Face swapping aims to optimize realistic facial image generation by leveraging the identity of a source face onto a target face while preserving pose, e

Caries DETR: Tooth Structure-aware Prior and Lesion-aware Dynamic Loss Refinement for DETR Based Caries Detection

ResearchDGX agent

arXiv:2604.23718v1 Announce Type: new Abstract: As dental caries appear as subtle, low-contrast lesions in intraoral imaging, existing deep learning models face significant challenges in the early det

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM

ApplicationsDGX agent

arXiv:2603.27507v2 Announce Type: replace Abstract: Recent advancements in multi-modal large language models (MLLMs) have shown strong potential for 3D scene understanding. However, existing methods s

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

Model ReleasesDGX agent

arXiv:2604.23781v1 Announce Type: new Abstract: Language-model agents are increasingly used as persistent coworkers that assist users across multiple working days. During such workflows, the surroundi

CLIP-Guided Data Augmentation for Night-Time Image Dehazing

ResearchDGX agent

arXiv:2604.05500v2 Announce Type: replace Abstract: Nighttime image dehazing faces a more complex degradation pattern than its daytime counterpart, as haze scattering couples with low illumination, no

CLLAP: Contrastive Learning-based LiDAR-Augmented Pretraining for Enhanced Radar-Camera Fusion

AgentsDGX agent

arXiv:2604.24044v1 Announce Type: new Abstract: Accurate 3D object detection is critical for autonomous driving, necessitating reliable, cost-effective sensors capable of operating in adverse weather

Coarse-to-Real: Generative Rendering for Populated Dynamic Scenes

ResearchDGX agent

arXiv:2601.22301v2 Announce Type: replace Abstract: Traditional rendering pipelines rely on complex assets, accurate materials and lighting, and substantial computational resources to produce realisti

Comparative Study of Weighted and Coupled Second- and Fourth-Order PDEs for Image Despeckling in Grayscale, Color, SAR, and Ultrasound

Model ReleasesDGX agent

arXiv:2604.23612v1 Announce Type: new Abstract: Partial Differential Equation (PDE)-based approaches have gained significant attention in image despeckling due to their strong capability to preserve s

Complexity of Linear Regions in Self-supervised Deep ReLU Networks

Model ReleasesDGX agent

arXiv:2604.24393v1 Announce Type: cross Abstract: There has been growing interest in studying the complexity of Rectified Linear Unit (ReLU) based activation networks. Recent work investigates the evo

Computer Vision-Based Early Detection of Container Loss at Sea

SafetyDGX agent

arXiv:2604.24193v1 Announce Type: new Abstract: Containerised shipping underpins global trade, yet container loss at sea remains a persistent safety, environmental, and economic challenge. Despite com

Contrastive Learning for Multimodal Human Activity Recognition with Limited Labeled Data

ApplicationsDGX agent

arXiv:2604.23281v1 Announce Type: cross Abstract: Human activity recognition serves as the foundation for various emerging applications. In recent years, researchers have used collaborative sensing of

Decoupling Wavelet Sub-bands for Single Source Domain Generalization in Fundus Image Segmentation

ResearchDGX agent

arXiv:2603.28463v2 Announce Type: replace Abstract: Domain generalization in fundus imaging is challenging due to variations in acquisition conditions across devices and clinical settings. The inabili

Deploy DINO with Many-to-Many Association

Model ReleasesDGX agent

arXiv:2604.23670v1 Announce Type: new Abstract: Motivated by the limited generalization of supervised image matching models to unseen image domains, we explore the zero-shot deployment of DINO feature

Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkage

SafetyDGX agent

arXiv:2511.22177v2 Announce Type: replace-cross Abstract: Most post-training methods for text-to-image samplers focus on model weights: either fine-tuning the backbone for alignment or distilling it f

Detecting and Evaluating Medical Hallucinations in Large Vision Language Models

Model ReleasesDGX agent

arXiv:2406.10185v2 Announce Type: replace Abstract: Large Vision Language Models (LVLMs) are increasingly integral to healthcare applications, including medical visual question answering and imaging r

DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning

SafetyDGX agent

arXiv:2601.16046v2 Announce Type: replace-cross Abstract: Language-driven dexterous grasp generation requires the models to understand task semantics, 3D geometry, and complex hand-object interactions

DGHMesh: A Large-scale Dual-radar mmWave Dataset and Generalization-focused Benchmark for Human Mesh Reconstruction

Model ReleasesDGX agent

arXiv:2604.22827v1 Announce Type: new Abstract: Millimeter-wave (mmWave) radar has shown great potential for contactless, privacy-preserving, and robust human sensing, yet existing mmWave-based human

DiffuSAM: Diffusion-Based Prompt-Free SAM2 for Few-Shot and Source-Free Medical Image Segmentation

ResearchDGX agent

arXiv:2604.24719v1 Announce Type: new Abstract: Segmentation models such as Segment Anything Model (SAM) and SAM2 achieve strong prompt-driven zero-shot performance. However, their training on natural

Diffusion Model as a Generalist Segmentation Learner

ResearchDGX agent

arXiv:2604.24575v1 Announce Type: new Abstract: Diffusion models are primarily trained for image synthesis, yet their denoising trajectories encode rich, spatially aligned visual priors. In this paper

Discriminator-Guided Adaptive Diffusion for Source-Free Test-Time Adaptation under Image Corruptions

ResearchDGX agent

arXiv:2604.23636v1 Announce Type: new Abstract: In this work, we study Source-Free Unsupervised Domain Adaptation under corruption-induced domain shifts, where performance degradation is caused by nat

Do Protective Perturbations Really Protect Portrait Privacy under Real-world Image Transformations?

ApplicationsDGX agent

arXiv:2604.23688v1 Announce Type: new Abstract: Proactive defense methods protect portrait images from unauthorized editing or talking face generation (TFG) by introducing pixel-level protective pertu

Don't Pause! Every prediction matters in a streaming video

Model ReleasesDGX agent

arXiv:2604.24317v1 Announce Type: new Abstract: Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largel

Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes

ResearchDGX agent

arXiv:2604.22847v1 Announce Type: new Abstract: We introduce Dream-Cubed, a large-scale dataset of Minecraft worlds at voxel resolution, and a family of models using cubes as powerful compositional un

DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning

Model ReleasesDGX agent

arXiv:2510.15050v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have made rapid progress, yet their reasoning ability often lags behind strong text-only LLMs. Bridging thi

DYMAPIA: A Multi-Domain Framework for Detecting AI-based Video Manipulation

Local AiDGX agent

arXiv:2604.24426v1 Announce Type: new Abstract: AI-generated media are advancing rapidly, raising pressing concerns for content authenticity and digital trust. We introduce DYMAPIA, a multi-domain Dee

DynProto: Dynamic Prototype Evolution for Out-of-Distribution Detection

Model ReleasesDGX agent

arXiv:2604.23729v1 Announce Type: new Abstract: Recent studies show that using potential out-of-distribution (OOD) labels from large corpora as auxiliary information can improve OOD detection in visio

EAGLE: Expert-Augmented Attention Guidance for Tuning-Free Industrial Anomaly Detection in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2602.17419v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) can enrich industrial anomaly detection with semantic descriptions and anomaly reasoning, but they still la

Easy Ensemble: Simple Deep Ensemble Learning for Sensor-Based Human Activity Recognition

Model ReleasesDGX agent

arXiv:2203.04153v2 Announce Type: replace Abstract: Sensor-based human activity recognition (HAR) is a paramount technology in the Internet of Things services. HAR using representation learning, which

Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing

ResearchDGX agent

arXiv:2604.23763v1 Announce Type: new Abstract: Large diffusion transformers (DiTs) follow global editing instructions well but consistently leak local edits into unrelated regions, because joint-atte

Efficient Image Annotation via Semi-Supervised Object Segmentation with Label Propagation

ResearchDGX agent

arXiv:2604.22992v1 Announce Type: new Abstract: Reliable object perception is necessary for general-purpose service robots. Open-vocabulary detectors struggle to generalize beyond a few classes and fu

← Previous
1…168169170171172…209
Next →