AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Local Ai

V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators

DGX agent

arXiv:2604.03307v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success, yet they remain prone to perception-related hallucinations in fine-graine

local-aiarxiv-cs-cv
17 Apr 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Vision-Based Safe Human-Robot Collaboration with Uncertainty Guarantees

DGX agent

arXiv:2604.15221v1 Announce Type: cross Abstract: We propose a framework for vision-based human pose estimation and motion prediction that gives conformal prediction guarantees for certifiably safe hu

safetyarxiv-cs-cv
17 Apr 2026
Research

VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models

DGX agent

arXiv:2604.15188v1 Announce Type: new Abstract: Visual token pruning methods effectively mitigate the quadratic computational growth caused by processing high-resolution images and video frames in vis

researcharxiv-cs-cv
17 Apr 2026
Research

WaveSFNet: A Wavelet-Based Codec and Spatial--Frequency Dual-Domain Gating Network for Spatiotemporal Prediction

DGX agent

arXiv:2603.23284v2 Announce Type: replace Abstract: Spatiotemporal predictive learning aims to forecast future frames from historical observations in an unsupervised manner, and is critical to a wide

researcharxiv-cs-cv
17 Apr 2026
Safety

When Fairness Metrics Disagree: Evaluating the Reliability of Demographic Fairness Assessment in Machine Learning

DGX agent

arXiv:2604.15038v1 Announce Type: cross Abstract: The evaluation of fairness in machine learning systems has become a central concern in high-stakes applications, including biometric recognition, heal

safetyarxiv-cs-cv
17 Apr 2026
Safety

Why Do Vision Language Models Struggle To Recognize Human Emotions?

DGX agent

arXiv:2604.15280v1 Announce Type: new Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans. Vision-language models (VLMs) have made trem

safetyarxiv-cs-cv
17 Apr 2026
Model Releases

WILD-SAM: Phase-Aware Expert Adaptation of SAM for Landslide Detection in Wrapped InSAR Interferograms

DGX agent

arXiv:2604.14540v1 Announce Type: new Abstract: Detecting slow-moving landslides directly from wrapped Interferometric Synthetic Aperture Radar (InSAR) interferograms is crucial for efficient geohazar

model-releasesarxiv-cs-cv
17 Apr 2026
Research

Zero-Ablation Overstates Register Content Dependence in DINO Vision Transformers

DGX agent

arXiv:2604.14433v1 Announce Type: new Abstract: Zero-ablation -- replacing token activations with zero vectors -- is widely used to probe token function in vision transformers. Register zeroing in DIN

researcharxiv-cs-cv
17 Apr 2026
Model Releases

Zero-Shot Retail Theft Detection via Orchestrated Vision Models: A Model-Agnostic, Cost-Effective Alternative to Trained Single-Model Systems

DGX agent

arXiv:2604.14846v1 Announce Type: new Abstract: Retail theft costs the global economy over 100 billion annually, yet existing AI-based detection systems require expensive custom model training on prop

model-releasesarxiv-cs-cv
17 Apr 2026
Research

3DRealHead: Few-Shot Detailed Head Avatar

DGX agent

arXiv:2604.13171v1 Announce Type: new Abstract: The human face is central to communication. For immersive applications, the digital presence of a person should mirror the physical reality, capturing t

researcharxiv-cs-cv
16 Apr 2026
Model Releases

4th Workshop on Maritime Computer Vision (MaCVi): Challenge Overview

DGX agent

arXiv:2604.13244v1 Announce Type: new Abstract: The 4th Workshop on Maritime Computer Vision (MaCVi) is organized as part of CVPR 2026. This edition features five benchmark challenges with emphasis on

model-releasesarxiv-cs-cv
16 Apr 2026
Research

A 3D SAM-Based Progressive Prompting Framework for Multi-Task Segmentation of Radiotherapy-induced Normal Tissue Injuries in Limited-Data Settings

DGX agent

arXiv:2604.13367v1 Announce Type: new Abstract: Radiotherapy-induced normal tissue injury is a clinically important complication, and accurate segmentation of injury regions from medical images could

researcharxiv-cs-cv
16 Apr 2026
Research

A Function-Centric Perspective on Flat and Sharp Minima

DGX agent

arXiv:2510.12451v2 Announce Type: replace-cross Abstract: Flat minima are strongly associated with improved generalisation in deep neural networks. However, this connection has proven nuanced in recen

researcharxiv-cs-cv
16 Apr 2026
Safety

A High-Resolution Landscape Dataset for Concept-Based XAI With Application to Species Distribution Models

DGX agent

arXiv:2604.13240v1 Announce Type: new Abstract: Mapping the spatial distribution of species is essential for conservation policy and invasive species management. Species distribution models (SDMs) are

safetyarxiv-cs-cv
16 Apr 2026
Model Releases

A Lightweight Multi-Metric No-Reference Image Quality Assessment Framework for UAV Imaging

DGX agent

arXiv:2604.13112v1 Announce Type: new Abstract: Reliable image quality assessment is essential in applications where large volumes of images are acquired automatically and must be filtered before furt

model-releasesarxiv-cs-cv
16 Apr 2026
Research

A Multi-Stage Optimization Pipeline for Bethesda Cell Detection in Pap Smear Cytology

DGX agent

arXiv:2604.13939v1 Announce Type: new Abstract: Computer vision techniques have advanced significantly in recent years, finding diverse and impactful applications within the medical field. In this pap

researcharxiv-cs-cv
16 Apr 2026
Research

A Multimodal Clinically Informed Coarse-to-Fine Framework for Longitudinal CT Registration in Proton Therapy

DGX agent

arXiv:2604.13397v1 Announce Type: new Abstract: Proton therapy offers superior organ-at-risk sparing but is highly sensitive to anatomical changes, making accurate deformable image registration (DIR)

researcharxiv-cs-cv
16 Apr 2026
Research

A Resource-Efficient Hybrid CNN-LSTM network for image-based bean leaf disease classification

DGX agent

arXiv:2604.13835v1 Announce Type: new Abstract: Accurate and resource-efficient automated diagnosis is a cornerstone of modern agricultural expert systems. While Convolutional Neural Networks (CNNs) h

researcharxiv-cs-cv
16 Apr 2026
Model Releases

A Study of Failure Modes in Two-Stage Human-Object Interaction Detection

DGX agent

arXiv:2604.13448v1 Announce Type: new Abstract: Human-object interaction (HOI) detection aims to detect interactions between humans and objects in images. While recent advances have improved performan

model-releasesarxiv-cs-cv
16 Apr 2026
Research

A Unified Conditional Flow for Motion Generation, Editing, and Intra-Structural Retargeting

DGX agent

arXiv:2604.13427v1 Announce Type: cross Abstract: Text-driven motion editing and intra-structural retargeting, where source and target share topology but may differ in bone lengths, are traditionally

researcharxiv-cs-cv
16 Apr 2026
Applications

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models

DGX agent

arXiv:2511.10946v3 Announce Type: replace Abstract: Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world

applicationsarxiv-cs-cv
16 Apr 2026
Safety

Action Images: End-to-End Policy Learning via Multiview Video Generation

DGX agent

arXiv:2604.06168v2 Announce Type: replace Abstract: World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model t

safetyarxiv-cs-cv
16 Apr 2026
Model Releases

Adaptive Multi-Scale Channel-Spatial Attention Aggregation Framework for 3D Indoor Semantic Scene Completion Toward Assisting Visually Impaired

DGX agent

arXiv:2602.16385v4 Announce Type: replace Abstract: Independent indoor mobility remains a critical challenge for individuals with visual impairments, largely due to the limited capability of existing

model-releasesarxiv-cs-cv
16 Apr 2026
Safety

ADP-DiT: Text-Guided Diffusion Transformer for Brain Image Generation in Alzheimer's Disease Progression

DGX agent

arXiv:2604.13495v1 Announce Type: new Abstract: Alzheimer's disease (AD) progresses heterogeneously across individuals, motivating subject-specific synthesis of follow-up magnetic resonance imaging (M

safetyarxiv-cs-cv
16 Apr 2026
Model Releases

Aerial Vision-Language Navigation with a Unified Framework for Spatial, Temporal and Embodied Reasoning

DGX agent

arXiv:2512.08639v3 Announce Type: replace Abstract: Aerial Vision-and-Language Navigation (VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and navigate c

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

AI Powered Image Analysis for Phishing Detection

DGX agent

arXiv:2604.13555v1 Announce Type: new Abstract: Phishing websites now rely heavily on visual imitation-copied logos, similar layouts, and matching colours-to avoid detection by text- and URL-based sys

model-releasesarxiv-cs-cv
16 Apr 2026
Research

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning

DGX agent

arXiv:2509.25699v2 Announce Type: replace Abstract: Interleaved-Modal Chain-of-Thought (I-MCoT) advances vision-language reasoning, such as Visual Question Answering (VQA). This paradigm integrates sp

researcharxiv-cs-cv
16 Apr 2026
Model Releases

An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning

DGX agent

arXiv:2211.16780v3 Announce Type: replace-cross Abstract: In online incremental learning, data continuously arrives with substantial distributional shifts, creating a significant challenge because pre

model-releasesarxiv-cs-cv
16 Apr 2026
Research

Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait Image

DGX agent

arXiv:2604.13856v1 Announce Type: new Abstract: Reconstructing a complete 3D head from a single portrait remains challenging because existing methods still face a sharp quality-speed trade-off: high-f

researcharxiv-cs-cv
16 Apr 2026
Research

Artificial intelligence application in lymphoma diagnosis with Vision Transformer using weakly supervised training

DGX agent

arXiv:2604.13795v1 Announce Type: new Abstract: Vision transformers (ViT) have been shown to allow for more flexible feature detection and can outperform convolutional neural network (CNN) when pre-tr

researcharxiv-cs-cv
16 Apr 2026
Model Releases

ASTER: Latent Pseudo-Anomaly Generation for Unsupervised Time-Series Anomaly Detection

DGX agent

arXiv:2604.13924v1 Announce Type: cross Abstract: Time-series anomaly detection (TSAD) is critical in domains such as industrial monitoring, healthcare, and cybersecurity, but it remains challenging d

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

ASTRA: Enhancing Multi-Subject Generation with Retrieval-Augmented Pose Guidance and Disentangled Position Embedding

DGX agent

arXiv:2604.13938v1 Announce Type: new Abstract: Subject-driven image generation has shown great success in creating personalized content, but its capabilities are largely confined to single subjects i

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

AudioX: A Unified Framework for Anything-to-Audio Generation

DGX agent

arXiv:2503.10522v4 Announce Type: replace-cross Abstract: Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a

model-releasesarxiv-cs-cv
16 Apr 2026
Hardware

Automatic Charge State Tuning of 300 mm FDSOI Quantum Dots Using Neural Network Segmentation of Charge Stability Diagram

DGX agent

arXiv:2604.13662v1 Announce Type: cross Abstract: Tuning of gate-defined semiconductor quantum dots (QDs) is a major bottleneck for scaling spin qubit technologies. We present a deep learning (DL) dri

hardwarearxiv-cs-cv
16 Apr 2026
Local Ai

Beyond Voxel 3D Editing: Learning from 3D Masks and Self-Constructed Data

DGX agent

arXiv:2604.13688v1 Announce Type: new Abstract: 3D editing refers to the ability to apply local or global modifications to 3D assets. Effective 3D editing requires maintaining semantic consistency by

local-aiarxiv-cs-cv
16 Apr 2026
Safety

Bias at the End of the Score

DGX agent

arXiv:2604.13305v1 Announce Type: new Abstract: Reward models (RMs) are inherently non-neutral value functions designed and trained to encode specific objectives, such as human preferences or text-ima

safetyarxiv-cs-cv
16 Apr 2026
Applications

Blind Bitstream-corrupted Video Recovery via Metadata-guided Diffusion Model

DGX agent

arXiv:2604.13906v1 Announce Type: new Abstract: Bitstream-corrupted video recovery aims to restore realistic content degraded during video storage or transmission. Existing methods typically assume th

applicationsarxiv-cs-cv
16 Apr 2026
Safety

C^2T: Captioning-Structure and LLM-Aligned Common-Sense Reward Learning for Traffic--Vehicle Coordination

DGX agent

arXiv:2604.13098v1 Announce Type: cross Abstract: State-of-the-art (SOTA) urban traffic control increasingly employs Multi-Agent Reinforcement Learning (MARL) to coordinate Traffic Light Controllers (

safetyarxiv-cs-cv
16 Apr 2026
Research

Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision

DGX agent

arXiv:2604.13304v1 Announce Type: new Abstract: Understanding the internal activations of Vision Transformers (ViTs) is critical for building interpretable and trustworthy models. While Sparse Autoenc

researcharxiv-cs-cv
16 Apr 2026
Safety

CausalDisenSeg: A Causality-Guided Disentanglement Framework with Counterfactual Reasoning for Robust Brain Tumor Segmentation Under Missing Modalities

DGX agent

arXiv:2604.13409v1 Announce Type: new Abstract: In clinical practice, the robustness of deep learning models for multimodal brain tumor segmentation is severely compromised by incomplete MRI data. Thi

safetyarxiv-cs-cv
16 Apr 2026
Safety

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling

DGX agent

arXiv:2604.13561v1 Announce Type: new Abstract: Vision-language models trained with contrastive learning on paired medical images and reports show strong zero-shot diagnostic capabilities, yet the eff

safetyarxiv-cs-cv
16 Apr 2026
Research

ClipGStream: Clip-Stream Gaussian Splatting for Any Length and Any Motion Multi-View Dynamic Scene Reconstruction

DGX agent

arXiv:2604.13746v1 Announce Type: new Abstract: Dynamic 3D scene reconstruction is essential for immersive media such as VR, MR, and XR, yet remains challenging for long multi-view sequences with larg

researcharxiv-cs-cv
16 Apr 2026
Safety

Context Sensitivity Improves Human-Machine Visual Alignment

DGX agent

arXiv:2604.13883v1 Announce Type: new Abstract: Modern machine learning models typically represent inputs as fixed points in a high-dimensional embedding space. While this approach has been proven pow

safetyarxiv-cs-cv
16 Apr 2026
Research

Creo: From One-Shot Image Generation to Progressive, Co-Creative Ideation

DGX agent

arXiv:2604.13956v1 Announce Type: cross Abstract: Text-to-image (T2I) systems enable rapid generation of high-fidelity imagery but are misaligned with how visual ideas develop. T2I systems generate ou

researcharxiv-cs-cv
16 Apr 2026
Research

Cyclic 2.5D Perceptual Loss for Cross-Modal 3D Medical Image Synthesis: T1w MRI to Tau PET

DGX agent

arXiv:2406.12632v3 Announce Type: replace-cross Abstract: Positron emission tomography (PET) provides molecular biomarkers for Alzheimer's disease and related dementias (ADRD) and is increasingly used

researcharxiv-cs-cv
16 Apr 2026
Model Releases

Decoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models

DGX agent

arXiv:2604.14044v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) excel in general vision-language tasks, their application to remote sensing change understanding is hinde

model-releasesarxiv-cs-cv
16 Apr 2026
Research

Deep Spatially-Regularized and Superpixel-Based Diffusion Learning for Unsupervised Hyperspectral Image Clustering

DGX agent

arXiv:2604.13307v1 Announce Type: new Abstract: An unsupervised framework for hyperspectral image (HSI) clustering is proposed that incorporates masked deep representation learning with diffusion-base

researcharxiv-cs-cv
16 Apr 2026
Research

Dehaze-then-Splat: Generative Dehazing with Physics-Informed 3D Gaussian Splatting for Smoke-Free Novel View Synthesis

DGX agent

arXiv:2604.13589v1 Announce Type: new Abstract: We present Dehaze-then-Splat, a two-stage pipeline for multi-view smoke removal and novel view synthesis developed for Track~2 of the NTIRE 2026 3D Rest

researcharxiv-cs-cv
16 Apr 2026
← Previous
1…238239240241242…261
Next →