AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
14 Apr 2026

MLLM-as-a-Judge Exhibits Model Preference Bias

Model ReleasesDGX agent

arXiv:2604.11589v1 Announce Type: new Abstract: Automatic evaluation using multimodal large language models (MLLMs), commonly referred to as MLLM-as-a-Judge, has been widely used to measure model perf

MMRareBench: A Rare-Disease Multimodal and Multi-Image Medical Benchmark

Model ReleasesDGX agent

arXiv:2604.10755v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced clinical tasks for common conditions, but their performance on rare diseases remains largely unte

MorphoFlow: Sparse-Supervised Generative Shape Modeling with Adaptive Latent Relevance

TutorialsDGX agent

arXiv:2604.11636v1 Announce Type: new Abstract: Statistical shape modeling (SSM) is central to population level analysis of anatomical variability, yet most existing approaches rely on densely annotat


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MosaicMRI: A Diverse Dataset and Benchmark for Raw Musculoskeletal MRI

Model ReleasesDGX agent

arXiv:2604.11762v1 Announce Type: new Abstract: Deep learning underpins a wide range of applications in MRI, including reconstruction, artifact removal, and segmentation. However, progress has been dr

Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model

ResearchDGX agent

arXiv:2505.23606v4 Announce Type: replace-cross Abstract: Unified generation models aim to handle diverse tasks across modalities -- such as text generation, image generation, and vision-language reas

Multi-Granularity Reasoning for Image Quality Assessment via Attribute-Aware Reinforcement Learning to Rank

SafetyDGX agent

arXiv:2604.09704v1 Announce Type: new Abstract: Recent advances in reasoning-induced image quality assessment (IQA) have demonstrated the power of reinforcement learning to rank (RL2R) for training vi

Multi-Head Attention based interaction-aware architecture for Bangla Handwritten Character Recognition: Introducing a Primary Dataset

Model ReleasesDGX agent

arXiv:2604.09717v1 Announce Type: new Abstract: Character recognition is the fundamental part of an optical character recognition (OCR) system. Word recognition, sentence transcription, document digit

Multi-modal, multi-scale representation learning for satellite imagery analysis just needs a good ALiBi

Model ReleasesDGX agent

arXiv:2604.10347v1 Announce Type: new Abstract: Vision foundation models have been shown to be effective at processing satellite imagery into representations fit for downstream tasks, however, creatin

MuPPet: Multi-person 2D-to-3D Pose Lifting

Local AiDGX agent

arXiv:2604.09715v1 Announce Type: new Abstract: Multi-person social interactions are inherently built on coherence and relationships among all individuals within the group, making multi-person localiz

MVAD: A Benchmark Dataset for Multimodal AI-Generated Video-Audio Detection

Model ReleasesDGX agent

arXiv:2512.00336v2 Announce Type: replace Abstract: The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content auth

Naka-GS: A Bionics-inspired Dual-Branch Naka Correction and Progressive Point Pruning for Low-Light 3DGS

SafetyDGX agent

arXiv:2604.11142v1 Announce Type: new Abstract: Low-light conditions severely hinder 3D restoration and reconstruction by degrading image visibility, introducing color distortions, and contaminating g

Near OOD Detection for Vision-Language Prompt Learning with Contrastive Logit Score

ResearchDGX agent

arXiv:2405.16091v2 Announce Type: replace Abstract: Prompt learning has emerged as an efficient and effective method for fine-tuning vision-language models such as CLIP. While many studies have explor

Neural Stochastic Processes for Satellite Precipitation Refinement

Model ReleasesDGX agent

arXiv:2604.10414v1 Announce Type: new Abstract: Accurate precipitation estimation is critical for flood forecasting, water resource management, and disaster preparedness. Satellite products provide gl

Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry

TutorialsDGX agent

arXiv:2406.04301v3 Announce Type: replace Abstract: Reconstructing accurate surfaces from sparse multi-view images remains challenging due to severe geometric ambiguity and occlusions. Existing genera

NeuVolEx: Implicit Neural Features for Volume Exploration

ResearchDGX agent

arXiv:2604.11172v1 Announce Type: cross Abstract: Direct volume rendering (DVR) aims to help users identify and examine regions of interest (ROIs) within volumetric data, and feature representations t

Ninja Codes: Neurally Generated Fiducial Markers for Stealthy 6-DoF Tracking

ApplicationsDGX agent

arXiv:2510.18976v2 Announce Type: replace Abstract: In this paper we describe Ninja Codes, neurally generated fiducial markers that can be made to naturally blend into various real-world environments.

NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild

ApplicationsDGX agent

arXiv:2604.11487v1 Announce Type: new Abstract: This paper presents an overview of the NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild, held in conjunction with the NTIRE works

NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models: Datasets, Methods and Results

Model ReleasesDGX agent

arXiv:2604.10551v1 Announce Type: new Abstract: This paper presents an overview of the NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models. This challenge utili

NTIRE 2026 Challenge on Single Image Reflection Removal in the Wild: Datasets, Results, and Methods

ApplicationsDGX agent

arXiv:2604.10321v1 Announce Type: new Abstract: In this paper, we review the NTIRE 2026 challenge on single-image reflection removal (SIRR) in the Wild. SIRR is a fundamental task in image restoration

NTIRE 2026 The 3rd Restore Any Image Model (RAIM) Challenge: AI Flash Portrait (Track 3)

Model ReleasesDGX agent

arXiv:2604.11230v1 Announce Type: new Abstract: In this paper, we present a comprehensive overview of the NTIRE 2026 3rd Restore Any Image Model (RAIM) challenge, with a specific focus on Track 3: AI

NTIRE 2026 The Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results

Model ReleasesDGX agent

arXiv:2604.10634v1 Announce Type: new Abstract: This paper presents an overview of the NTIRE 2026 Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images. Building upon the success

Observe Less, Understand More: Cost-aware Cross-scale Observation for Remote Sensing Understanding

Model ReleasesDGX agent

arXiv:2604.11415v1 Announce Type: new Abstract: Remote sensing understanding inherently requires multi-resolution observation, since different targets and application tasks demand different levels of

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video

Model ReleasesDGX agent

arXiv:2604.11102v1 Announce Type: new Abstract: Current multimodal large language models (MLLMs) have demonstrated remarkable capabilities in short-form video understanding, yet translating long-form

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

Model ReleasesDGX agent

arXiv:2604.11804v1 Announce Type: new Abstract: In this work, we study Human-Object Interaction Video Generation (HOIVG), which aims to synthesize high-quality human-object interaction videos conditio

On The Application of Linear Attention in Multimodal Transformers

ResearchDGX agent

arXiv:2604.10064v1 Announce Type: new Abstract: Multimodal Transformers serve as the backbone for state-of-the-art vision-language models, yet their quadratic attention complexity remains a critical b

On the Effectiveness of Textual Prompting with Lightweight Fine-Tuning for SAM3 Remote Sensing Segmentation

SafetyDGX agent

arXiv:2512.15564v2 Announce Type: replace Abstract: Remote sensing (RS) image segmentation is constrained by the limited availability of annotated data and a gap between overhead imagery and natural i

Online Reasoning Video Object Segmentation

Model ReleasesDGX agent

arXiv:2604.11411v1 Announce Type: new Abstract: Reasoning video object segmentation predicts pixel-level masks in videos from natural-language queries that may involve implicit and temporally grounded

Optimization-Guided Diffusion for Interactive Scene Generation

SafetyDGX agent

arXiv:2512.07661v3 Announce Type: replace Abstract: Realistic and diverse multi-agent driving scenes are crucial for evaluating autonomous vehicles, but safety-critical events which are essential for

PA-SFM: Tracker-free differentiable acoustic radiation for freehand 3D photoacoustic imaging

HardwareDGX agent

arXiv:2604.09643v1 Announce Type: new Abstract: Three-dimensional (3D) handheld photoacoustic tomography typically relies on bulky and expensive external positioning sensors to correct motion artifact

PACO: Proxy-Task Alignment and Online Calibration for On-the-Fly Category Discovery

SafetyDGX agent

arXiv:2604.11484v1 Announce Type: new Abstract: On-the-Fly Category Discovery (OCD) requires a model, trained on an offline support set, to recognize known classes while discovering new ones from an o

Pair2Scene: Learning Local Object Relations for Procedural Scene Generation

Local AiDGX agent

arXiv:2604.11808v1 Announce Type: new Abstract: Generating high-fidelity 3D indoor scenes remains a significant challenge due to data scarcity and the complexity of modeling intricate spatial relation

PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion

ResearchDGX agent

arXiv:2601.07447v2 Announce Type: replace Abstract: Existing image foundation models are not optimized for spherical images having been trained primarily on perspective images. PanoSAMic integrates th

Parameter Efficient Fine-tuning for Domain-specific Gastrointestinal Disease Recognition

Model ReleasesDGX agent

arXiv:2604.10451v1 Announce Type: new Abstract: Despite recent advancements in the field of medical image analysis with the use of pretrained foundation models, the issue of distribution shifts betwee

Particle Diffusion Matching: Random Walk Correspondence Search for the Alignment of Standard and Ultra-Widefield Fundus Images

SafetyDGX agent

arXiv:2604.10085v1 Announce Type: new Abstract: We propose a robust alignment technique for Standard Fundus Images (SFIs) and Ultra-Widefield Fundus Images (UWFIs), which are challenging to align due

PASTA: Vision Transformer Patch Aggregation for Weakly Supervised Target and Anomaly Segmentation

TutorialsDGX agent

arXiv:2604.09701v1 Announce Type: new Abstract: Detecting unseen anomalies in unstructured environments presents a critical challenge for industrial and agricultural applications such as material recy

PERCEPT-Net: A Perceptual Loss Driven Framework for Reducing MRI Artifact Tissue Confusion

ResearchDGX agent

arXiv:2604.10439v1 Announce Type: new Abstract: Purpose: Existing deep learning-based MRI artifact correction models exhibit poor clinical generalization due to inherent artifact-tissue confusion, fai

Perceptual Inductive Bias Is What You Need Before Contrastive Learning

SafetyDGX agent

arXiv:2506.01201v2 Announce Type: replace Abstract: David Marr's seminal theory of human perception stipulates that visual processing is a multi-stage process, prioritizing the derivation of boundary

PhyMix: Towards Physically Consistent Single-Image 3D Indoor Scene Generation with Implicit--Explicit Optimization

Model ReleasesDGX agent

arXiv:2604.10125v1 Announce Type: new Abstract: Existing single-image 3D indoor scene generators often produce results that look visually plausible but fail to obey real-world physics, limiting their

Physics-Informed Synthetic Dataset and Denoising TIE-Reconstructed Phase Maps in Transient Flows Using Deep Learning

ResearchDGX agent

arXiv:2604.10610v1 Announce Type: cross Abstract: High-speed quantitative phase imaging enables non-intrusive visualization of transient compressible gas flows and energetic phenomena. However, phase

Point2Pose: Occlusion-Recovering 6D Pose Tracking and 3D Reconstruction for Multiple Unknown Objects Via 2D Point Trackers

Model ReleasesDGX agent

arXiv:2604.10415v1 Announce Type: new Abstract: We present Point2Pose, a model-free method for causal 6D pose tracking of multiple rigid objects from monocular RGB-D video. Initialized only from spars

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs

ApplicationsDGX agent

arXiv:2604.11627v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the ra

PointSplat: Efficient Geometry-Driven Pruning and Transformer Refinement for 3D Gaussian Splatting

ResearchDGX agent

arXiv:2604.09903v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently unlocked real-time, high-fidelity novel view synthesis by representing scenes using explicit 3D primitives. Ho

Precision Synthesis of Multi-Tracer PET via VLM-Modulated Rectified Flow for Stratifying Mild Cognitive Impairment

Model ReleasesDGX agent

arXiv:2604.11176v1 Announce Type: new Abstract: The biological definition of Alzheimer's disease (AD) relies on multi-modal neuroimaging, yet the clinical utility of positron emission tomography (PET)

PrefPaint: Enhancing Medical Image Inpainting through Expert Human Feedback

ResearchDGX agent

arXiv:2506.21834v2 Announce Type: replace Abstract: Inpainting, the process of filling missing or corrupted image parts, has broad applications in medical imaging. However, generating anatomically acc

Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation

SafetyDGX agent

arXiv:2603.20725v2 Announce Type: replace Abstract: Text-to-image generation has advanced rapidly, yet it still struggles to capture the nuanced user preferences. Existing approaches typically rely on

Preventing Latent Rehearsal Decay in Online Continual SSL with SOLAR

TutorialsDGX agent

arXiv:2604.10586v1 Announce Type: cross Abstract: This paper explores Online Continual Self-Supervised Learning (OCSSL), a scenario in which models learn from continuous streams of unlabeled, non-stat

Prints in the Magnetic Dust: Robust Similarity Search in Legacy Media Images Using Checksum Count Vectors

ResearchDGX agent

arXiv:2604.09657v1 Announce Type: new Abstract: Digitizing magnetic media containing computer data is only the first step towards the preservation of early home computing era artifacts. The audio tape

ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction

ApplicationsDGX agent

arXiv:2604.02003v2 Announce Type: replace Abstract: Generating ground-level views and coherent 3D site models from aerial-only imagery is challenging due to extreme viewpoint changes, missing intermed

Progressive Deep Learning for Automated Spheno-Occipital Synchondrosis Maturation Assessment

ResearchDGX agent

arXiv:2604.10945v1 Announce Type: new Abstract: Accurate assessment of spheno-occipital synchondrosis (SOS) maturation is a key indicator of craniofacial growth and a critical determinant for orthodon

Progressively Texture-Aware Diffusion for Contrast-Enhanced Sparse-View CT

ResearchDGX agent

arXiv:2604.11559v1 Announce Type: new Abstract: Diffusion-based sparse-view CT (SVCT) imaging has achieved remarkable advancements in recent years, thanks to its more stable generative capability. How

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation

SafetyDGX agent

arXiv:2604.10030v1 Announce Type: new Abstract: Video diffusion models have achieved remarkable progress in generating high-quality videos. However, these models struggle to represent the temporal suc

ProPhy: Progressive Physical Alignment for Dynamic World Simulation

Local AiDGX agent

arXiv:2512.05564v2 Announce Type: replace Abstract: Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to pro

PSF-Med: Measuring and Explaining Paraphrase Sensitivity in Medical Vision Language Models

Model ReleasesDGX agent

arXiv:2602.21428v2 Announce Type: replace Abstract: Medical Vision Language Models (VLMs) can change their answers when clinicians rephrase the same question, a failure mode that threatens deployment

Quantization Robustness to Input Degradations for Object Detection

ApplicationsDGX agent

arXiv:2508.19600v2 Announce Type: replace Abstract: Post-training quantization (PTQ) is crucial for deploying efficient object detection models, like YOLO, on resource-constrained devices. However, th

Quantum-Gated Task-interaction Knowledge Distillation for Pre-trained Model-based Class-Incremental Learning

TutorialsDGX agent

arXiv:2604.11112v1 Announce Type: cross Abstract: Class-incremental learning (CIL) aims to continuously accumulate knowledge from a stream of tasks and construct a unified classifier over all seen cla

R2E-VID: Two-Stage Robust Routing via Temporal Gating for Elastic Edge-Cloud Video Inference

ResearchDGX agent

arXiv:2604.09681v1 Announce Type: cross Abstract: With the rapid growth of large-scale video analytics applications, edge-cloud collaborative systems have become the dominant paradigm for real-time in

RADA: Region-Aware Dual-encoder Auxiliary learning for Barely-supervised Medical Image Segmentation

ResearchDGX agent

arXiv:2604.11164v1 Announce Type: new Abstract: Deep learning has greatly advanced medical image segmentation, but its success relies heavily on fully supervised learning, which requires dense annotat

Radiology Report Generation for Low-Quality X-Ray Images

Model ReleasesDGX agent

arXiv:2604.10188v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have significantly advanced automated Radiology Report Generation (RRG). However, existing methods implicitly assume high-

Real-Time Human Reconstruction and Animation using Feed-Forward Gaussian Splatting

ResearchDGX agent

arXiv:2604.10259v1 Announce Type: new Abstract: We present a generalizable feed-forward Gaussian splatting framework for human 3D reconstruction and real-time animation that operates directly on multi

ReaLiTy and LADS: A Unified Framework and Dataset Suite for LiDAR Adaptation Across Sensors and Adverse Weather Conditions

ResearchDGX agent

arXiv:2604.10213v1 Announce Type: cross Abstract: Reliable LiDAR perception requires robustness across sensors, environments, and adverse weather. However, existing datasets rarely provide physically

← Previous
1…197198199200201…207
Next →