MLLM-as-a-Judge Exhibits Model Preference Bias
arXiv:2604.11589v1 Announce Type: new Abstract: Automatic evaluation using multimodal large language models (MLLMs), commonly referred to as MLLM-as-a-Judge, has been widely used to measure model perf
Knowledge catalogue
arXiv:2604.11589v1 Announce Type: new Abstract: Automatic evaluation using multimodal large language models (MLLMs), commonly referred to as MLLM-as-a-Judge, has been widely used to measure model perf
arXiv:2604.10755v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced clinical tasks for common conditions, but their performance on rare diseases remains largely unte
arXiv:2604.11636v1 Announce Type: new Abstract: Statistical shape modeling (SSM) is central to population level analysis of anatomical variability, yet most existing approaches rely on densely annotat
arXiv:2604.11762v1 Announce Type: new Abstract: Deep learning underpins a wide range of applications in MRI, including reconstruction, artifact removal, and segmentation. However, progress has been dr
arXiv:2505.23606v4 Announce Type: replace-cross Abstract: Unified generation models aim to handle diverse tasks across modalities -- such as text generation, image generation, and vision-language reas
arXiv:2604.09704v1 Announce Type: new Abstract: Recent advances in reasoning-induced image quality assessment (IQA) have demonstrated the power of reinforcement learning to rank (RL2R) for training vi
arXiv:2604.09717v1 Announce Type: new Abstract: Character recognition is the fundamental part of an optical character recognition (OCR) system. Word recognition, sentence transcription, document digit
arXiv:2604.10347v1 Announce Type: new Abstract: Vision foundation models have been shown to be effective at processing satellite imagery into representations fit for downstream tasks, however, creatin
arXiv:2604.09715v1 Announce Type: new Abstract: Multi-person social interactions are inherently built on coherence and relationships among all individuals within the group, making multi-person localiz
arXiv:2512.00336v2 Announce Type: replace Abstract: The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content auth
arXiv:2604.11142v1 Announce Type: new Abstract: Low-light conditions severely hinder 3D restoration and reconstruction by degrading image visibility, introducing color distortions, and contaminating g
arXiv:2405.16091v2 Announce Type: replace Abstract: Prompt learning has emerged as an efficient and effective method for fine-tuning vision-language models such as CLIP. While many studies have explor
arXiv:2604.10414v1 Announce Type: new Abstract: Accurate precipitation estimation is critical for flood forecasting, water resource management, and disaster preparedness. Satellite products provide gl
arXiv:2406.04301v3 Announce Type: replace Abstract: Reconstructing accurate surfaces from sparse multi-view images remains challenging due to severe geometric ambiguity and occlusions. Existing genera
arXiv:2604.11172v1 Announce Type: cross Abstract: Direct volume rendering (DVR) aims to help users identify and examine regions of interest (ROIs) within volumetric data, and feature representations t
arXiv:2510.18976v2 Announce Type: replace Abstract: In this paper we describe Ninja Codes, neurally generated fiducial markers that can be made to naturally blend into various real-world environments.
arXiv:2604.11487v1 Announce Type: new Abstract: This paper presents an overview of the NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild, held in conjunction with the NTIRE works
arXiv:2604.10551v1 Announce Type: new Abstract: This paper presents an overview of the NTIRE 2026 Challenge on Short-form UGC Video Restoration in the Wild with Generative Models. This challenge utili
arXiv:2604.10321v1 Announce Type: new Abstract: In this paper, we review the NTIRE 2026 challenge on single-image reflection removal (SIRR) in the Wild. SIRR is a fundamental task in image restoration
arXiv:2604.11230v1 Announce Type: new Abstract: In this paper, we present a comprehensive overview of the NTIRE 2026 3rd Restore Any Image Model (RAIM) challenge, with a specific focus on Track 3: AI
arXiv:2604.10634v1 Announce Type: new Abstract: This paper presents an overview of the NTIRE 2026 Second Challenge on Day and Night Raindrop Removal for Dual-Focused Images. Building upon the success
arXiv:2604.11415v1 Announce Type: new Abstract: Remote sensing understanding inherently requires multi-resolution observation, since different targets and application tasks demand different levels of
arXiv:2604.11102v1 Announce Type: new Abstract: Current multimodal large language models (MLLMs) have demonstrated remarkable capabilities in short-form video understanding, yet translating long-form
arXiv:2604.11804v1 Announce Type: new Abstract: In this work, we study Human-Object Interaction Video Generation (HOIVG), which aims to synthesize high-quality human-object interaction videos conditio
arXiv:2604.10064v1 Announce Type: new Abstract: Multimodal Transformers serve as the backbone for state-of-the-art vision-language models, yet their quadratic attention complexity remains a critical b
arXiv:2512.15564v2 Announce Type: replace Abstract: Remote sensing (RS) image segmentation is constrained by the limited availability of annotated data and a gap between overhead imagery and natural i
arXiv:2604.11411v1 Announce Type: new Abstract: Reasoning video object segmentation predicts pixel-level masks in videos from natural-language queries that may involve implicit and temporally grounded
arXiv:2512.07661v3 Announce Type: replace Abstract: Realistic and diverse multi-agent driving scenes are crucial for evaluating autonomous vehicles, but safety-critical events which are essential for
arXiv:2604.09643v1 Announce Type: new Abstract: Three-dimensional (3D) handheld photoacoustic tomography typically relies on bulky and expensive external positioning sensors to correct motion artifact
arXiv:2604.11484v1 Announce Type: new Abstract: On-the-Fly Category Discovery (OCD) requires a model, trained on an offline support set, to recognize known classes while discovering new ones from an o
arXiv:2604.11808v1 Announce Type: new Abstract: Generating high-fidelity 3D indoor scenes remains a significant challenge due to data scarcity and the complexity of modeling intricate spatial relation
arXiv:2601.07447v2 Announce Type: replace Abstract: Existing image foundation models are not optimized for spherical images having been trained primarily on perspective images. PanoSAMic integrates th
arXiv:2604.10451v1 Announce Type: new Abstract: Despite recent advancements in the field of medical image analysis with the use of pretrained foundation models, the issue of distribution shifts betwee
arXiv:2604.10085v1 Announce Type: new Abstract: We propose a robust alignment technique for Standard Fundus Images (SFIs) and Ultra-Widefield Fundus Images (UWFIs), which are challenging to align due
arXiv:2604.09701v1 Announce Type: new Abstract: Detecting unseen anomalies in unstructured environments presents a critical challenge for industrial and agricultural applications such as material recy
arXiv:2604.10439v1 Announce Type: new Abstract: Purpose: Existing deep learning-based MRI artifact correction models exhibit poor clinical generalization due to inherent artifact-tissue confusion, fai
arXiv:2506.01201v2 Announce Type: replace Abstract: David Marr's seminal theory of human perception stipulates that visual processing is a multi-stage process, prioritizing the derivation of boundary
arXiv:2604.10125v1 Announce Type: new Abstract: Existing single-image 3D indoor scene generators often produce results that look visually plausible but fail to obey real-world physics, limiting their
arXiv:2604.10610v1 Announce Type: cross Abstract: High-speed quantitative phase imaging enables non-intrusive visualization of transient compressible gas flows and energetic phenomena. However, phase
arXiv:2604.10415v1 Announce Type: new Abstract: We present Point2Pose, a model-free method for causal 6D pose tracking of multiple rigid objects from monocular RGB-D video. Initialized only from spars
arXiv:2604.11627v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the ra
arXiv:2604.09903v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently unlocked real-time, high-fidelity novel view synthesis by representing scenes using explicit 3D primitives. Ho
arXiv:2604.11176v1 Announce Type: new Abstract: The biological definition of Alzheimer's disease (AD) relies on multi-modal neuroimaging, yet the clinical utility of positron emission tomography (PET)
arXiv:2506.21834v2 Announce Type: replace Abstract: Inpainting, the process of filling missing or corrupted image parts, has broad applications in medical imaging. However, generating anatomically acc
arXiv:2603.20725v2 Announce Type: replace Abstract: Text-to-image generation has advanced rapidly, yet it still struggles to capture the nuanced user preferences. Existing approaches typically rely on
arXiv:2604.10586v1 Announce Type: cross Abstract: This paper explores Online Continual Self-Supervised Learning (OCSSL), a scenario in which models learn from continuous streams of unlabeled, non-stat
arXiv:2604.09657v1 Announce Type: new Abstract: Digitizing magnetic media containing computer data is only the first step towards the preservation of early home computing era artifacts. The audio tape
arXiv:2604.02003v2 Announce Type: replace Abstract: Generating ground-level views and coherent 3D site models from aerial-only imagery is challenging due to extreme viewpoint changes, missing intermed
arXiv:2604.10945v1 Announce Type: new Abstract: Accurate assessment of spheno-occipital synchondrosis (SOS) maturation is a key indicator of craniofacial growth and a critical determinant for orthodon
arXiv:2604.11559v1 Announce Type: new Abstract: Diffusion-based sparse-view CT (SVCT) imaging has achieved remarkable advancements in recent years, thanks to its more stable generative capability. How
arXiv:2604.10030v1 Announce Type: new Abstract: Video diffusion models have achieved remarkable progress in generating high-quality videos. However, these models struggle to represent the temporal suc
arXiv:2512.05564v2 Announce Type: replace Abstract: Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to pro
arXiv:2602.21428v2 Announce Type: replace Abstract: Medical Vision Language Models (VLMs) can change their answers when clinicians rephrase the same question, a failure mode that threatens deployment
arXiv:2508.19600v2 Announce Type: replace Abstract: Post-training quantization (PTQ) is crucial for deploying efficient object detection models, like YOLO, on resource-constrained devices. However, th
arXiv:2604.11112v1 Announce Type: cross Abstract: Class-incremental learning (CIL) aims to continuously accumulate knowledge from a stream of tasks and construct a unified classifier over all seen cla
arXiv:2604.09681v1 Announce Type: cross Abstract: With the rapid growth of large-scale video analytics applications, edge-cloud collaborative systems have become the dominant paradigm for real-time in
arXiv:2604.11164v1 Announce Type: new Abstract: Deep learning has greatly advanced medical image segmentation, but its success relies heavily on fully supervised learning, which requires dense annotat
arXiv:2604.10188v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have significantly advanced automated Radiology Report Generation (RRG). However, existing methods implicitly assume high-
arXiv:2604.10259v1 Announce Type: new Abstract: We present a generalizable feed-forward Gaussian splatting framework for human 3D reconstruction and real-time animation that operates directly on multi
arXiv:2604.10213v1 Announce Type: cross Abstract: Reliable LiDAR perception requires robustness across sensors, environments, and adverse weather. However, existing datasets rarely provide physically