Model Releases
FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
arXiv:2512.18073v2 Announce Type: replace Abstract: Multimodal LLMs (MLLMs) are capable of performing complex data analysis, visual question answering, generation, and reasoning tasks. However, their
arXiv:2512.18073v2 Announce Type: replace Abstract: Multimodal LLMs (MLLMs) are capable of performing complex data analysis, visual question answering, generation, and reasoning tasks. However, their ability to analyze biometric data is relatively underexplored. In this work, we investigate the effectiveness of MLLMs in understanding fine structural and textural details present in fingerprint images. To this end, we design a comprehensive benchmark, FPBench, to evaluate 20 MLLMs (open-source and proprietary models) across 7 real and synthetic datasets on a suite of 8 biometric and forensic tasks (e.g., pattern analysis, fingerprint verification, real versus synthetic classification, etc.) using zero-shot and chain-of-thought prompting strategies. We further fine-tune vision and language encoders on a subset of open-source MLLMs to demonstrate domain adaptation. FPBench is a novel benchmark designed as a first step towards developing foundation models in fingerprints. Our findings indicate fine-tuning of vision and language encoders improves the performance by 7%-39%. Our codes are available at https://github.com/Ektagavas/FPBench.
Related
- Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
- Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
- MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
- Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios
- MMRareBench: A Rare-Disease Multimodal and Multi-Image Medical Benchmark
Source: arXiv cs.CV | 2026-04-14