Model Releases
ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence
arXiv:2604.20719v1 Announce Type: cross Abstract: Omnimodal Notation Processing (ONP) represents a unique frontier for omnimodal AI due to the rigorous, multi-dimensional alignment required across aud
arXiv:2604.20719v1 Announce Type: cross Abstract: Omnimodal Notation Processing (ONP) represents a unique frontier for omnimodal AI due to the rigorous, multi-dimensional alignment required across auditory, visual, and symbolic domains. Current research remains fragmented, focusing on isolated transcription tasks that fail to bridge the gap between superficial pattern recognition and the underlying musical logic. This landscape is further complicated by severe notation biases toward Western staff and the inherent unreliability of "LLM-as-a-judge" metrics, which often mask structural reasoning failures with systemic hallucinations. To establish a more rigorous standard, we introduce ONOTE, a multi-format benchmark that utilizes a deterministic pipeline--grounded in canonical pitch projection--to eliminate subjective scoring biases across diverse notation systems. Our evaluation of leading omnimodal models exposes a fundamental disconnect between perceptual accuracy and music-theoretic comprehension, providing a necessary framework for diagnosing reasoning vulnerabilities in complex, rule-constrained domains.
Related
- OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
- VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories
- PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers
- General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks
Source: arXiv cs.AI | 2026-04-23