Model Releases
TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2
arXiv:2606.19259v2 Announce Type: replace-cross Abstract: Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation model
arXiv:2606.19259v2 Announce Type: replace-cross Abstract: Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation models become increasingly capable of synthesizing realistic textual content and structured visual designs, detecting AI-generated text-rich images has become an important challenge for digital trust and content authenticity. Existing benchmarks, however, largely focus on object-centric images and provide limited coverage of scenarios where textual semantics and layout organization are central. In this paper, we introduce TextRich, a multi-domain benchmark for detecting text-rich images generated by OpenAI's GPT-Image-2. The benchmark contains 12,095 images across six representative categories: commercial posters, infographic charts, academic posters, receipts, tables, and UI screenshots. Using this benchmark, we evaluate five representative AI-generated image detectors under a zero-shot setting and further explore the capability of a multimodal vision-language model for this task. Our results reveal substantial performance variations across text-rich domains, where existing AI-generated image detectors exhibit distinct strengths and failure modes. Although the strongest detector achieves competitive overall performance, it remains ineffective on certain structured categories and highly sensitive to JPEG compression. Vision-language models provide a promising complementary approach, but still struggle with highly structured text-rich images. These findings highlight the need for text- and layout-aware detection methods for modern AI-generated images. Our dataset is released at https://huggingface.co/datasets/Shuyiww/TextRich.
Related
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- COCO-Inpaint: A Benchmark for Detecting and Localizing Inpainting-Based Image Manipulations
- ELIQ: A Label-Free Framework for Quality Assessment of Evolving AI-Generated Images
- FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence
Source: arXiv cs.AI | 2026-07-28