Safety
PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation
arXiv:2607.24353v1 Announce Type: new Abstract: Text-to-image generation models can synthesize high-quality images from natural language descriptions, but their performance remains highly sensitive to
arXiv:2607.24353v1 Announce Type: new Abstract: Text-to-image generation models can synthesize high-quality images from natural language descriptions, but their performance remains highly sensitive to prompt formulation. Existing prompt optimization methods mainly rely on text-side rewriting, prompt expansion, or external reward signals, offering limited image-grounded diagnosis and weak support for learning reusable optimisation policies. In this paper, we propose PRISM, a Prompt Refinement framework via Image-grounded Self-rewarding Mechanism. PRISM closes the prompt-image-feedback loop by interpreting generated images with structured visual diagnosis and scoring them along semantic consistency, aesthetic quality, and human preference alignment. It first initializes a unified VLM through multi-task supervised fine-tuning, and then improves the prompt policy via self-rewarding optimization with a hybrid ideal-point and Chebyshev reward. Extensive experiments show that PRISM improves holistic image quality and fine-grained semantic alignment, while providing interpretable feedback for targeted prompt refinement. The code is available at https://anonymous.4open.science/r/PRISM-FF81.
Related
- Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization
- Boosting Text-to-Image Diffusion Models via Core Token Attention-Based Seed Selection
- STEDiff: Strengthening Text Embedding for Text-to-Image Alignment in Diffusion Model
- PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
Source: arXiv cs.CV | 2026-07-28