ResearchPod Summary
Text-to-image (T2I) generation models are highly sensitive to prompt formulation, yet existing optimization methods often rely on text-only rewriting or scalar rewards that lack deep visual understanding. This paper asks: can we create a closed-loop system where a model learns to diagnose its own generated images and use that visual feedback to iteratively refine prompts for better semantic alignment and aesthetic quality?
The authors propose PRISM, a framework that integrates prompt rewriting and visual feedback assessment into a single Vision-Language Model (VLM). The process occurs in two stages:
Experimental results on benchmarks like BeautifulPrompt and T2I-CompBench demonstrate that PRISM consistently outperforms existing text-only and visual-feedback-based methods. By incorporating structured visual diagnosis, the model produces prompts that lead to higher CLIP scores, better aesthetic ratings, and improved human preference alignment. The study highlights that learning a reusable optimization policy via self-rewarding is more effective than relying on test-time-only feedback or fragmented evaluation metrics.
PRISM bridges the gap between text-based prompt engineering and visual evaluation. By enabling models to "see" and critique their own output, it provides a scalable way to improve image generation quality without needing constant human intervention or external reward models. This approach offers a path toward more controllable and faithful T2I generation systems that can adapt to specific user preferences.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.