ResearchPod Summary
Scene Text Recognition (STR) has historically focused on regular, horizontal text found in natural environments. However, artistic text—often seen in posters, magazines, and advertisements—features highly customized fonts, complex textures, and non-standard layouts. Existing STR models struggle with these variations because they are typically trained on datasets lacking artistic diversity and rely on fixed-template inputs that distort irregular text shapes. This paper addresses these bottlenecks by advancing both the data foundation and the model architecture for WordArt-oriented scene text recognition (WATER).
To overcome the scarcity of high-quality artistic text data, the authors developed WATER-S, a 2M-image synthetic dataset composed of two complementary subsets:
To better handle the visual complexity of artistic text, the authors propose WATERec. Unlike traditional STR models that force images into fixed-size templates, WATERec uses a Vision Transformer encoder that supports arbitrary aspect ratios. By incorporating Rotary Positional Embeddings (RoPE), the model effectively captures spatial relationships regardless of the input shape. The architecture also employs an autoregressive (AR) decoder, which is better suited for the non-linear reading orders and complex layouts frequently encountered in artistic designs.
Experimental results demonstrate that WATERec significantly outperforms existing STR methods and vision-language models on the WordArt-Bench, achieving 90.40% accuracy. The study highlights that combining both tool-based and generative synthetic data is essential for robust generalization. By providing a scalable data foundation and a flexible model design, this work establishes a new state-of-the-art baseline for artistic text recognition, proving that specialized architectural choices are necessary for handling stylized text.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.