ResearchPod Summary
Vision-Text Compression (VTC) is a technique that renders long text into images to bypass the high computational costs of processing long text sequences in Multimodal Large Language Models (MLLMs). However, these models often perform significantly worse when reading rendered text compared to native text. The authors investigate whether this performance gap is due to information loss during rendering or a fundamental representational mismatch, and they seek a way to align these two input paths without requiring external human annotations.
The authors identify a phenomenon they call "cross-path inconsistency," where the vision encoder (ViT) treats rendered text as visual patterns (glyphs, layout) rather than linguistic content. To address this, they introduce SPIRAL (Self-improving Path Integration and Realignment), a framework that uses the model's own text-path behavior as a teacher to supervise its image-path behavior. SPIRAL employs two complementary alignment strategies:
SPIRAL significantly closes the performance gap between rendered-image inputs and native-text inputs. On the VTCBench, the framework improved the Qwen3-VL-8B model's performance from 35.10 to 54.02, nearly reaching the native text-input performance of 55.60. The authors found that OPD is particularly effective for retrieval tasks due to its fine-grained, sample-efficient nature, while DPO excels at reasoning and memory tasks by providing a more holistic, sequence-level signal. These benefits generalize across different backbones and out-of-domain benchmarks, confirming that the primary bottleneck in VTC is representational alignment rather than information density.
This work reframes the challenge of long-context multimodal modeling from "how to compress more" to "how to align visual representations with linguistic semantics." By demonstrating that models can self-correct their cross-modal inconsistencies without external supervision, the authors provide a scalable path for deploying efficient, high-performance VTC systems that maintain the reasoning capabilities of native-text models.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.