ResearchPod Summary
In industrial e-commerce recommendation systems, multimodal representations (images and titles) are increasingly used to augment traditional ID-based models. The standard industry practice is a two-stage paradigm: pre-training a multimodal encoder on scenario-specific data, then freezing it to extract features for the downstream CTR prediction model. While researchers have explored end-to-end training (E2EM) to better align these encoders with CTR objectives, this paper reports that E2EM fails to improve upon well-pretrained encoders, often leading to performance degradation.
The authors identify two primary reasons for the failure of E2EM. First, raw CTR data is inherently noisy; user clicks are driven by a mix of multimodal semantics and non-multimodal factors (e.g., price, position bias, interest fatigue). Second, co-training an encoder with an ID-based CTR model does not automatically decouple these signals. Because the shared CTR loss propagates gradients from all samples, the multimodal encoder is forced to learn from non-semantic noise, which disrupts its fine-grained semantic understanding. Furthermore, E2EM is computationally expensive, making it impractical for large-scale industrial deployment.
To overcome these challenges, the authors propose a "Mine-Then-Train" framework. Instead of training the encoder on all raw CTR data, they first train a lightweight multimodal annotation model to identify click preferences that are explicitly explainable by multimodal content. They then use this model to mine high-quality, multimodally interpretable triplets from the CTR data. These triplets are filtered to ensure they contain clear semantic relevance and provide information gain beyond existing ID features. The multimodal encoder is then fine-tuned on these curated triplets using a triplet margin loss, while maintaining an SCL (Semantic-aware Contrastive Learning) regularization term to prevent the loss of general semantic knowledge. This approach effectively isolates high-quality supervision, allowing the encoder to learn representations that are natively aligned with user click preferences.
This research provides a robust, efficient alternative to end-to-end training that respects the constraints of large-scale industrial systems. By demonstrating that data quality is more critical than joint optimization for multimodal CTR models, the authors offer a practical roadmap for improving recommendation performance. Online A/B tests on Taobao’s display advertising system confirmed the effectiveness of this approach, yielding a 1.5% improvement in CTR and a 0.5% increase in revenue per mille (RPM).
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.