ResearchPod Summary
In online advertising, click-through rate (CTR) models benefit significantly from analyzing long user interaction histories. However, these models face a critical deployment bottleneck: ad auctions require scoring within a few hundred milliseconds, making it computationally impossible to run large, full-history transformer models at request time. This paper investigates how to maintain the predictive power of long-range sequential modeling while adhering to strict production latency budgets.
The researchers introduce a multi-stage architecture that splits the workload between offline and online components. A high-capacity offline transformer asynchronously processes a user's full cross-surface interaction history and caches a compact, fixed-size representation in a feature store. At serving time, a lightweight runtime transformer processes only the most recent user events. The final CTR prediction is generated by a two-tower model that combines the cached long-term representation with the fresh short-term output from the runtime encoder.
The model is pre-trained using a dual-objective approach—feedback prediction and next-item prediction—on a year of cross-surface interaction logs. This ensures the transformer learns a robust, unified representation of user intent across organic search, product galleries, and advertising networks. During fine-tuning, the system is optimized for CTR using a DCNv2-based two-tower structure, which allows for efficient scoring against candidate banners.
The split architecture successfully recovers 72–80% of the quality of a full-history transformer that would otherwise be too expensive to deploy. In production A/B testing, this design delivered significant performance improvements: a +2.77% increase in the primary ranking metric for search advertising and a +2.1% increase for the Yandex Advertising Network. These gains were achieved without increasing serving latency, demonstrating that the cached offline representation is sufficiently robust to handle the inherent staleness of batch-processed data.
This work provides a scalable blueprint for deploying large-scale sequential models in latency-sensitive environments. By demonstrating that long-range behavioral signals can be effectively distilled into cached embeddings, the authors offer a practical solution to the 'latency-quality' trade-off that currently limits the adoption of advanced transformer architectures in real-time recommendation and advertising systems.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.