ResearchPod Summary
Large-scale neural recommender systems often use a softmax cross-entropy objective, which requires computing logits for every item in the catalog. This is computationally expensive and memory-intensive. To mitigate this, practitioners use sampled softmax, which approximates the objective by considering only a small subset of negative items. However, when operating under a fixed memory budget, it is unclear how to optimally allocate resources between the batch size (n) and the number of sampled negative items (k). This paper investigates this trade-off to determine which configuration leads to faster convergence and better model performance.
The authors model the sampled-softmax training process as a stochastic optimization problem. They analyze the variance of the stochastic gradient estimator, which is influenced by two distinct sources of noise: the mini-batch sampling of users and the class-sampling of negative items. By treating the final classification layer as a multi-class logistic regression model, the authors derive a theoretical framework to quantify how the gradient variance depends on the choice of n and k. They then test their derived allocation rule across multiple sequential recommendation benchmarks, including MovieLens-20M and Gowalla, using both SGD and Adam optimizers.
The theoretical analysis reveals that the gradient variance is minimized when the number of training examples (batch size) is maximized. The authors conclude that, under a fixed memory constraint, the most efficient allocation is to prioritize larger batches over a larger number of negative samples. While the theoretical optimum suggests k ~ 1, the authors propose a practical implementation rule of n = k = sqrt(B) for computational efficiency. Experiments confirm that this configuration consistently achieves faster convergence and superior final recommendation quality compared to imbalanced alternatives that prioritize more negative samples at the expense of batch size.
This research provides a rigorous, actionable guideline for training large-scale recommender systems, moving away from heuristic-based hyperparameter tuning. By understanding the interplay between batch size and negative sampling, practitioners can optimize their training pipelines to be more efficient and effective, ultimately reducing the computational cost of training state-of-the-art recommendation models.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.