ResearchPod Summary
Synthetic data is typically scaled by either enlarging the source (adding seed questions or changing the teacher) or by increasing the generation budget for a fixed source. While these two approaches are often conflated, this paper disentangles them into Source Expansion (SE) and Fixed-Source Synthesis (FSS). The authors investigate whether FSS provides a reliable scaling axis, how it compares to SE at matched total-sample budgets, and whether common synthesis protocols (like persona prompting or trace-level repair) actually improve performance when the source is held constant.
The authors isolate FSS by fixing the seed-question pool and the teacher model, then varying only the per-question response budget under Rejection Sampling (RS). They derive a three-parameter scaling law for FSS based on the coverage of latent features within the fixed source. To compare SE and FSS, they match the total number of training examples across different allocations of seed-question counts and response budgets. Finally, they audit various synthesis protocols—such as temperature tuning, embedding-based selection, and trace-level repair—by comparing them against a plain RS baseline under the same fixed-source control.
This work provides a principled framework for compute allocation in synthetic data pipelines. By showing that FSS is a bounded axis, the authors suggest that researchers should prioritize expanding the source (SE) rather than over-investing in complex synthesis protocols or excessive sampling from a limited seed set. This distinction helps clarify why some synthetic data scaling efforts succeed while others hit performance ceilings.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.