ResearchPod Summary
Knowledge distillation allows smaller models to replicate the capabilities of powerful, proprietary teacher models by training on their outputs. To protect intellectual property, providers have introduced output perturbation defenses—methods that modify the teacher's responses to degrade the quality of the distilled student model. However, these defenses are often evaluated against a narrow, implicit baseline: an attacker who queries each prompt exactly once and uses the raw output for training. This paper argues that this lack of a standardized threat model makes it impossible to compare different defenses or assess their true robustness against sophisticated adversaries.
The authors propose a formal framework to categorize attackers along three dimensions:
By separating the query budget from the data budget, the framework accounts for strategic behaviors—such as re-querying prompts that yielded poor results—that are ignored in standard evaluations. The interface profile further captures how specific API design choices (like whether a provider allows a user to 'prefill' the start of a model's response) can fundamentally change the attacker's ability to bypass a defense.
To demonstrate the framework, the authors evaluate Antidistillation Sampling (ADS), a defense that perturbs the teacher's output distribution. By varying only the interface profile—specifically, whether the attacker is permitted to inject a response prefix—the authors show that the effectiveness of ADS fluctuates significantly. This proves that a defense's 'success' is not a static value but a conditional outcome. The authors warn that deploying such defenses without a rigorous threat model creates a false sense of security, potentially leading organizations to rely on ineffective protections for their intellectual property.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.