ResearchPod Summary
Deep neural networks are vulnerable to adversarial perturbations, and while adversarial training improves empirical robustness, it often makes formal verification difficult. Conversely, certified training methods that use Interval Bound Propagation (IBP) produce verifiable models but often suffer from lower standard accuracy. This paper investigates whether transferring adversarial knowledge from a robust teacher to a certified student can bridge this performance gap.
The authors introduce AD-CERT, a training objective that combines adversarial distillation with IBP-based certification. The student model is trained to match the predictive distribution of a fixed, empirically robust teacher on adversarial examples (the distillation component) while simultaneously minimizing a sound upper bound on the worst-case loss using IBP (the certification component). By distilling at the logit level, the student learns to mimic the teacher's robust behavior without requiring the teacher's complex training process to be directly verifiable.
AD-CERT achieves state-of-the-art certified accuracy for ReLU-based architectures across standard benchmarks, including MNIST, CIFAR-10, and TinyImageNet. The authors demonstrate that logit-level distillation is more effective than feature-space distillation for this purpose, providing a 5.40 percentage point improvement in certified accuracy in unified experimental setups. The distillation branch acts as a smooth, teacher-guided surrogate for the adversarial lower bound, which helps the model maintain high standard accuracy while remaining amenable to formal verification.
This work provides a practical and effective way to improve the certified robustness of neural networks. By decoupling the empirical robustness (provided by the teacher) from the formal certification (provided by the student's IBP loss), researchers can leverage existing robust models to enhance the performance of verifiable systems. This is particularly relevant for safety-critical applications where both high accuracy and formal guarantees are required.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.