ResearchPod Summary
Deploying transformer models (like BERT, ViT, or LLaMA) under Fully Homomorphic Encryption (FHE) is computationally expensive because non-linear operations—such as softmax, normalization, and activation functions—are incompatible with FHE's polynomial arithmetic. Current solutions typically replace these functions with polynomial approximations using a uniform configuration across all layers. This approach is rigid and inefficient, as it fails to exploit the fact that different layers have varying tolerances for approximation error. However, manually tuning these parameters for each layer is impossible due to the astronomical size of the search space (up to 10^225 configurations).
ATLAS addresses this by automating the selection of per-layer approximation hyperparameters. It treats the problem as a multi-objective optimization task, balancing end-to-end inference latency against predictive accuracy. To manage the complexity, ATLAS employs a two-stage evolutionary search strategy that progressively relaxes constraints and uses surrogate models to bypass the high cost of evaluating every candidate configuration. The framework is model-agnostic, meaning it can be applied as a post-processing step to any existing FHE-compatible transformer without requiring retraining or modifications to the underlying encryption or packing schemes.
Empirical evaluations on BERT-base, LLaMA3-8B, and ViT-base demonstrate that ATLAS significantly outperforms expert-designed, uniform-configuration baselines. By assigning heterogeneous approximation settings to each layer, ATLAS reduces multiplicative depth by 17-35% and end-to-end inference latency by 20-25%, all while maintaining negligible accuracy loss. Notably, the entire optimization process completes in under one hour, making it a practical tool for deploying secure, private inference models.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.