ResearchPod Summary
Tabular foundation models (TFMs) have become powerful tools for in-context learning (ICL), allowing users to perform predictions on structured data without task-specific training. However, because these models rely on sensitive records placed directly in the prompt context, they are vulnerable to membership inference attacks (MIAs) that can leak private information. This paper investigates whether formal privacy protections can be applied to tabular ICL without the common, yet often unrealistic, requirement of having access to public in-distribution data for knowledge transfer.
The authors first demonstrate that tabular ICL is indeed vulnerable to various MIAs, ranging from passive grey-box attacks to active context-manipulation attacks. To address this, they introduce TabPATE, a framework based on the Private Aggregation of Teacher Ensembles (PATE). TabPATE partitions private data among multiple teacher models and aggregates their predictions on synthetic queries to produce a safe, labeled student context. Unlike prior PATE-style methods that require public data to generate these queries, TabPATE leverages the structured nature of tabular data—specifically, known feature types and bounded ranges—to generate synthetic queries directly from the data schema or from lightly privatized marginal statistics.
TabPATE successfully provides a practical path to private tabular ICL. Across standard benchmarks, the framework maintains competitive utility at moderate privacy budgets while reducing the success of membership inference attacks to near-random levels. The authors show that while simple data-independent queries are often sufficient for balanced classification, allocating a small portion of the privacy budget to release private marginal statistics significantly improves performance on more complex or imbalanced datasets. Furthermore, the approach extends effectively to regression tasks, providing a robust, public-data-free alternative to existing private ICL methods.
This work provides a critical privacy-preserving mechanism for sensitive domains like healthcare and finance, where tabular data is prevalent but public datasets are often unavailable or restricted. By removing the dependency on public data, TabPATE makes differentially private in-context learning more accessible and applicable to real-world scenarios where data privacy is paramount.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.