ResearchPod Summary
As tabular foundation models (TFMs) gain popularity for analyzing microbiome data, a critical question arises: how robust are these models when the query data distribution differs from the support set provided during in-context learning? Because microbiome data is characterized by compositionality and high zero-inflation, the authors investigate whether TFMs can maintain performance when faced with realistic shifts in these structural properties.
The authors introduce a benchmark using six gut microbiome datasets across four disease contexts. They employ an in-context learning evaluation protocol where the support set remains unperturbed, while the query samples undergo three types of biologically inspired perturbations: feature removal, increased zero-inflation (sparsification), and zero-imputation. To ensure the findings reflect robustness to distribution shift rather than the loss of predictive signal, the authors use a data-driven method to identify and protect the most discriminative taxa from these perturbations.
The study reveals that protecting discriminative features is insufficient to guarantee model stability. All three perturbation strategies consistently degrade performance across all tested TFMs. Zero-imputation—filling in zero values with observed abundances—proved to be the most harmful, indicating that TFMs are highly sensitive to changes in the global feature structure. Furthermore, TFMs showed greater sensitivity to sparsification (increased zero-inflation) compared to a classical random forest baseline, suggesting that the inductive biases of current foundation models may be poorly suited to the extreme sparsity inherent in microbiome data.
These results highlight a significant vulnerability in applying current tabular foundation models to metagenomic data. The findings suggest that while TFMs perform well on standard benchmarks, their reliance on global data patterns makes them fragile in the face of the technical and biological variability common in microbiome studies. This underscores the need for more robust, domain-aware architectures that can handle the unique compositional and sparse nature of taxonomic abundance data.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.