Md Aminur Hossain, Omkumar Vaghasiya, Rajeev Ranjan Dwivedi, Vinod Kurmi, Biplab Banerjee
5 min
Abstract
Remote sensing foundation models (RSFMs) are commonly evaluated using aggregate metrics, which can hide systematic performance disparities across ecological regions. We introduce FairRSFM, a biome-aware benchmark for evaluating ecological group robustness in RSFMs. FairRSFM maps georeferenced samples from 14 terrestrial biome classes into six ecologically meaningful macro-groups and evaluates models under a unified frozen-backbone evaluation protocol. The benchmark covers four downstream datasets: m-EuroSAT, m-BigEarthNet, m-SA-Crop-Type, and MMEarth20K with Dynamic World label maps. Using Prithvi-EO-2.0, SatMAE, and DOFA across three random seeds, we show that aggregate performance consistently masks biome-dependent disparities across architectures and tasks. For example, Prithvi-EO-2.0 reaches 90.98% overall macro-F1 on m-EuroSAT but a mean worst-group score of only 83.72%, while m-SA-Crop-Type drops from 27.30% overall mIoU to 18.47% in the Xeric and Mineralogical group. We further evaluate Biome-Orthogonal Linear Probing (BOLP), Dynamic Biome Reweighting (DBR), and GroupDRO as complementary mitigation baselines. Their effectiveness is model- and task-dependent; for example, BOLP improves Prithvi-EO-2.0 worst-group F1@opt on m-BigEarthNet from 46.12% to 50.27% without updating the RSFM backbone. FairRSFM provides a reusable protocol for diagnosing and mitigating ecological robustness gaps in remote sensing foundation models. Code and datasets are available at: https://github.com/aminurhossain/FairRSFM.
Alex: DBR operates purely at the loss level. It rebalances the objective using validation performance by biome, whereas BOLP changes the representation space itself.
Sam: So one is a soft constraint on training and the other is a hard constraint on the features. Does DBR share the information-loss problem?
Alex: Less so. It discards nothing and only changes the gradient weight for particular groups. It tends to be more stable, but it is less effective at removing deep-seated representation bias. So BOLP is the more surgical option when the backbone is clearly skewed, and DBR is the more conservative one for general fine-tuning.
Sam: Across both methods, though, the authors describe efficacy as inconsistent. It varies with architecture and task.
Alex: Yes, and that limits how much the mitigation results can support. I would treat the benchmark's diagnostic contribution as the load-bearing part. The gaps are systematic, and they are invisible to aggregate metrics like mean intersection-over-union or top-one accuracy. Reporting group-robustness metrics such as Normalized Failure Range and Equalized Odds Disparity makes them visible.
Sam: And the grouping itself is shaky, isn't it? A patch on an ecotone gets one biome label.
Alex: The authors acknowledge it. They use a centroid-based spatial join, which is not pixel-perfect but gives a consistent protocol for diagnosing gaps across datasets. The noise is real. The consistency is what makes cross-dataset comparison possible. The authors are also clear that biome awareness captures only one slice of geographic bias. Sensor-level shifts, temporal acquisition effects, and uneven imagery coverage are still unaddressed.
Sam: So the practical point is that these models shouldn't be treated as universal encoders. One trained mostly on temperate landscapes may be unreliable in arid or cryospheric regions, and a mean score won't warn you.
Alex: Right. The benchmark shifts the question from whether a model works to where it fails. BOLP is a useful option for groups without the compute to retrain a large vision transformer, but its value rests on that diagnostic framing more than on any claim that the bias is solved.
Sam: That seems the defensible reading to me. It is a more transparent audit of what these models know about the planet, rather than what they know about their training distribution.
Alex: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Sam: Thanks for listening.