Remote sensing foundation models (RSFMs) are commonly evaluated using aggregate metrics, which can hide systematic performance disparities across ecological regions. We introduce FairRSFM, a biome-aware benchmark for evaluating ecological group robustness in RSFMs. FairRSFM maps georeferenced samples from 14 terrestrial biome classes into six ecologically meaningful macro-groups and evaluates models under a unified frozen-backbone evaluation protocol. The benchmark covers four downstream datasets: m-EuroSAT, m-BigEarthNet, m-SA-Crop-Type, and MMEarth20K with Dynamic World label maps. Using Prithvi-EO-2.0, SatMAE, and DOFA across three random seeds, we show that aggregate performance consistently masks biome-dependent disparities across architectures and tasks. For example, Prithvi-EO-2.0 reaches 90.98% overall macro-F1 on m-EuroSAT but a mean worst-group score of only 83.72%, while m-SA-Crop-Type drops from 27.30% overall mIoU to 18.47% in the Xeric and Mineralogical group. We further evaluate Biome-Orthogonal Linear Probing (BOLP), Dynamic Biome Reweighting (DBR), and GroupDRO as complementary mitigation baselines. Their effectiveness is model- and task-dependent; for example, BOLP improves Prithvi-EO-2.0 worst-group F1@opt on m-BigEarthNet from 46.12% to 50.27% without updating the RSFM backbone. FairRSFM provides a reusable protocol for diagnosing and mitigating ecological robustness gaps in remote sensing foundation models. Code and datasets are available at: https://github.com/aminurhossain/FairRSFM.
Alex: A benchmark for remote sensing foundation models, called FairRSFM, suggests that strong aggregate scores can hide large gaps across ecological regions. In crop-type mapping, performance drops by nearly ten points when the model moves into xeric or mineralogical biomes.
Sam: That undercuts global accuracy as a reliability signal. Is the underlying issue that frozen backbones extract features biased toward the common landscapes in their training data?
Alex: That is the working explanation. The model picks up an ecological accent, and the dominant biomes drive the validation score. Minority biomes can fail without moving the mean much. To address this without retraining the backbone, the authors propose Biome-Orthogonal Linear Probing, or BOLP.
Sam: What does it actually do to the embeddings?
Alex: It builds a matrix of per-biome mean embeddings, centers it, and runs a singular value decomposition to find the directions along which biomes differ most. It then projects the features onto the orthogonal complement of the top directions, so the linear probe never sees them. The aim is a classifier that relies on class-discriminative features rather than cues tied to a particular forest or desert.
Sam: Which raises the obvious cost question. If you remove those directions, are you throwing away useful signal?
Alex: There is a trade-off. BOLP improves worst-group robustness, but the paper reports a slight dip in aggregate performance. The authors treat that as a deliberate choice of equity across ecological regimes over raw mean accuracy.
Sam: I want to push on that, because in remote sensing the ecological context is often part of the class identity. Say a wetland vegetation type exists in only one biome. Its signature could live in exactly the directions you project out.
Alex: It could, and that is the central tension. The authors' defence is that the SVD targets mean differences between biomes, not class-specific variance, so it filters the accent rather than the class. But a class that is distinguishable only through its biome is arguably fragile under distribution shift anyway.
Sam: That is a reasonable argument, but it is an argument rather than a demonstration. The benchmark can't tell you in advance which biome-restricted classes are genuinely fragile and which are simply being erased.
Alex: Fair. What the approach does is trade confidence on well-represented training biomes for better behaviour on underrepresented ones. Whether that trade is right for a given class is something the group-level numbers can't fully resolve.
Sam: The paper also has a second mitigation, Dynamic Biome Reweighting. How does it differ?
Alex: DBR operates purely at the loss level. It rebalances the objective using validation performance by biome, whereas BOLP changes the representation space itself.
Sam: So one is a soft constraint on training and the other is a hard constraint on the features. Does DBR share the information-loss problem?
Alex: Less so. It discards nothing and only changes the gradient weight for particular groups. It tends to be more stable, but it is less effective at removing deep-seated representation bias. So BOLP is the more surgical option when the backbone is clearly skewed, and DBR is the more conservative one for general fine-tuning.
Sam: Across both methods, though, the authors describe efficacy as inconsistent. It varies with architecture and task.
Alex: Yes, and that limits how much the mitigation results can support. I would treat the benchmark's diagnostic contribution as the load-bearing part. The gaps are systematic, and they are invisible to aggregate metrics like mean intersection-over-union or top-one accuracy. Reporting group-robustness metrics such as Normalized Failure Range and Equalized Odds Disparity makes them visible.
Sam: And the grouping itself is shaky, isn't it? A patch on an ecotone gets one biome label.
Alex: The authors acknowledge it. They use a centroid-based spatial join, which is not pixel-perfect but gives a consistent protocol for diagnosing gaps across datasets. The noise is real. The consistency is what makes cross-dataset comparison possible. The authors are also clear that biome awareness captures only one slice of geographic bias. Sensor-level shifts, temporal acquisition effects, and uneven imagery coverage are still unaddressed.
Sam: So the practical point is that these models shouldn't be treated as universal encoders. One trained mostly on temperate landscapes may be unreliable in arid or cryospheric regions, and a mean score won't warn you.
Alex: Right. The benchmark shifts the question from whether a model works to where it fails. BOLP is a useful option for groups without the compute to retrain a large vision transformer, but its value rests on that diagnostic framing more than on any claim that the bias is solved.
Sam: That seems the defensible reading to me. It is a more transparent audit of what these models know about the planet, rather than what they know about their training distribution.
Alex: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Sam: Thanks for listening.