Francesco Leofante, Daniel Neider, Mustafa Yalçıner
8 min
Abstract
Existing algorithms for generating Counterfactual Explanations (CXs) for Machine Learning (ML) typically assume fully specified inputs. However, real-world data often contains missing values, and the impact of these incomplete inputs on the performance of existing CX methods remains unexplored. To address this gap, we systematically evaluate recent CX generation methods on their ability to provide valid and plausible counterfactuals when inputs are incomplete. As part of this investigation, we hypothesize that robust CX generation methods will be better suited to address the challenge of providing valid and plausible counterfactuals when inputs are incomplete. Our findings reveal that while robust CX methods achieve higher validity than non-robust ones, all methods struggle to find valid counterfactuals. These results motivate the need for new CX methods capable of handling incomplete inputs.
Alex: Okay, so they separate the usual check—does the change work on the guessed data?—from the real test: does it work on the actual data? That makes sense for catching where guesses go wrong.
Sam: Precisely. The standard check is just counterfactual validity: does the suggested point get classified right? But recourse validity goes further—it tests if adding the suggested tweaks to the true input flips it correctly, even though the true input was never seen during explanation. The study compared non-robust methods against robust ones designed to handle wobbles—like if the model gets retrained on new data, or if actions have a bit of noise when applied.
Alex: Wait, so these robust methods are built for those kinds of model uncertainties, and the paper suggests that preparation helps with bad guesses too?
Sam: Yes—the logic is that imputation creates uncertainty like those perturbations, so robust methods tolerate it better and deliver higher recourse validity. In tests across datasets, robust approaches succeeded about twice as often as non-robust ones on true inputs after imputation. This challenges the idea that good guessing alone fixes explanations.
Alex: Huh. So it's not just patching the data—it's about explanations that bend without breaking under real messiness. How did they actually measure it across different setups—like what other checks besides recourse validity?
Sam: They used four common tabular datasets with sensor-like continuous features, such as wine quality ratings or diabetes risk factors. For each, they trained a basic two-layer neural network classifier, then tested on complete examples while randomly hiding one, two, or three feature values per example. They imputed them, generated explanations with ten methods—five non-robust baselines and five robust ones plus ARMIN tuned for gaps—and scored four metrics.
Alex: Got it—so hiding a few values mimics real sensor drops. What were the key metrics?
Sam: Standard validity checks if the suggested change flips the prediction on the imputed input. Recourse validity tests the true input. They also tracked cost—the total tweak size across features—and plausibility, which checks how realistic the change seems by comparing it to the training data crowd. If it sticks out like a sore thumb, the score drops.
Alex: Right, so low outlier score means the tweak feels natural. What did the results show?
Sam: Robust methods showed higher recourse validity, holding up even with three missing values—a significant edge confirmed by statistical tests. ARMIN matched robust levels, thanks to its handling of multiple guesses. All methods saw drops as gaps grew, and costs were similar between groups, meaning robustness adds reliability without bigger changes.
Alex: Huh—so the robustness pays off in trustworthiness, even if tweaks stay practical-sized, but nothing's foolproof with more missing data. What does that drop look like, and why?
Sam: As the number of hidden values rises from one to three, recourse validity falls for every method. Even top performers fail a notable share of the time. Each extra guess adds more uncertainty, pushing suggested changes further from what would truly work on the real input.
Alex: Right, more guesses compound the errors. And nearly all methods hit full counterfactual validity on the guessed data itself?
Sam: Yes, nearly all hit 100% there. The exception is Wachter, which lands lower because its step-by-step adjustments can get stuck midway.
Alex: Huh, so reliability isn't free, but it's not wildly expensive either. How did plausibility play out, and did data traits matter?
Sam: Most robust methods score well on plausibility, staying close to typical training examples. Data-hunting non-robust ones do fine too. ARMIN lags because its multi-guess caution pulls it toward unusual spots. Across datasets, power proves toughest—likely from its tighter feature ranges making guesses riskier.
Alex: So data traits matter a lot—power's like a narrow path where slips hurt more. Any tweaks to methods help?
Sam: For Wachter, tuning its step size boosts validity some, especially on easier data. But power stays challenging. Overall, the paper notes robust methods hold a meaningful validity lead, yet all falter enough to call for explanations built to handle gaps natively.
Alex: Yeah, those limits seem universal. What does the paper suggest as the path forward?
Sam: The study underscores that imputation is no full fix for trustworthy explanations in messy real-world settings. Robust methods offer a clear improvement, but validity drops for all as gaps multiply. This points to a need for explanation systems designed from the ground up to work directly with incomplete data, without relying on flawed guesses.
Alex: So, like building advice that accounts for the gaps upfront? That ties back to the power plant operator—who can't afford to act on untrustworthy tweaks.
Sam: Precisely. Future tools should incorporate uncertainty natively, ensuring recourses stay valid even when sensors drop readings. The paper's tests expose these shared weaknesses, highlighting robustness as a meaningful but partial bridge.
Alex: Makes sense—this work clarifies why current advice can mislead quietly, and pushes for better handling of real data flaws. That's a solid step in making machine learning explanations dependable. Thanks for joining, everyone.