Existing algorithms for generating Counterfactual Explanations (CXs) for Machine Learning (ML) typically assume fully specified inputs. However, real-world data often contains missing values, and the impact of these incomplete inputs on the performance of existing CX methods remains unexplored. To address this gap, we systematically evaluate recent CX generation methods on their ability to provide valid and plausible counterfactuals when inputs are incomplete. As part of this investigation, we hypothesize that robust CX generation methods will be better suited to address the challenge of providing valid and plausible counterfactuals when inputs are incomplete. Our findings reveal that while robust CX methods achieve higher validity than non-robust ones, all methods struggle to find valid counterfactuals. These results motivate the need for new CX methods capable of handling incomplete inputs.
Alex: Welcome to another episode of ResearchPod. Sam, I've been thinking about how machine learning models give advice in the real world—like what to tweak in a power plant to hit energy targets—but what if a sensor fails and drops key data?
Sam: This episode draws from a study by Leofante, Neider, and Yalçıner called "Evaluating Counterfactual Explanation Methods on Incomplete Inputs." The central puzzle it tackles is whether explanations from machine learning models remain trustworthy when real-world sensor data has gaps, like a missing humidity reading in a power plant.
Alex: So this paper is basically asking if the usual ways to generate helpful "what-if" changes for a model's predictions hold up when inputs are incomplete?
Sam: Yes, exactly. Machine learning models predict outcomes—like whether a combined cycle power plant produces enough electricity based on temperature, pressure, humidity, and steam pressure from sensors. But sensors can fail, leaving gaps. Standard explanations assume complete data, so they break down.
Alex: Right, and those explanations are supposed to show the smallest tweaks to flip a bad prediction to a good one, like adjusting conditions to meet the power target.
Sam: Researchers call these "counterfactual explanations," or CXs for short—they're minimal changes to inputs that change the model's output to something desirable. The trouble starts with imputation: when data's missing, you guess the value using patterns from similar past data, like averaging or finding close matches. But those guesses are often off, so the suggested change works on the guessed input but fails in reality, as shown in the paper's power plant example where a faulty humidity sensor misleads the advice.
Alex: So the operator follows the model's tip based on a bad guess, and the plant still underperforms?
Sam: That's the core problem. The study hypothesizes that some explanation methods, built to handle slight wobbles in the model or data—like robust CX methods—might tolerate these imputation errors better than basic ones. They test this across ten methods, imputing missing values in various ways, then checking if the advice actually flips the prediction on the true, unseen input—what they term "recourse validity."
Alex: Okay, so even robust ones help more, but none are perfect yet. What exactly did the researchers do to test that—how did they set up the comparison?
Sam: They started by taking real datasets with complete sensor data, like the power plant readings, then artificially hid some values to create incomplete inputs. For each incomplete case, they tried several ways to guess the missing numbers—simple averages from similar past data, or more advanced matching to nearby examples. Then, for each guessed input, they ran ten different explanation methods to suggest changes that should flip the prediction. Finally, they checked if those changes actually worked when applied back to the true, hidden input—that's the recourse validity.
Alex: Okay, so they separate the usual check—does the change work on the guessed data?—from the real test: does it work on the actual data? That makes sense for catching where guesses go wrong.
Sam: Precisely. The standard check is just counterfactual validity: does the suggested point get classified right? But recourse validity goes further—it tests if adding the suggested tweaks to the true input flips it correctly, even though the true input was never seen during explanation. The study compared non-robust methods against robust ones designed to handle wobbles—like if the model gets retrained on new data, or if actions have a bit of noise when applied.
Alex: Wait, so these robust methods are built for those kinds of model uncertainties, and the paper suggests that preparation helps with bad guesses too?
Sam: Yes—the logic is that imputation creates uncertainty like those perturbations, so robust methods tolerate it better and deliver higher recourse validity. In tests across datasets, robust approaches succeeded about twice as often as non-robust ones on true inputs after imputation. This challenges the idea that good guessing alone fixes explanations.
Alex: Huh. So it's not just patching the data—it's about explanations that bend without breaking under real messiness. How did they actually measure it across different setups—like what other checks besides recourse validity?
Sam: They used four common tabular datasets with sensor-like continuous features, such as wine quality ratings or diabetes risk factors. For each, they trained a basic two-layer neural network classifier, then tested on complete examples while randomly hiding one, two, or three feature values per example. They imputed them, generated explanations with ten methods—five non-robust baselines and five robust ones plus ARMIN tuned for gaps—and scored four metrics.
Alex: Got it—so hiding a few values mimics real sensor drops. What were the key metrics?
Sam: Standard validity checks if the suggested change flips the prediction on the imputed input. Recourse validity tests the true input. They also tracked cost—the total tweak size across features—and plausibility, which checks how realistic the change seems by comparing it to the training data crowd. If it sticks out like a sore thumb, the score drops.
Alex: Right, so low outlier score means the tweak feels natural. What did the results show?
Sam: Robust methods showed higher recourse validity, holding up even with three missing values—a significant edge confirmed by statistical tests. ARMIN matched robust levels, thanks to its handling of multiple guesses. All methods saw drops as gaps grew, and costs were similar between groups, meaning robustness adds reliability without bigger changes.
Alex: Huh—so the robustness pays off in trustworthiness, even if tweaks stay practical-sized, but nothing's foolproof with more missing data. What does that drop look like, and why?
Sam: As the number of hidden values rises from one to three, recourse validity falls for every method. Even top performers fail a notable share of the time. Each extra guess adds more uncertainty, pushing suggested changes further from what would truly work on the real input.
Alex: Right, more guesses compound the errors. And nearly all methods hit full counterfactual validity on the guessed data itself?
Sam: Yes, nearly all hit 100% there. The exception is Wachter, which lands lower because its step-by-step adjustments can get stuck midway.
Alex: Huh, so reliability isn't free, but it's not wildly expensive either. How did plausibility play out, and did data traits matter?
Sam: Most robust methods score well on plausibility, staying close to typical training examples. Data-hunting non-robust ones do fine too. ARMIN lags because its multi-guess caution pulls it toward unusual spots. Across datasets, power proves toughest—likely from its tighter feature ranges making guesses riskier.
Alex: So data traits matter a lot—power's like a narrow path where slips hurt more. Any tweaks to methods help?
Sam: For Wachter, tuning its step size boosts validity some, especially on easier data. But power stays challenging. Overall, the paper notes robust methods hold a meaningful validity lead, yet all falter enough to call for explanations built to handle gaps natively.
Alex: Yeah, those limits seem universal. What does the paper suggest as the path forward?
Sam: The study underscores that imputation is no full fix for trustworthy explanations in messy real-world settings. Robust methods offer a clear improvement, but validity drops for all as gaps multiply. This points to a need for explanation systems designed from the ground up to work directly with incomplete data, without relying on flawed guesses.
Alex: So, like building advice that accounts for the gaps upfront? That ties back to the power plant operator—who can't afford to act on untrustworthy tweaks.
Sam: Precisely. Future tools should incorporate uncertainty natively, ensuring recourses stay valid even when sensors drop readings. The paper's tests expose these shared weaknesses, highlighting robustness as a meaningful but partial bridge.
Alex: Makes sense—this work clarifies why current advice can mislead quietly, and pushes for better handling of real data flaws. That's a solid step in making machine learning explanations dependable. Thanks for joining, everyone.