Ryan Shandler, Michael L. Gross, Yahli Shereshevsky
4 min
As artificial intelligence (AI) becomes increasingly integrated into military targeting, a central concern is whether human operators will succumb to automation bias—the tendency to uncritically accept algorithmic suggestions. This study challenges the assumption that automation bias is an inevitable consequence of AI integration. By deploying a high-fidelity replica of an AI decision-support system (DSS) and conducting two experiments with 2,015 Israeli military personnel, the authors examine how soldiers actually process and respond to AI-generated targeting recommendations in the context of active conflict.
Contrary to the fear that soldiers will blindly defer to machines, the researchers found strong evidence of algorithmic aversion. Participants were consistently more skeptical of strike recommendations when they were attributed to an AI system compared to those attributed to human intelligence analysts. This aversion was not uniform; it was most pronounced in high-stakes scenarios where the expected collateral damage was significant. The study suggests that when the moral and operational consequences of a decision are high, military personnel are more likely to favor human judgment, likely due to a perceived need for social accountability and context-awareness.
In a second experiment, the authors tested whether "explainable AI" (XAI) could mitigate this aversion. By providing participants with a brief, partial explanation of the key data inputs (e.g., intelligence reports or surveillance signals) that informed the AI's recommendation, the researchers found that algorithmic aversion was entirely eliminated. This intervention did not require full transparency of the underlying "black box" logic; rather, it provided enough insight for operators to feel they could exercise meaningful judgment. This suggests that transparency—even in a limited form—is a powerful tool for fostering calibrated trust and restoring a sense of agency to human operators.
This research shifts the debate from the inevitability of automation bias to the importance of human-AI interaction design. It demonstrates that military personnel are not passive recipients of algorithmic output but active, cautious evaluators. By showing that explainability can bridge the trust gap, the study offers a practical pathway for military institutions to leverage AI capabilities while maintaining the necessary human oversight to ensure legal and ethical compliance in warfare.
How is AI transforming decision-making in modern conflict? This study provides a unique empirical window into that question by deploying a high-fidelity replica of an AI decision-support system (DSS) used in military targeting. After reconstructing the interface and functionality of the real-world system, we tested its impact on combat decisions in two experiments involving 2,015 Israeli military personnel. Contrary to widespread fears of automation bias, we find strong evidence of algorithmic aversion, especially in scenarios involving high collateral damage. Yet we also show that integrating “explainable AI” features reduces algorithmic aversion and promotes more thoughtful evaluations of algorithmic recommendations. These findings challenge prevailing assumptions, revealing that trust in military AI is dynamic, varying with individual predispositions, perceived operational stakes, and the informational features of the interface. By grounding normative concerns in empirical evidence, our study offers critical insight into the integration of AI in warfare and underscores the enduring importance of human agency in high-stakes military decision-making.
Sam: So it's not about explaining the math underneath the model. It's giving the operator a causal hook that mimics how they'd reason through the problem themselves, so they can mentally check the logic instead of just accepting or rejecting a verdict.
Alex: That's the mechanism the authors point to. Showing the inputs — rather than the computation — turns the AI from an opaque oracle into something closer to a partner offering a rationale. That shift reduced the aversion and, notably, encouraged more calibrated evaluation rather than blanket rejection or blanket acceptance.
Sam: A referee would probably still ask how far that generalizes, though — a three-input display is a fairly minimal intervention, and this is one simulated system built around one military's targeting workflow.
Alex: That's a fair pushback, and it's the main constraint on the result. The high fidelity of the simulation is also what limits it — it's built around one operational context, with participants who may not represent every branch or every command culture. Whether a similar transparency feature would work the same way in a different decision-support setting, or with less binary stakes than a strike recommendation, is untested here. [[RP_SECTION:implications-for-human-control|Implications for human control]]
Sam: Still, it reframes what "meaningful human control" should mean in practice. The lever isn't just making the algorithm more accurate — it's making sure the interface lets the operator's own reasoning process stay engaged.
Alex: That's the core implication. Oversight isn't just an algorithm-performance problem, it's an interface-design problem — and the two get evaluated very differently.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.