Ralf Beuthan, Megan Coffee, Heejin Kim, Na Yeon Kim, Pedro Kringen, Elisabeth Hildt, Haekyung Lee, Seunggeun Lee, Emilie Wiinblad Mathez, Sira Maliphol, Vadim Pak, Yuna Park, Stephan Sonnenberg, Jesmin Jahan Tithi, Magnus Westerlund, Roberto V. Zicari
9 min
Abstract
The polygenic risk scores (PRS) have emerged as an important methodology for quantifying genetic predisposition to complex traits and clinical disease. Significant progress has been made in applying PRS to conditions such as obesity, cancer, and type 2 diabetes (T2DM). Studies have demonstrated that PRS can effectively identify individuals at high risk, thereby enabling early screening, personalized treatment, and targeted interventions for diseases with a genetic predisposition. One current limitation of PRS, however, is the lack of interpretability tools. To address this problem for T2DM, researchers at the Graduate School of Data Science at the Seoul National University introduced eXplainable PRS (XPRS). This visualization tool decomposes PRSs into gene-level and single-nucleotide polymorphism (SNP) contribution scores via Shapley Additive Explanations (SHAP), providing granular insights into the specific genetic factors driving an individual's risk profile. We used a co-design approach to assess XPRS trustworthiness by considering legal, medical, ethical, and technical robustness during early design and potential clinical use. For that, we used Z-inspection, an ethically aligned Trustworthy AI co-design methodology, and piloted the Council of Europe's Human Rights, Democracy, and the Rule of Law Impact Assessment for AI Systems (HUDERIA) (Council of Europe (CAI) 2025). The findings of this use-case comprise a comprehensive set of ethical, legal, and technical lessons learned. These insights, identified by a multidisciplinary team of experts (ethics, legal, human rights, computer science, and medical), serve as a framework for designers to navigate future challenges with this and other AI systems. The findings also provide a useful reference for researchers developing explainability frameworks for PRS in diverse clinical contexts.
Alex: Technology readiness level—that's like stages from idea to everyday use, right? So if accuracy is shaky across groups, how does that affect trust in the breakdowns?
Sam: Exactly—levels go from basic concepts at 1 up to proven in operations at 9. The groups stressed tailoring displays with confidence levels so doctors can inform patients clearly, avoiding confusion that leads to wrong advice. A core medical worry was clinical utility—figuring out exactly when and how doctors would use it, like for screening or lifestyle talks. Without clear scenarios, it risks inconsistent decisions.
Alex: Right, because a fancy chart is useless if it baffles the doctor mid-appointment. And with risks based on probabilities from population data, not personal guarantees, how do they avoid over-relying on it?
Sam: They suggested showing prediction confidence, like wider ranges for unsure cases, and framing it as support for judgment, not a final call. This guards against automation bias, where people follow the tool blindly. Other points included fairness across ancestries, where performance might drop outside East Asian data, and selection bias from research cohorts not matching real patients.
Alex: Huh... so probabilistic means chances from averages, not certainties. That could mislead without those cues.
Sam: Yes. These tensions—like more details aiding explanation but complicating interfaces—highlight needs for validation before scaling.
Alex: So the assessments point to real steps needed, without overstating readiness. Measured progress makes sense, given the medical flags. But how does this fit into the bigger legal picture for tools like XPRS in Korea?
Sam: Korea's new AI Basic Act, passed in late 2024, sets rules for what's called high-impact AI. These are systems that could seriously affect people's lives, safety, or basic rights—like tools in healthcare that guide decisions on monitoring or treatments. For XPRS, which processes genetic data for diabetes risk, it might qualify if used clinically, requiring operators to ensure transparency, safety plans, and impact checks. It mirrors the EU AI Act's high-risk category.
Alex: Okay, so high-impact means real-world stakes. What about human rights specifically—did the legal group spot positives or tensions?
Sam: The group highlighted positives: XPRS could advance the right to health and the right to share in science's benefits without discrimination. But tensions emerged in physical and mental integrity—like if inaccurate risks lead to unnecessary stress or invasive checks. They recommend high accuracy, patient choice, and clear info. Privacy is key too, as genetic info reveals health traits others shouldn't access without consent. The paper urges encryption and purpose limits to prevent misuse, say in insurance.
Alex: So legal guardrails push for caution in deployment. It's a solid framework to build on. But the legal group touched on human rights—what about equality and non-discrimination? With genetic data mostly from certain groups, doesn't that risk unfairness?
Sam: The working group noted patterns of bias in PRS training data, often from Korean and Japanese sources. This means predictions might not work as well for other ethnic groups. They map this to fairness pillars, stressing avoidance of unfair bias and technical reliability, since lower accuracy could lead to wrong diagnoses or skipped care. For children, PRS mixes thousands of gene tweaks in hard-to-see ways, and kids change fast biologically, so early risks might not hold.
Alex: Those ripples make sense. How did they tie this to bigger ethics?
Sam: They mapped issues to four pillars: respect for human autonomy, like consent; prevention of harm, covering data privacy; fairness, tackling bias; and explicability, ensuring clear communication. For XPRS, explainability helps but needs checks on accuracy trade-offs.
Alex: Cautious rollout ties back to everything we've discussed. Does the paper look at how tools like XPRS might apply beyond diabetes—to other diseases?
Sam: Yes, it points out where breaking down genetic risk scores works well. For diseases caused by hundreds of small DNA changes that add up gradually—like heart disease or high blood pressure—showing each change's role helps doctors understand the main pathways involved. This guides custom prevention, much like sorting puzzle pieces to see the full picture. The paper calls these polygenic diseases, from "poly" meaning many genes.
Alex: So for those buildup diseases, it's a fit. But what about ones driven by a single big genetic hit?
Sam: In contrast, diseases mainly from one rare, powerful DNA flaw don't gain much from this breakdown. Take inherited high cholesterol or cystic fibrosis—doctors focus on spotting that one flaw and treating it directly, like fixing a broken pipe rather than tallying tiny leaks everywhere. These are known as monogenic disorders, meaning one gene dominates.
Alex: Huh, that distinction clarifies when to use it. Pulling it all together, what are the big lessons from this evaluation?
Sam: The study stresses trustworthy AI isn't just about good predictions—it's about clear use cases, known limits, and explanations that fit doctors and patients. Interpretability depends on who’s using it; what helps researchers might confuse a clinic visit. Key steps forward include defining users early, listing "don't use" situations, and testing if breakdowns truly aid decisions. With East Asian data limits and no real-clinic tests yet, it's at lab-validation stage.
Alex: Makes sense—a solid step in making genetic tools practical and honest. Thanks, Sam, for breaking this down. Thanks for listening to ResearchPod.