ResearchPod Summary
Genetic birth defects present a significant diagnostic challenge due to incomplete fetal phenotypes, heterogeneous clinical evidence, and the need to distinguish causal variants from thousands of sequencing-derived candidates. The authors sought to develop a computational workflow that effectively integrates diverse data sources—ranging from clinical rules and molecular predictors to protein structure and pathway context—to improve the prioritization of causal variants in prenatal and early-infant settings.
DeepBD utilizes a grounded agentic architecture that divides the interpretation process into four distinct layers:
The system was trained and evaluated on a large-scale in-house cohort of 18,622 fetal and infant cases, allowing the model to learn the specific distribution of competing variants encountered in clinical practice.
DeepBD demonstrated superior performance in causal variant prioritization compared to established baselines, including Exomiser, DeepRare, and prompted LLM reranking. With a Recall@1 of 0.658 and Recall@10 of 0.929, the system effectively handles diverse molecular consequences, such as missense, frameshift, and splice-site variants. Ablation studies confirmed that the integration of rule-based evidence, mechanistic context, and specialist refinement provides complementary signals, validating the necessity of a multi-layered, agentic approach over simpler, monolithic models.
This work addresses the post-sequencing bottleneck in clinical genomics by moving beyond simple rule-based filtering or unconstrained LLM reasoning. By grounding agentic capabilities in a stable, trainable evidence engine, DeepBD provides a scalable and auditable framework for clinicians to navigate the complexities of prenatal and neonatal genetic diagnosis, ultimately supporting faster and more accurate clinical decision-making.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.