ResearchPod Summary
Metal-organic frameworks are highly promising porous materials for gas storage, separation, and catalysis, with discovery increasingly driven by high-throughput computational screening. However, many experimentally reported crystal structures contain structural defects such as missing atoms, crystallographic disorder, or charge imbalance. These unreasonable structures severely distort predicted properties like electronic energy, band gaps, and gas adsorption. While various validation strategies exist—ranging from heuristic rules and geometry checks to graph-based machine learning models—they often suffer from limited interpretability, require proprietary access, or act as black boxes. This study investigates whether large language models can overcome these limitations by not only identifying unreasonable metal-organic frameworks, but also providing interpretable diagnostic rationales for their failure.
General-purpose language models cannot naturally interpret raw crystallographic information files because structural validity depends simultaneously on local coordination, framework connectivity, and charge distribution. To determine the optimal representation, the researchers benchmarked nine distinct descriptors spanning non-text, semi-text, and full-text formats across four large language models. Pre-trained models struggled to extract useful information from raw crystal text alone. However, after fine-tuning, models utilizing specialized, chemically informed descriptors such as mof2text and robocry achieved superior predictive performance. This demonstrates that successful validation depends heavily on organizing structural information into a linguistically learnable format rather than simply providing raw structural completeness.
When paired with mof2text, fine-tuned language models achieve validation accuracy and performance comparable to established graph-based models, while significantly outperforming geometry-based heuristics. Beyond binary classification, the text-based approach facilitates the generation of interpretable rationales. Through embedding analysis and fine-tuned multi-class classification, the models successfully identified specific defect origins such as abnormal bonding, coordination errors, connectivity issues, and charge imbalances. This bridges the gap between automated database curation and actionable structural refinement.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.