ResearchPod Summary
Medical image classification, particularly for distinguishing keloids from hypertrophic scars, often suffers from limited labeled data and significant variations in image acquisition across clinical sites. While large vision-language models (VLMs) show promise, they often require sending sensitive patient data to external servers and function as opaque black boxes. This paper asks: Can an LLM contribute clinical knowledge to medical image classification without observing patient images or making the final diagnosis, while ensuring the resulting system is auditable and data-efficient?
The authors introduce ScaFE (Scar Feature Engineering), a framework that treats the LLM as a knowledge-driven feature engineer rather than a classifier. Instead of diagnosing images, the LLM retrieves clinical evidence and writes Python code to extract specific, visually assessable scar attributes (e.g., color, texture, morphology). These programs are executed in a restricted, local environment where raw images never leave the facility. A validation-guided loop allows the LLM to refine the code based on aggregate performance metrics and SHAP feature importance, ensuring the final features are both clinically grounded and effective for a lightweight Random Forest classifier.
ScaFE demonstrates superior performance in leave-one-site-out evaluations across 600 clinical photographs. It achieved 81.0% site-macro balanced accuracy, outperforming the strongest baseline, BiomedCLIP, by 10.0 percentage points. The framework is notably data-efficient, retaining 72.0% balanced accuracy even when using only 10% of the available development data. Furthermore, the iterative refinement process significantly improved the reliability of the generated code, raising the executable-program rate from 66.7% to 95.0% and ensuring that 91.7% of the final features are backed by verified clinical evidence.
This approach provides a pathway for deploying AI in clinical settings where data privacy and auditability are paramount. By separating clinical knowledge (LLM) from measurement (local code) and decision-making (Random Forest), ScaFE avoids the risks associated with opaque, end-to-end deep learning models. It demonstrates that LLMs can be effectively utilized to transfer medical expertise into robust, interpretable, and locally executable diagnostic tools without requiring massive, centralized datasets.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.