Bin Cao, Huichi Zhou, Runyu Yang, Jingsong Li, Shuchen Sun, Yan Song, Hanyu Gao, Zhongwei Yu, Tong-Yi Zhang, Jun Wang
9 min
Abstract
A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by diagnosing failures, revising skill instructions and code, and validating revisions before reuse, without retraining the language model or changing the underlying physical models. Skills selected using development data and frozen before held-out evaluation achieve higher refinement scores than the original expert-designed skills across FullProf, GSAS-II and PyWPEM. The agent resolves strongly overlapping reflections, quantifies a five-phase ancient Egyptian cosmetic, tracks lattice evolution in an operating battery and compares atomic configurations in a disordered oxide catalyst. On DeltaXRDbench, it leads the evaluated methods in single- and multiphase identification across simulated and experimental data. Without supplied composition, single-phase top-1 accuracies reach 96.30\%, 81.78\% and 40.83\% on MP500, RRUFF and opXRD, respectively, compared with 58.00\%, 58.47\% and 26.45\% for the strongest comparator. These results demonstrate how an integrated scientific tool ecosystem can support agents that extract structural knowledge from measurements while accumulating validated analytical expertise that transfers to new samples.
Sam: Give me a change that would matter at the laboratory bench, rather than just a better-written instruction.
Alex: In a training episode using GSAS-II, an established refinement program, changing several parameter groups together worsened the fit. Refinement means adjusting a proposed structural model to reproduce the measured pattern. The revised skill emphasizes starting from defaults and changing one group at a time.
Sam: So adjusting a crystal dimension improves one sample’s model. Changing when the procedure allows that adjustment could improve later analyses.
Alex: That’s the operational distinction the manuscript makes. Other revisions address software reliability, including checks that requested settings were actually passed through and saved outputs match reported diagnostics. Learning here includes knowing how to use the tools correctly.
Sam: A revised rule could fix its original failure and break something else. How do they avoid calling that improvement?
Alex: Training episodes generate proposed changes. Paired development evaluations decide which changes to retain. Then the selected procedure is frozen and compared with the original expert-designed workflow on held-out samples.
Sam: What is the clearest result from that frozen comparison?
Alex: For FullProf, another refinement program, the mean held-out score rose from about 47 to about 72. The score summarizes agreement between measured and calculated profiles. Failed or timed-out evaluations count as zero.
Sam: That isn’t the percentage of structures correctly identified, then. And did the same pattern appear with other software?
Alex: It is not an identification accuracy. The final workflows also improved mean held-out scores for GSAS-II and PyWPEM, the authors’ physics-constrained modelling engine. But the permitted changes differed between branches, so these aren’t identical interventions replicated across programs.
Sam: I’d separate that evidence from the agent’s overall identification performance. A good identification benchmark doesn’t, by itself, show that learning caused the advantage.
Alex: The manuscript separates those comparisons too. Frozen workflow tests assess transfer of revised procedures. The identification benchmark assesses the integrated system’s ability to find structures. It doesn’t isolate how much each individual skill edit contributed.
Sam: Before the benchmark results, how does the agent decide whether a proposed structure deserves refinement?
Alex: It retrieves candidate structures and compares their predicted diffraction with the measurement. Acceptance requires a favourable inspection verdict, applicable composition checks, and required peak-evidence checks. Agreement between retrieval tools can prioritize inspection, but cannot override failed evidence checks.
Sam: What happens when a single structure doesn’t explain the pattern? Does the agent just add more phases?
Alex: In automatic mode, completed rejection of the available single-phase hypotheses triggers mixture analysis. A failed tool call alone does not establish that the sample contains several phases. That distinction prevents a computing problem from becoming a scientific conclusion.
Sam: There’s also a danger in preprocessing: the convenient input for a prediction model might replace the actual measurement.
Alex: They keep those representations separate. Retrieval can use a resampled, normalized pattern, while peak analysis and refinement retain the full measured range and original intensity scale. The candidate is judged against experimental evidence, not merely against the representation convenient for retrieval.
Sam: How broad is the identification test, and which comparison should listeners remember?
Alex: DeltaXRDbench contains about 62,000 patterns, combining simulated data, experimental mineral measurements, and another experimental collection called opXRD. For the mineral collection, without supplied composition, Gan Jiang’s first candidate was correct about 82 percent of the time. AutoXRD, the strongest comparator there, reached about 58 percent.
Sam: Those are substantial differences, but experimental collections can vary. Was performance equally strong across them?
Alex: No. On opXRD, first-candidate accuracy without composition was about 41 percent, although it still led the evaluated methods. Many measurements in that collection cover incomplete angular ranges. A better ranking does not mean the structural problem is solved.
Sam: Mixtures sound harder still, because finding one component isn’t the same as finding every component.
Alex: The benchmark distinguishes partial recovery from getting the entire phase set right, with no extra phase. Without composition, Gan Jiang recovered the exact set in only about three percent of opXRD mixtures. It led the comparisons, but complete recovery remained difficult.
Sam: And those mixtures were assembled from source patterns, rather than all being measurements of physically mixed specimens.
Alex: Yes. Mixtures built from experimental profiles retain measured peak complexity, but don’t reproduce every effect of physical mixing. The application cases add concrete demonstrations, including an ancient Egyptian cosmetic, an operating battery, and competing atomic arrangements in an oxide catalyst.
Sam: What should we take from those demonstrations without overreading them?
Alex: They show different analytical tasks: quantifying overlapping components, following structural change across successive scans, and comparing candidate atomic configurations. The catalyst result is a diffraction-supported model within the configurations tested. More generally, the inspector’s acceptance means consistency with the profile, not proof of structural uniqueness.
Sam: And the phrase “self-learning” could suggest continuous improvement after deployment. Has that actually been demonstrated?
Alex: The experiments demonstrate transfer after controlled revision and freezing. They provide a basis for deployment-time learning, with regression checks, versioning, monitoring, and rollback. They do not provide a longitudinal demonstration of open-ended improvement during routine use.
Sam: Who should read the full manuscript, and where should they start?
Alex: Scientific-agent builders should start with “Self-improvement transfers analytical skills to unseen samples,” then the methods on evaluating that improvement. Diffraction researchers should read the identification benchmark comparisons and the application closest to their work. Pay attention to missing angular coverage and complete mixture recovery, not just which method ranks first.
Sam: And what should everyone else carry away?
Alex: Fixing one answer is not the same as improving the next procedure. Gan Jiang makes that second claim testable by revising external skills and checking them on unseen samples.
Sam: Keep that distinction in mind when you hear an agent called self-learning.