AutoXRDAutonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis

Verifiable, stepwise powder X-ray diffraction analysis with executable crystallographic backends and scientific checks.

134XRD tasks
10recent LLMs
1,340model–task runs
625.6Mtrajectory tokens
Overview

From plausible answers to verifiable XRD workflows

Powder X-ray diffraction is central to materials characterization, yet reliable end-to-end automation remains challenging. AutoXRD organizes powder-XRD analysis into stepwise refinement procedures, grounds actions in observed evidence, and applies deterministic crystallographic and physical checks before accepting a result.

XRDBench evaluates the resulting agents through 100 diagnostic questions and 34 executable workflows across eleven XRD task families. Across ten models, performance decreases from 61.9 on diagnostic QA to 53.7 on E2E workflows, exposing persistent limitations in quantitative reasoning, indexing, Rietveld refinement, and controlled termination.

Method

AutoXRD: evidence-preserving scientific execution

Planning, refinement backends, physical validation, and auditable trajectories in one agent framework.

Architecture of AutoXRD and XRDBench
AutoXRD overview. Skill-guided planning, executable XRD tools, deterministic checks, and evidence-preserving refinement trajectories. Click the figure to open its vector PDF.
Benchmark

XRDBench evaluates reasoning and execution

Two complementary tracks test isolated scientific decisions and their composition into executable XRD workflows.

XRDBench-QA100

Diagnostic tasks

Refinement decisions, residual interpretation, parameter recovery, phase analysis, and refinement-history assessment.

  • 30 easy · 40 medium · 30 hard
  • Maximum 20 tool calls
  • Exact, objective, and scientific evaluation
XRDBench-E2E34

Executable workflows

State-changing analyses across eleven XRD task families with artifacts, execution gates, and physical validation.

  • Measured diffraction patterns
  • Maximum 50 agent steps
  • Artifact 15% · Metrics 55% · Judge 30%
XRDBench task composition across QA and E2E tracks
XRDBench composition. The two tracks contain 134 tasks spanning diagnostic capabilities and eleven executable XRD task families.
Evaluation

Large-scale evaluation of ten recent LLMs

1,340 model–task runs reveal a clear gap between bounded scientific reasoning and autonomous workflow execution.

Overall mean57.8

out of 100 across the two tracks

QA mean61.9

bounded diagnostic reasoning

E2E mean53.7

executable XRD workflows

Best overall81.1

GPT-5.6 Sol

Overall, QA, and E2E performance of ten language models
Overall performance. Scores combine XRDBench-QA and XRDBench-E2E with equal track weight. Click the figure to open its vector PDF.
Fine-grained analysis

Capability profiles expose different scientific bottlenecks

Fine-grained capability radar charts for QA and E2E tasks
Capability profiles. Models are relatively strong at refinement-history assessment and result acceptance, but remain weak at action selection, phase quantification, indexing, and Rietveld refinement.
Efficiency

Accuracy, execution time, tokens, steps, and cost

Efficiency and cost-effectiveness trade-offs across ten models
Efficiency and cost-effectiveness. Resource means pool all 134 tasks per model; GPT-5.6 Luna provides the strongest score–cost trade-off.
Ablation study

Which AutoXRD components matter?

All six evaluated components improve GPT-5.6 Luna on their targeted XRDBench-QA settings, with the largest gains from result-acceptance checks and final scientific review.

Score improvements from six AutoXRD components in GPT-5.6 Luna ablations
Component effects. Bars report the full-system score minus the score after removing each component. Click the figure to open its vector PDF.
Findings

Where current XRD agents still fail

01

Scientific outcome failures

Agents may execute a workflow successfully yet return inaccurate cells, HKL assignments, or structurally invalid refinements.

02

Execution and evidence failures

Missing code, invalid backend outputs, or incomplete artifacts prevent scoring and scientific verification.

03

Termination failures

Some trajectories consume nearly the full 50-step budget without producing a valid deliverable.

Citation

AutoXRD

BibTeX will be added when the public paper record is available.