fig2

When code does not run: reproducibility challenges in materials machine learning benchmarks

Figure 2. Operational workflow used in this study to assess practical reproducibility on MatBench. LLM: Large language model.