fig3

Computational materials agents: from task demonstrations to executable scientific workflows

Figure 3. Common evaluation forms for computational materials agents. Local component checks assess formal properties of individual artifacts or execution steps. An end-to-end case study places these checks within the complete record of a selected agent run, including the task and inputs, software setup, actions and decisions, execution record, and interpreted results. Structured benchmarks evaluate performance across a predefined task set through repeated runs conducted under a common protocol and environment, with scoring based on metrics, rubrics, or reference results. The figure was assembled by the authors using Microsoft PowerPoint. Icons were adapted from Lucide Icons and used under the ISC License and, where applicable, the MIT License for Feather-derived icons.