fig5
Figure 5. Evolution of LLM-driven knowledge extraction paradigms in materials science. The input “literature corpus” denotes corpus-scale processing (103-104 papers, yielding ~1.6 × 105 nodes from a decade of literature[62]). Across all paradigms, the LLM operates in either (i) prompted, in-context inference without weight updates; or (ii) parameter fine-tuning or domain-adaptive pretraining[8,69]. The progression from schema-based to schema-free to agentic systems reflects increasing autonomy and decreasing direct human control; in practice, these paradigms are hybridized through schemas, retrieval, tools, and verification loops. LLM: Large language model; MOF: metal-organic framework; APIs: application programming interfaces.






