Reproducing a computational result requires knowing the parameters that produced it, and in materials science those parameters are usually reported as prose in a methods section. This practice addresses the reuse of density-functional theory data and computational workflows reported in published literature by transferring it into knowledge graphs. The pipeline works in four stages: a literature search narrowed by successive filters; retrieval of the relevant passages, by section headers where papers follow a conventional structure and by semantic similarity where they do not; extraction into structured records using prompt-engineered language models; and alignment of those records to established materials ontologies, extended where existing vocabularies fall short. The demonstration targets stacking fault energy calculations in magnesium and its alloys, chosen because these defects govern ductility and are computed by several competing protocols. The output is a set of ontology-aligned knowledge graph representations of literature-reported DFT workflows, covering 711 data points, accompanied by extraction pipeline scripts, model outputs, and computational workflows.
SuperCode
AI-assisted optimisation of scientific software for sustainable computing