An initiative of the European Commission

Ontology-Aligned Structuring and Reuse of Multimodal Materials Data and Workflows Toward Automatic Reproduction

Details

Practice contact

Markus Stricker

Organisation(s)

Ruhr University Bochum, Federal Institute for Materials Research and Testing Berlin, Forschungszentrum Jülich, Max Planck Institute for Sustainable Materials Düsseldorf

Country

Germany

Scientific Domain

Materials Sciences

Context

Institutional Type

Research institute & Higher education institution

Data Governance

Non-applicable

AI Use

Reproducing a computational result requires knowing the parameters that produced it, and in materials science those parameters are usually reported as prose in a methods section. This practice addresses the reuse of density-functional theory data and computational workflows reported in published literature by transferring it into knowledge graphs. The pipeline works in four stages: a literature search narrowed by successive filters; retrieval of the relevant passages, by section headers where papers follow a conventional structure and by semantic similarity where they do not; extraction into structured records using prompt-engineered language models; and alignment of those records to established materials ontologies, extended where existing vocabularies fall short. The demonstration targets stacking fault energy calculations in magnesium and its alloys, chosen because these defects govern ductility and are computed by several competing protocols. The output is a set of ontology-aligned knowledge graph representations of literature-reported DFT workflows, covering 711 data points, accompanied by extraction pipeline scripts, model outputs, and computational workflows.

Enabling Conditions

Reliance on pre-trained language models accessed through prompt engineering removes the need for model training infrastructure; two models are used, one commercial and one open-weight, with sampling temperature set and reported for every call. Filtering confines AI to a targeted extraction task: an initial query returning 21,080 publications is narrowed through successive criteria to 23 studies, then to eight benchmark papers by manual screening. Earlier work on a broader corpus produced schema fragmentation severe enough to make unification impractical, so this narrowing is a condition of the pipeline working at all. Alignment with existing materials science ontologies provides the semantic scaffolding into which extracted values are placed, with extensions issued as versioned releases where established vocabularies fall short. Alignment is performed manually and kept conservative: only fields clearly defined and contributing to reproducibility are mapped. The repository is publicly shared under an open licence and contains the pipeline scripts, model outputs, ontology-aligned representations and computational workflows; the source articles cannot be redistributed. Reporting includes disclosure of LLM use at each stage and comparisons against manually curated reference data. Using the knowledge graph presumes familiarity with ontologies, SPARQL querying and DFT provenance concepts.

Outcomes

The practice produced ontology-aligned knowledge graphs representing 711 stacking-fault-energy data points and their associated density-functional-theory workflows from eight publications. These structured representations make results easier to discover and compare across papers and provide a foundation for future automated reproduction of reported calculations. The semi-automated pipeline can reduce the burden of fully manual curation while retaining human oversight.

Evaluation against manually curated reference data showed both the value and limitations of LLM-assisted extraction. GPT-4.1 achieved 91.9% precision, 90.4% recall and 98.4% coverage, whereas the open-weight Ministral-3-14B model achieved 85.5% precision and 81.6% coverage. Errors in values, compositions, slip planes and stacking-fault labels demonstrate that model outputs should not be treated as authoritative without validation.

Openness and transferability are supported through a public repository containing pipeline scripts, model outputs, ontology-aligned representations and computational workflows. Ontology extensions remain accessible through versioned persistent URLs. The documented prompts, model configurations and comparison with human-curated data make the role and performance of AI inspectable. However, copyright restrictions prevent redistribution of the source publications, and adaptation to other scientific domains would require new domain schemas, ontological mappings and validation datasets.

Sources

Sepideh Baghaee Ravari, Abril Azocar Guzman, Sarath Menon, Stefan Sandfeld, Tilmann Hickel, Markus Stricker, Ontology-Aligned Structuring and Reuse of Multimodal Materials Data and Workflows Toward Automatic Reproduction, Advanced Engineering Materials 2026, 0, e202600008. https://doi.org/10.1002/adem.202600008

https://gitlab.ruhr-uni-bochum.de/icams-mids/ontostruct

Similar Good Practices

  • Good Practices
  • AI Science Community
  • AI Research

SuperCode

AI-assisted optimisation of scientific software for sustainable computing

  • Good Practices
  • AI Science Community
  • AI Research

WorldCereal

Open, retrainable crop-mapping workflows with emerging foundation-model integration

  • Good Practices
  • AI Science Community
  • AI Research

YieldSAT

A multimodal benchmark dataset for field and subfield crop yield prediction

  • Good Practices
  • AI Science Community
  • AI Research

TESSERA / GeoTessera

Open reuse of geospatial foundation-model embeddings for Earth observation

  • Good Practices
  • AI Science Community
  • AI Research

Semantic workflows for atomistic simulations

Toward Knowledge-Based Workflows: A Semantic Approach to Atomistic Simulations for Mechanical and Thermodynamic Properties

  • Good Practices
  • AI Science Community
  • AI Research

Bonding Analysis Database and Machine Learning Framework

  • Good Practices
  • AI Science Community
  • AI Research

Ontology-Aligned Structuring and Reuse of Multimodal Materials Data and Workflows Toward Automatic Reproduction

  • Good Practices
  • AI Science Community
  • AI Research

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

polySCOUT

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

pyMarAI

Human-in-the-Loop Deep-Learning Toolchain for Tumor Spheroid Delineation

  • Good Practices
  • AI Science Community
  • AI Research

Segmentation of various organelles of microalgae in free-living cell and symbiotic forms in large 3D electron microscopy images

Atlas of microalgae in plankton symbioses revealed by 3D electron microscopy

  • Good Practices
  • AI Science Community
  • AI Research

Implicit neural image field for biological microscopy image compression