An initiative of the European Commission

Fully Open Meditron

Auditable and reproducible development of clinical decision-support LLMs

Details

Practice contact

Xavier Theimer-Lienhard

Organisation(s)

EPFL, Laboratory for Intelligent Global Health and Humanitarian Response Technologies (LiGHT)

Country

Switzerland

Scientific Domain

Life Sciences

Context

Institutional Type

Research institute & Higher education institution

Data Governance

Open

Resource Conditions

High-resource

AI Use

Fully Open Meditron is an end-to-end open pipeline for developing large language models intended for research on clinical decision support. It combines a clinician-audited training corpus, reproducible data construction and model-training procedures, and evaluation protocols designed around clinical use. The corpus brings together eight public medical question-answering datasets and extends them with synthetic exam-style questions, guideline-grounded questions derived from clinical practice guidelines, and clinical vignettes. AI is used both to generate parts of the training corpus and as the basis for the resulting MeditronFO models, it is also used to identify coverage gaps, and for evaluation of open-ended clinical answers. Before training, the corpus is systematically decontaminated against evaluation benchmarks, while synthetic generations with known answers are checked against the gold label. The pipeline is applied to five fully open base models, producing domain-adapted MeditronFO variants whose training data, construction procedures and evaluation workflow can be inspected and reproduced.

Enabling Conditions

The practice extends the lab’s earlier Meditron models, motivated by the lack of any fully open medical specialist model. Clinical input is incorporated during data construction and validation: synthetic extensions are clinician-vetted, the overall pipeline undergoes review by a physician panel, and the open-ended clinical evaluation protocol is calibrated against human raters. This human involvement is complemented by technical controls designed to make the training process auditable. The public framework separates and documents synthetic data generation, benchmark decontamination, model training, medical and general benchmark evaluation, HealthBench, and pairwise clinical evaluation. The pipeline builds on existing fully open foundation models rather than developing base models from scratch. Five such models are fine-tuned (Apertus-70B/8B-Instruct, OLMo-2-32B-SFT, EuroLLM-22B/9B-Instruct), with Gemma-3-27B-IT as an open-weight control; GPT-OSS-120B generates the synthetic data, Qwen3-32B labels the corpus to identify coverage gaps, and Qwen3-235B serves as the LLM judge for HealthBench and Auto-MOOVE. Openness is treated as a condition across the whole workflow: source-data provenance, generation procedures, model configurations, decontamination settings and evaluation scripts are released together. The models and data are explicitly positioned for research rather than approved clinical deployment. It draws on assets from prior and parallel work: expert-written vignettes and pairwise physician ratings from the MOOVE initiative, the decontamination pipeline developed for Apertus, and existing fully open foundation models. The work was carried out as project #27 of the Swiss AI Initiative, funded through an ETH Domain grant, with compute provided by the Swiss National Supercomputing Centre (CSCS) on the Alps infrastructure.

Outcomes

The pipeline shows that opening the full development process can substantially narrow the performance gap with closed-data medical models. All five MeditronFO variants are preferred over their fully open base models in LLM-judged pairwise clinical evaluation, and Apertus-70B-MeditronFO raises the aggregate medical benchmark score from 47.2% to 53.8%, a 6.6-point gain that sets a new fully open state of the art. Applied to an open-weight control, Gemma-3-27B-MeditronFO is preferred over MedGemma in 58.6% of LLM-judged comparisons and scores 58.0% versus 55.9% on HealthBench. MedGemma nonetheless still leads on the aggregate benchmark average (60.7% versus 58.6%), and the best fully open model narrows but does not close this gap. A principal outcome is therefore not only a set of improved medical LLMs, but a reproducible way of constructing and evaluating them. The released framework makes corpus construction, synthetic-data generation, decontamination, training and evaluation available for inspection and reuse. The Auto-MOOVE protocol also makes open-ended clinical evaluation scalable: once validated against existing ratings from 204 human raters, an LLM judge can compare models without new rounds of human annotation. Trustworthiness is supported through clinician review, benchmark decontamination and human-calibrated clinical evaluation.The models are not presented as approved clinical tools, and downstream use requires application-specific safety assessment. The emphasis on openness is grounded in evidence that opaque adaptation pipelines are vulnerable to data-poisoning attacks that survive safety evaluations, and that narrow-domain fine-tuning can cause broad misalignment. Benchmark decontamination remains syntactic rather than semantic, and openness, while enabling third-party auditing, also allows the recipe to be reproduced without equivalent audits. The main limitation is resource intensity: reproducing training of models up to 70B parameters requires substantial GPU infrastructure. Smaller models, however, gained the most from the recipe (Apertus-8B improved by 12.8 points), suggesting that useful domain adaptation does not depend on the largest scale.

Sources

Publication: https://doi.org/10.48550/arXiv.2605.16215

Models: https://huggingface.co/collections/EPFLiGHT/meditronfo

Data: https://huggingface.co/datasets/EPFLiGHT/fully-open-meditron

Github: https://github.com/EPFLiGHT/FullyOpenMeditron

Similar Good Practices

  • Good Practices
  • AI Science Community
  • AI Research

SuperCode

AI-assisted optimisation of scientific software for sustainable computing

  • Good Practices
  • AI Science Community
  • AI Research

WorldCereal

Open, retrainable crop-mapping workflows with emerging foundation-model integration

  • Good Practices
  • AI Science Community
  • AI Research

YieldSAT

A multimodal benchmark dataset for field and subfield crop yield prediction

  • Good Practices
  • AI Science Community
  • AI Research

TESSERA / GeoTessera

Open reuse of geospatial foundation-model embeddings for Earth observation

  • Good Practices
  • AI Science Community
  • AI Research

Semantic workflows for atomistic simulations

Toward Knowledge-Based Workflows: A Semantic Approach to Atomistic Simulations for Mechanical and Thermodynamic Properties

  • Good Practices
  • AI Science Community
  • AI Research

Bonding Analysis Database and Machine Learning Framework

  • Good Practices
  • AI Science Community
  • AI Research

Ontology-Aligned Structuring and Reuse of Multimodal Materials Data and Workflows Toward Automatic Reproduction

  • Good Practices
  • AI Science Community
  • AI Research

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

polySCOUT

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

pyMarAI

Human-in-the-Loop Deep-Learning Toolchain for Tumor Spheroid Delineation

  • Good Practices
  • AI Science Community
  • AI Research

Segmentation of various organelles of microalgae in free-living cell and symbiotic forms in large 3D electron microscopy images

Atlas of microalgae in plankton symbioses revealed by 3D electron microscopy

  • Good Practices
  • AI Science Community
  • AI Research

Implicit neural image field for biological microscopy image compression