An initiative of the European Commission

Integrating public Bioimage Data and AI Models

Integration of public bioimage data with public AI models for bioimage analysis

Details

Practice contact

Teresa Zulueta-Coarasa, Matthew Hartley

Organisation(s)

BioImage Archive, EMBL-EBI

Country

UK

Scientific Domain

Life Sciences

Context

Institutional Type

Research institute & Higher education institution; Research infrastructure

AI Use

This practice connects data from two public repositories: AI models for image analysis published in the BioImage Model Zoo and biological image data held in EMBL-EBI’s BioImage Archive (BIA). Pre-trained models are applied to archived images from selected datasets. The models perform a range of tasks, including segmentation, molecule localisation prediction, and denoising.

When the BIA holds reference annotations for segmentation tasks, four standard metrics: precision, recall, intersection over union, and Dice, are automatically computed and stored alongside the predictions. Results are published on a browsable webpage and as a structured results table in an open repository.

The purpose is to evaluate model performance across different datasets: model developers gain evidence of how well their models generalise beyond their training data, while life scientists gain a basis for choosing models that are suitable for their own images.

Enabling Conditions

The feasibility of this work relies on both public resources. Models can be retrieved using the BioImage Model Zoo core library, while the BioImage Archive provides images in a cloud-ready format (OME-Zarr), allowing them to be streamed without downloading them locally. The pipeline is packaged as an integrator library and a benchmarking script. It can be run locally in a Conda environment, in Docker, or on a compute cluster using Singularity and Nextflow. This allows the same code to support both individual exploratory runs and large-scale batch processing.

A training notebook hosted on Google Colab further lowers the barrier to entry, allowing users to benchmark a model against archive data directly in a browser without any local installation. This is particularly useful for a service intended for life scientists as well as users with more technical expertise.

Reusing existing public models rather than training new ones keeps the computational requirements proportionate to the purpose of the project. For segmentation tasks, model performance is evaluated using standard metrics computed against reference annotations already deposited in the archive.

Outcomes

Benchmarking metrics are computed and stored automatically for every model and image combination. Batch runs can process many combinations at once, and the results are published as webpages on the BioImage Archive website.

Two views of the results are provided for different audiences. A model-centric view allows model developers to assess how well their models generalise across different image types. A data-centric view helps life scientists compare models and choose those that are most suitable for their own data. Publishing predictions together with their measured agreement against reference annotations documents model performance in a form that is open to public scrutiny, including cases where models perform poorly. The code, pipeline, and tutorial notebook are all openly available, allowing others to reproduce or extend the benchmarking.

Sources

The practice results are available here: https://beta.bioimagearchive.org/bioimage-archive/galleries/ai/models, https://beta.bioimagearchive.org/bioimage-archive/galleries/ai/models

The code used to run the models on BIA data is publicly available on GitHub: https://github.com/BioImage-Archive/bia-bmz-integration

We also provide a Google Colab notebook for users who want to learn how to run the models on BIA data themselves:, https://colab.research.google.com/github/BioImage-Archive/bia-training/blob/main/notebooks/BMZ_benchmarking_with_BIA_data.ipynb

Similar Good Practices

  • Good Practices
  • AI Science Community
  • AI Research

SuperCode

AI-assisted optimisation of scientific software for sustainable computing

  • Good Practices
  • AI Science Community
  • AI Research

WorldCereal

Open, retrainable crop-mapping workflows with emerging foundation-model integration

  • Good Practices
  • AI Science Community
  • AI Research

YieldSAT

A multimodal benchmark dataset for field and subfield crop yield prediction

  • Good Practices
  • AI Science Community
  • AI Research

TESSERA / GeoTessera

Open reuse of geospatial foundation-model embeddings for Earth observation

  • Good Practices
  • AI Science Community
  • AI Research

Semantic workflows for atomistic simulations

Toward Knowledge-Based Workflows: A Semantic Approach to Atomistic Simulations for Mechanical and Thermodynamic Properties

  • Good Practices
  • AI Science Community
  • AI Research

Bonding Analysis Database and Machine Learning Framework

  • Good Practices
  • AI Science Community
  • AI Research

Ontology-Aligned Structuring and Reuse of Multimodal Materials Data and Workflows Toward Automatic Reproduction

  • Good Practices
  • AI Science Community
  • AI Research

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

polySCOUT

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

pyMarAI

Human-in-the-Loop Deep-Learning Toolchain for Tumor Spheroid Delineation

  • Good Practices
  • AI Science Community
  • AI Research

Segmentation of various organelles of microalgae in free-living cell and symbiotic forms in large 3D electron microscopy images

Atlas of microalgae in plankton symbioses revealed by 3D electron microscopy

  • Good Practices
  • AI Science Community
  • AI Research

Implicit neural image field for biological microscopy image compression