An initiative of the European Commission

pyMarAI

Human-in-the-Loop Deep-Learning Toolchain for Tumor Spheroid Delineation

Details

Practice contact

Jens Maus

Organisation(s)

Helmholtz-Zentrum Dresden-Rossendorf (HZDR), Institute of Radiopharmaceutical Cancer Research

Country

Germany

Scientific Domain

Life Sciences – preclinical cancer research and biomedical image analysis

Context

Institutional Type

Research institutes & higher education institutions

Data Governance

Managed

Resource Conditions

Moderate-resource

AI Use

Tumor spheroid growth assays evaluate cancer treatments in three-dimensional cell cultures. Microscopic images acquired over time require spheroid delineation to quantify growth and treatment response. Treatment-induced debris can make conventional thresholding unreliable and manual correction time-consuming.

pyMarAI integrates an nnU-Net-based convolutional neural network into this image-analysis workflow. Its graphical interface organizes microscopy images, runs segmentation on institutional GPU resources, and displays the resulting two-dimensional masks for review. Researchers visually assess delineation quality, label results GOOD or BAD, and use external software for manual correction when needed. The reviewed masks provide the basis for measuring spheroid area and assessing growth and treatment response. Reviewed and corrected delineations can also be used as training data for future model retraining. Researchers remain responsible for quality control and scientific interpretation.

Enabling Conditions

The practice arose from an image-analysis bottleneck: treatment-induced debris made threshold-based tumor spheroid delineation and manual correction increasingly time-consuming. At HZDR, expertise in radiopharmaceutical cancer research, microscopy, quantitative image analysis, AI methods and research software engineering was combined with access to existing mid-cost institutional GPU resources.

Existing institutional experiments supplied microscopy images and threshold-based delineations that had been manually corrected when needed. After suitability screening, 38,090 of 58,380 images were used for model development and evaluation. Screening used a small sample of images from each plate. Complete experiments were assigned either to training or independent testing. Within cross-validation, all images of a given spheroid remained within the same fold.

Models were trained from scratch using nnU-Net and PyTorch on an institutional Ubuntu system equipped with four NVIDIA Tesla V100S GPUs. nnU-Net’s self-configuration reduced the need for manual network design and hyperparameter selection. A PyQt5 graphical interface makes image handling, inference and review accessible to laboratory staff without requiring them to operate nnU-Net from the command line or expertise in AI implementation. Manual correction using external software is supported, while a structured directory layout enables collection of additional data for later retraining.

Outcomes

Agreement with manual delineation was high, with median Dice coefficients of 0.979 (Q1–Q3: 0.961–0.991) in cross-validation and 0.974 (Q1–Q3: 0.955–0.983) in independent testing. In one evaluated experiment, dose-response curves derived from uncorrected network output, with errors retained, produced closely comparable treatment-effect estimates at two selected time points. The remaining segmentation differences therefore did not materially alter the assessed treatment response in that experiment.

For approximately 1,000 images, estimated delineation time fell from roughly ten hours over several days to about two hours, including review and correction. Inference takes approximately one second per image. Routine use since August 2025 demonstrates the transition from a validated AI method to an integrated laboratory tool.

AI output is not treated as automatically authoritative. Researchers inspect masks, record GOOD or BAD quality labels and correct unsuitable results. Approximately 7–8% of images in cross-validation and independent testing had Dice coefficients below 0.9; 0.5–0.9% showed no overlap with the manual ground truth. However, failure analysis identified object-selection mismatches, incomplete delineation and oversegmentation by AI. A close review also exposed a single incorrect manual reference annotation. On 7,103 additional images from another microscope and other cell lines at HZDR, only 22% of delineations required correction. This demonstrates transfer while setting a clear boundary on unattended use and outlines the requirement for retraining for a specific group of images and better generalization.

The open publication, documented GitHub code and DOI-identified model weights support discovery and reuse in line with FAIR objectives including Open licenses support adaptation and sharing. For SCIANCE and RAISE, pyMarAI illustrates how routine life-science AI adoption requires more than model performance:  it depends on curated data, suitable institutional compute, interdisciplinary skills, an accessible and easy to use interface and explicit human oversight. These conditions can inform other image-intensive research workflows even when their scientific objects and models differ.

Sources

Scientific publication: https://doi.org/10.1021/acsmeasuresciau.5c00172
Software and documentation: https://github.com/hzdr-MedImaging/pyMarAI
Versioned model artefact: https://doi.org/10.14278/rodare.4198
Helmholtz Imaging solution entry: https://connect.helmholtz-imaging.de/solution/135

Similar Good Practices

  • Good Practices
  • AI Science Community
  • AI Research

SuperCode

AI-assisted optimisation of scientific software for sustainable computing

  • Good Practices
  • AI Science Community
  • AI Research

WorldCereal

Open, retrainable crop-mapping workflows with emerging foundation-model integration

  • Good Practices
  • AI Science Community
  • AI Research

YieldSAT

A multimodal benchmark dataset for field and subfield crop yield prediction

  • Good Practices
  • AI Science Community
  • AI Research

TESSERA / GeoTessera

Open reuse of geospatial foundation-model embeddings for Earth observation

  • Good Practices
  • AI Science Community
  • AI Research

Semantic workflows for atomistic simulations

Toward Knowledge-Based Workflows: A Semantic Approach to Atomistic Simulations for Mechanical and Thermodynamic Properties

  • Good Practices
  • AI Science Community
  • AI Research

Bonding Analysis Database and Machine Learning Framework

  • Good Practices
  • AI Science Community
  • AI Research

Ontology-Aligned Structuring and Reuse of Multimodal Materials Data and Workflows Toward Automatic Reproduction

  • Good Practices
  • AI Science Community
  • AI Research

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

polySCOUT

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

pyMarAI

Human-in-the-Loop Deep-Learning Toolchain for Tumor Spheroid Delineation

  • Good Practices
  • AI Science Community
  • AI Research

Segmentation of various organelles of microalgae in free-living cell and symbiotic forms in large 3D electron microscopy images

Atlas of microalgae in plankton symbioses revealed by 3D electron microscopy

  • Good Practices
  • AI Science Community
  • AI Research

Implicit neural image field for biological microscopy image compression