Human-in-the-Loop Deep-Learning Toolchain for Tumor Spheroid Delineation
SuperCode
AI-assisted optimisation of scientific software for sustainable computing
Human-in-the-Loop Deep-Learning Toolchain for Tumor Spheroid Delineation
Jens Maus
Helmholtz-Zentrum Dresden-Rossendorf (HZDR), Institute of Radiopharmaceutical Cancer Research
Germany
Life Sciences – preclinical cancer research and biomedical image analysis
Research institutes & higher education institutions
Managed
Moderate-resource
Tumor spheroid growth assays evaluate cancer treatments in three-dimensional cell cultures. Microscopic images acquired over time require spheroid delineation to quantify growth and treatment response. Treatment-induced debris can make conventional thresholding unreliable and manual correction time-consuming.
pyMarAI integrates an nnU-Net-based convolutional neural network into this image-analysis workflow. Its graphical interface organizes microscopy images, runs segmentation on institutional GPU resources, and displays the resulting two-dimensional masks for review. Researchers visually assess delineation quality, label results GOOD or BAD, and use external software for manual correction when needed. The reviewed masks provide the basis for measuring spheroid area and assessing growth and treatment response. Reviewed and corrected delineations can also be used as training data for future model retraining. Researchers remain responsible for quality control and scientific interpretation.
The practice arose from an image-analysis bottleneck: treatment-induced debris made threshold-based tumor spheroid delineation and manual correction increasingly time-consuming. At HZDR, expertise in radiopharmaceutical cancer research, microscopy, quantitative image analysis, AI methods and research software engineering was combined with access to existing mid-cost institutional GPU resources.
Existing institutional experiments supplied microscopy images and threshold-based delineations that had been manually corrected when needed. After suitability screening, 38,090 of 58,380 images were used for model development and evaluation. Screening used a small sample of images from each plate. Complete experiments were assigned either to training or independent testing. Within cross-validation, all images of a given spheroid remained within the same fold.
Models were trained from scratch using nnU-Net and PyTorch on an institutional Ubuntu system equipped with four NVIDIA Tesla V100S GPUs. nnU-Net’s self-configuration reduced the need for manual network design and hyperparameter selection. A PyQt5 graphical interface makes image handling, inference and review accessible to laboratory staff without requiring them to operate nnU-Net from the command line or expertise in AI implementation. Manual correction using external software is supported, while a structured directory layout enables collection of additional data for later retraining.
Agreement with manual delineation was high, with median Dice coefficients of 0.979 (Q1–Q3: 0.961–0.991) in cross-validation and 0.974 (Q1–Q3: 0.955–0.983) in independent testing. In one evaluated experiment, dose-response curves derived from uncorrected network output, with errors retained, produced closely comparable treatment-effect estimates at two selected time points. The remaining segmentation differences therefore did not materially alter the assessed treatment response in that experiment.
For approximately 1,000 images, estimated delineation time fell from roughly ten hours over several days to about two hours, including review and correction. Inference takes approximately one second per image. Routine use since August 2025 demonstrates the transition from a validated AI method to an integrated laboratory tool.
AI output is not treated as automatically authoritative. Researchers inspect masks, record GOOD or BAD quality labels and correct unsuitable results. Approximately 7–8% of images in cross-validation and independent testing had Dice coefficients below 0.9; 0.5–0.9% showed no overlap with the manual ground truth. However, failure analysis identified object-selection mismatches, incomplete delineation and oversegmentation by AI. A close review also exposed a single incorrect manual reference annotation. On 7,103 additional images from another microscope and other cell lines at HZDR, only 22% of delineations required correction. This demonstrates transfer while setting a clear boundary on unattended use and outlines the requirement for retraining for a specific group of images and better generalization.
The open publication, documented GitHub code and DOI-identified model weights support discovery and reuse in line with FAIR objectives including Open licenses support adaptation and sharing. For SCIANCE and RAISE, pyMarAI illustrates how routine life-science AI adoption requires more than model performance: it depends on curated data, suitable institutional compute, interdisciplinary skills, an accessible and easy to use interface and explicit human oversight. These conditions can inform other image-intensive research workflows even when their scientific objects and models differ.
Scientific publication: https://doi.org/10.1021/acsmeasuresciau.5c00172
Software and documentation: https://github.com/hzdr-MedImaging/pyMarAI
Versioned model artefact: https://doi.org/10.14278/rodare.4198
Helmholtz Imaging solution entry: https://connect.helmholtz-imaging.de/solution/135