Open benchmarking of generative AI for fast calorimeter simulation
SuperCode
AI-assisted optimisation of scientific software for sustainable computing
Open benchmarking of generative AI for fast calorimeter simulation
Vladimir Vava Gligorov
CERN; Austrian Academy of Sciences (Marietta Blau Institute for Particle Physics, MBI); INFN; University of Hamburg; Lawrence Berkeley National Laboratory; Rutgers University
Switzerland; Austria; Italy; Germany; France; United States
Astronomy & Physics
Research institute & Higher education institution; Research infrastructure
Open
High-resource
CaloChallenge establishes a common, open benchmark for evaluating generative AI models that approximate computationally expensive calorimeter simulations in high-energy physics. Detailed GEANT4 simulations model how particles such as electrons, photons and pions produce showers of energy inside particle detectors, but generating these events at the scale required by current and future experiments is a major computing bottleneck. CaloChallenge provides shared calorimeter-shower datasets of increasing dimensionality and asks generative models to learn the conditional distribution of energy deposits given the energy of the incoming particle. The challenge received 59 submissions across its four datasets, including variational autoencoders, generative adversarial networks, normalizing flows, diffusion models and conditional flow-matching approaches.. All models are evaluated against common GEANT4 reference samples using the same physics-sensitive and machine-learning-based metrics, together with measures of generation time and model size. The principal output is therefore a comparative evidence base for assessing when generative AI can provide fast simulation without unacceptable loss of physical fidelity.
The practice was organised as an inter-experimental community challenge so that research groups developing generative fast-simulation methods could compare them under shared conditions instead of using different datasets and evaluation procedures. The core organising team brought together researchers from CERN, the Marietta Blau Institute for Particle Physics (MBI), INFN, the University of Hamburg, Lawrence Berkeley National Laboratory and Rutgers University, with a much larger international community contributing models and evaluation results. A central enabling condition is the provision of openly accessible, standardised reference data. CaloChallenge released datasets ranging from relatively simple calorimeter geometries with hundreds of voxels to highly granular showers containing more than 40,000 voxels. Data are distributed through Zenodo with persistent identifiers and use a common HDF5 structure; detector geometry files, loading utilities and evaluation scripts are publicly available. Participants generate samples in the same format as the reference data, allowing the same evaluation pipeline to be applied across architectures. Quality is assessed through several complementary measures, including physical shower observables, histogram comparisons, classifier-based tests and distributional metrics, while computational characteristics such as generation time and memory requirements are also considered. This combination of shared data, common output formats and independent multi-metric evaluation makes direct methodological comparison possible.
CaloChallenge produced a large-scale comparative evaluation of contemporary generative approaches to calorimeter fast simulation. The final study analyses 59 submissions across four datasets: some groups submitted to a single dataset, others to all four, and some contributed several approaches or variants of them, such as distilled models. The study shows that model performance involves trade-offs between physical fidelity, generation speed and model complexity; no single architecture is uniformly superior across all criteria. The resulting study provides a common evidence base for assessing generative calorimeter fast simulation and contributes more broadly to the scientific evaluation of generative models. Trustworthiness is strengthened by requiring models to be evaluated against common GEANT4 references through multiple complementary metrics rather than relying on visual plausibility or a single aggregate score. Computational efficiency is also part of the comparison, allowing faster simulation to be considered together with scientific quality. Openness and reuse are particularly strong: the challenge datasets are deposited on Zenodo, evaluation code is public, submitted samples and model repositories are linked, and numerical results and notebooks used to reproduce published figures are shared. This allows subsequent researchers to benchmark new methods against the same reference results and reduces duplication in dataset preparation and evaluation. The challenge does not demonstrate a single quantified reduction in energy use for particle-physics experiments; its contribution is a transparent framework for identifying models that may reduce simulation cost while maintaining acceptable physical fidelity.
1. Main paper: Krause, C., Faucci Giannelli, M., Kasieczka, G., Nachman, B., Salamani, D., Shih, D., Zaborowska, A., Amram, O., Borras, K., Buckley, M.R. and Buhmann, E., 2025. CaloChallenge 2022: a community challenge for fast calorimeter simulation. Reports on Progress in Physics, 88(11), p.116201. https://iopscience.iop.org/article/10.1088/1361-6633/ae1304
2. Homepage: https://calochallenge.github.io/homepage/