An initiative of the European Commission

CaloChallenge

Open benchmarking of generative AI for fast calorimeter simulation

Details

Practice contact

Vladimir Vava Gligorov

Organisation(s)

CERN; Austrian Academy of Sciences (Marietta Blau Institute for Particle Physics, MBI); INFN; University of Hamburg; Lawrence Berkeley National Laboratory; Rutgers University

Country

Switzerland; Austria; Italy; Germany; France; United States

Scientific Domain

Astronomy & Physics

Context

Institutional Type

Research institute & Higher education institution; Research infrastructure

Data Governance

Open

Resource Conditions

High-resource

AI Use

CaloChallenge establishes a common, open benchmark for evaluating generative AI models that approximate computationally expensive calorimeter simulations in high-energy physics. Detailed GEANT4 simulations model how particles such as electrons, photons and pions produce showers of energy inside particle detectors, but generating these events at the scale required by current and future experiments is a major computing bottleneck. CaloChallenge provides shared calorimeter-shower datasets of increasing dimensionality and asks generative models to learn the conditional distribution of energy deposits given the energy of the incoming particle. The challenge received 59 submissions across its four datasets, including variational autoencoders, generative adversarial networks, normalizing flows, diffusion models and conditional flow-matching approaches.. All models are evaluated against common GEANT4 reference samples using the same physics-sensitive and machine-learning-based metrics, together with measures of generation time and model size. The principal output is therefore a comparative evidence base for assessing when generative AI can provide fast simulation without unacceptable loss of physical fidelity.

Enabling Conditions

The practice was organised as an inter-experimental community challenge so that research groups developing generative fast-simulation methods could compare them under shared conditions instead of using different datasets and evaluation procedures. The core organising team brought together researchers from CERN, the Marietta Blau Institute for Particle Physics (MBI), INFN, the University of Hamburg, Lawrence Berkeley National Laboratory and Rutgers University, with a much larger international community contributing models and evaluation results. A central enabling condition is the provision of openly accessible, standardised reference data. CaloChallenge released datasets ranging from relatively simple calorimeter geometries with hundreds of voxels to highly granular showers containing more than 40,000 voxels. Data are distributed through Zenodo with persistent identifiers and use a common HDF5 structure; detector geometry files, loading utilities and evaluation scripts are publicly available. Participants generate samples in the same format as the reference data, allowing the same evaluation pipeline to be applied across architectures. Quality is assessed through several complementary measures, including physical shower observables, histogram comparisons, classifier-based tests and distributional metrics, while computational characteristics such as generation time and memory requirements are also considered. This combination of shared data, common output formats and independent multi-metric evaluation makes direct methodological comparison possible.

Outcomes

CaloChallenge produced a large-scale comparative evaluation of contemporary generative approaches to calorimeter fast simulation. The final study analyses 59 submissions across four datasets: some groups submitted to a single dataset, others to all four, and some contributed several approaches or variants of them, such as distilled models. The study shows that model performance involves trade-offs between physical fidelity, generation speed and model complexity; no single architecture is uniformly superior across all criteria. The resulting study provides a common evidence base for assessing generative calorimeter fast simulation and contributes more broadly to the scientific evaluation of generative models. Trustworthiness is strengthened by requiring models to be evaluated against common GEANT4 references through multiple complementary metrics rather than relying on visual plausibility or a single aggregate score. Computational efficiency is also part of the comparison, allowing faster simulation to be considered together with scientific quality. Openness and reuse are particularly strong: the challenge datasets are deposited on Zenodo, evaluation code is public, submitted samples and model repositories are linked, and numerical results and notebooks used to reproduce published figures are shared. This allows subsequent researchers to benchmark new methods against the same reference results and reduces duplication in dataset preparation and evaluation. The challenge does not demonstrate a single quantified reduction in energy use for particle-physics experiments; its contribution is a transparent framework for identifying models that may reduce simulation cost while maintaining acceptable physical fidelity.

Sources

1. Main paper: Krause, C., Faucci Giannelli, M., Kasieczka, G., Nachman, B., Salamani, D., Shih, D., Zaborowska, A., Amram, O., Borras, K., Buckley, M.R. and Buhmann, E., 2025. CaloChallenge 2022: a community challenge for fast calorimeter simulation. Reports on Progress in Physics, 88(11), p.116201. https://iopscience.iop.org/article/10.1088/1361-6633/ae1304
2. Homepage: https://calochallenge.github.io/homepage/

Similar Good Practices

  • Good Practices
  • AI Science Community
  • AI Research

SuperCode

AI-assisted optimisation of scientific software for sustainable computing

  • Good Practices
  • AI Science Community
  • AI Research

WorldCereal

Open, retrainable crop-mapping workflows with emerging foundation-model integration

  • Good Practices
  • AI Science Community
  • AI Research

YieldSAT

A multimodal benchmark dataset for field and subfield crop yield prediction

  • Good Practices
  • AI Science Community
  • AI Research

TESSERA / GeoTessera

Open reuse of geospatial foundation-model embeddings for Earth observation

  • Good Practices
  • AI Science Community
  • AI Research

Semantic workflows for atomistic simulations

Toward Knowledge-Based Workflows: A Semantic Approach to Atomistic Simulations for Mechanical and Thermodynamic Properties

  • Good Practices
  • AI Science Community
  • AI Research

Bonding Analysis Database and Machine Learning Framework

  • Good Practices
  • AI Science Community
  • AI Research

Ontology-Aligned Structuring and Reuse of Multimodal Materials Data and Workflows Toward Automatic Reproduction

  • Good Practices
  • AI Science Community
  • AI Research

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

polySCOUT

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

pyMarAI

Human-in-the-Loop Deep-Learning Toolchain for Tumor Spheroid Delineation

  • Good Practices
  • AI Science Community
  • AI Research

Segmentation of various organelles of microalgae in free-living cell and symbiotic forms in large 3D electron microscopy images

Atlas of microalgae in plankton symbioses revealed by 3D electron microscopy

  • Good Practices
  • AI Science Community
  • AI Research

Implicit neural image field for biological microscopy image compression