Enhancing particle physics analysis reproducibility with agentic systems
SuperCode
AI-assisted optimisation of scientific software for sustainable computing
Enhancing particle physics analysis reproducibility with agentic systems
Caterina Doglioni
University of Manchester, University College London
United Kingdom
Astronomy & Physics
Research infrastructure
Open
Moderate-resource
Particle-physics measurements remain a public asset long after the experiment that produced them, but only if their analysis definitions survive in a form other researchers and models can reuse. The community’s standard for doing this is Rivet: executable code that encodes exactly how a measurement was defined. As of today, only 40% of published LHC measurements include this code, meaning that most measurements cannot easily be compared against new theories. AgentRivet is a multi-agent AI system exploiting large language models that produces these missing routines directly from the published paper. An Analyst agent reads the paper and extracts the physics definitions; a Coder agent turns them into working C++ code; dedicated review agents then check both the software and the physics before release. The output is a ready-to-use Rivet routine, which turns a publication into a resource that the whole field can build on.
AgentRivet was started as a response to a known gap in the high energy physics community: although there are agreed standards and frameworks (HEPData and Rivet) for preserving a measurement, routine coverage has stayed near 39% for years. This is mainly limited by the time and recognition that preservation work receives, rather than by interest in the topic. The project team pairs expertise in particle physics at the University of Manchester with expertise in research software engineering at University College London. The team includes active developers and users of HEPData and Rivet with expertise in ML and generative AI methods.
The system was designed around what a working analysis routine actually needs to contain, rather than around the AI method in isolation. It builds directly on Rivet and HEPData, the extraction task a well-defined target. Currently, AgentRivet uses pre-trained commercial language models from three independent providers (OpenAI, Google, Anthropic) accessed through their public APIs, avoiding dependence on any single provider’s service. It splits the work across specialised agents, with structured, validated data passed between them to keeps software and physics checks separate. Both reasoning and outputs are auditable and cached, avoiding re-computation of expensive intermediate results. The resulting code, and the full record of how it was produced, is released openly on GitLab and via pypy, and it can be ran on any machine as it includes a virtual environment that fetches the necessary packages.
AgentRivet’s first tests, on recent ATLAS and CMS measurements with no existing Rivet routine, show that current AI models can read a physics paper and turn it into working analysis code, at a cost of roughly USD 1.20–2.20 per routine. The generated code compiled with few errors, but its physics content varied by model, by observable, and even between repeat runs of the same model; the review stages caught and corrected a substantial share of both software and physics mistakes before outputting an analysis. Where problems did get through, most traced back to ambiguities in how the original paper described the measurement. This drives the workflow towards needing a human-in-the-loop for a judgment call, adding a preserved audit trail that makes it possible to find such decision points.
The code is released publicly on GitLab, so other groups can reuse, inspect, or adapt the workflow. Currently, the outputs are intended to be checked by a domain expert before formal release, and a new version is in preparation adding further automated steps to close this gap. In terms of frugality, an LLM-based approach was judged proportionate because the task can’t be handled by simpler rule-based extraction. Within that choice we tested whether smaller, cheaper models were used wherever adequate rather than defaulting to the most capable, most expensive option throughout. Future work will move towards smaller, fine-tuned open-weights models that can run locally, with the workflow’s energy use and carbon footprint measured directly rather than inferred from cost.
Together, these first results show that expanding analysis preservation with AI is both technically workable and affordable, with review and traceability that keep it trustworthy treated as part of the method itself.
Code: https://gitlab.com/hepcedar/AgentRivet
Paper: Costa, A.J., Doglioni, C., Gütschow, C., Pilkington, A.D. and Sinha, S., 2026. AgentRivet: an automated system for producing Rivet routines from journal publications. arXiv preprint arXiv:2606.13535. https://arxiv.org/abs/2606.13535