A multimodal benchmark dataset for field and subfield crop yield prediction
SuperCode
AI-assisted optimisation of scientific software for sustainable computing
A multimodal benchmark dataset for field and subfield crop yield prediction
Miro Miranda; Deepak Pathak
RPTU Kaiserslautern-Landau; German Research Center for Artificial Intelligence (DFKI); Vision Impulse GmbH; University of Groningen
Germany; Netherlands
Earth Sciences
Research institute & Higher education institution; Industry & Innovation
Managed access
High-resource
The ability to predict crop yield using satellite imagery is limited by the scarcity of field-level ground truth. Yield measurements are expensive to collect, inconsistent in quality, and frequently subject to access restrictions, resulting in most existing datasets being limited to individual crops or regions. YieldSAT provides a multimodal benchmark for machine-learning research at field and subfield levels. Yield measurements recorded by GPS-equipped combine harvesters are cleaned and corrected for moisture, then rasterised onto the 10 m Sentinel-2 grid. This process produces approximately 12.2 million labelled pixels from 2,173 fields. Each field is linked to satellite observations and weather, soil and topographic variables over the growing season. The data cover four crops in four countries between 2016 and 2024. The benchmark supports several AI methods from the computer vision domain that can process multimodal and multispectral temporal data, from simple linear regression to LSTM, Transformer and advanced fusion models. YieldSAT supports image regression and multimodal learning experiments, providing common evaluation settings and baseline results against which new crop yield prediction methods can be assessed. For the publication, the models were run on the institute’s infrastructure; the published data are intended to be used on each user’s own infrastructure.
YieldSAT was developed by researchers from RPTU Kaiserslautern-Landau, DFKI, Vision Impulse, and the University of Groningen, prompted by limits in existing data, curation and compute, and by the lack of powerful prediction models: traditional methods operating on physical knowledge lack accuracy and scalability, and the global approach was new. The project was partly funded via the AI4EO Solution Factory of the ESA InCubed Programme, with industry collaboration and support. The interdisciplinary team of engineers, data scientists, AI engineers, agricultural scientists and Earth-observation scientists provided the broad competences required, with external training partly involved. Access to combine-harvester records enabled ground truth to be established at a spatial resolution that is rarely available for crop yield modelling.
Considerable work went into ensuring the quality of these measurements. Agricultural experts manually inspected all 2,173 fields and assigned quality levels using documented guidelines. The raw measurements then pass through a specified sequence of transformations, including the removal of zero and biologically infeasible yields, the filtering of statistical outliers, the correction of wet yields to dry yields, and the spatial aggregation onto the satellite grid. The resulting raster data retains information about the number and variability of the measurements that contribute to individual pixels.
Sources of weather, soil and topography were selected according to their relevance to crop growth, public availability, global coverage and spatial resolution. YieldSAT is distributed in two forms: a training-ready version with 24 temporal observations and predefined multimodal fusion, and a flexible version that preserves the modalities at their original temporal, spatial, and spectral resolutions. Models, both built from scratch and foundation models, were trained on institutional GPU clusters using open-source deep-learning libraries such as PyTorch. The design follows community standards and partly the EU AI Act, and the asset provides its own standards, representation and models. Using it requires AI and domain expertise; tutorials, documentation and example code hosted on GitHub help bridge potential gaps, while the dataset is maintained on internal infrastructure and accessed on request.
YieldSAT establishes a common benchmark for high-resolution crop-yield prediction using directly measured harvester yields, supporting research in Earth science, climate science, agriculture, multimodal learning and computer vision. Its coverage spans four crops, four countries, nine growing seasons and approximately 12.2 million labelled 10 m pixels. As the first resource of its kind, it has attracted high interest and demand from academia and industry since its launch. Baseline experiments across a range of AI models, from expensive to very cheap options, provide reference results for subsequent studies, including results for models that did not perform well, while the two released data representations support both standardised comparison and research on alternative multimodal architectures.
The benchmark also quantifies how performance changes outside the conditions represented in training data. In the Argentina soybean experiments highlighted by the authors, holding out an entire year reduced explained variance by 22 percentage points relative to standard cross-validation; holding out a geographical region reduced it by 8 points. Experiments with ensembles and spatial inductive biases show that these losses can be reduced, making robustness to distribution shift a measurable research problem within the benchmark.
The resource also makes the quality of its ground truth visible. Field-level quality assessments and published examples distinguish stronger and weaker yield maps instead of presenting all observations as equivalent. This allows researchers to account for measurement quality when developing and comparing models. Uncertainty of model predictions is not part of the publication but has been addressed in related studies.
The benchmark is publicly documented, with tutorials and example code available through the project website. Access to the underlying dataset is managed because the field-level agricultural records are subject to data-sharing restrictions: since geolocations may be used to retrace natural persons, a licence that respects personal rights permits data access only for scientific purposes. Training and evaluation consumed substantial compute, justified by the impact and transparency of the study, and compute metrics were shared with the publishing institution.
Miranda, M. et al. (2026). YieldSAT: A Multimodal Benchmark Dataset for High-Resolution Crop Yield Prediction. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026.
Project and dataset access: https://yieldsat.github.io/
Code and tutorials: https://yieldsat.github.io/