Large-scale data foundation for AI in Earth observation
SuperCode
AI-assisted optimisation of scientific software for sustainable computing
Large-scale data foundation for AI in Earth observation
Gencer Sumbul
Technische Universität Berlin; German Research Center for Artificial Intelligence (DFKI); BIFOLD
Germany
Earth Sciences
Research institute & Higher education institution
Open
Moderate-resource
BigEarthNet is a large-scale, openly reusable benchmark dataset designed specifically for developing and evaluating AI methods for Earth observation. The original release contains 590,326 non-overlapping Sentinel-2 image patches from 125 tiles acquired across ten European countries. Each patch retains the multispectral structure of Sentinel-2 and is assigned one or more land-cover labels derived from CORINE Land Cover 2018, reflecting the fact that satellite scenes usually contain several land-cover types at once. AI is not used to generate the labels; instead, the practice provides domain-specific data for training and testing machine-learning and deep-learning models. To demonstrate its value, the authors trained AI models from scratch on BigEarthNet and compared them with models pretrained on natural images. The key output is therefore a common, scientifically grounded benchmark that supports comparable AI research on multispectral remote-sensing imagery. The original release is later extended with Sentinel-1 image patches for multi-modal AI development and testing, a new nomenclature of land-cover classes aligned with the properties of remote sensing imagery, and pixel-level reference maps.
The practice was enabled by combining openly available Copernicus Sentinel-2 imagery with the established CORINE Land Cover 2018 inventory. The team selected 125 low-cloud Sentinel-2 tiles acquired between June 2017 and May 2018, applied atmospheric correction, excluded the band not containing surface information, and divided the imagery into patches at the native 10 m, 20 m, and 60 m band resolutions. CORINE classes were then associated with each patch as multi-label annotations. This design addressed two limitations of earlier remote-sensing benchmarks: their small size and their use of single RGB labels that do not represent the multispectral, multi-label character of satellite imagery. Quality control was explicit. Visual inspection identified 70,987 patches fully covered by seasonal snow, cloud, or cloud shadow, and the authors published lists of these patches and recommended excluding them for relevant training and testing tasks. The dataset and documentation were made publicly available so that research groups could train and compare models on a shared reference resource rather than construct separate local datasets. For further enrichment of BigEarthNet after initial release, each Sentinel-2 patch is paired with a Sentinel-1 patch generated from publicly available 325 Sentinel-1 Ground Range Detected (GRD) products with close temporal proximity to the original Sentinel-2 tiles. In addition, to reflect the updates in the CORINE inventory and to align the annotations with the characteristics of remote sensing imagery, a new class nomenclature of 19 classes with updated annotations was provided.
BigEarthNet established a reusable and large-scale data foundation for AI research in Earth observation and demonstrated the value of domain-specific benchmark data. In the original study, a shallow CNN trained from scratch on all available Sentinel-2 spectral bands achieved an F1 score of 0.7098, compared with 0.4988 for the ImageNet-pretrained Inception-v2 baseline; the reported improvements were statistically significant. The result showed that models trained on large, multispectral remote-sensing data can outperform transfer from conventional computer-vision benchmarks whose image characteristics differ substantially from satellite observations. In the follow-up study, the same conclusion was also validated for multi-modal multi-label classification. Openness and reuse are central outcomes: BigEarthNet provides a common dataset for training, evaluation, and comparison of methods, reducing duplication in data preparation and improving comparability across studies. The dataset has subsequently been refined in later releases, including new data modality, updated labels and data splits, showing that the benchmark is maintained as a research resource rather than remaining a one-off archive. The main limitation is that the original labels from the CORINE inventory inherit uncertainties and class imbalance; these limitations are documented and addressed through quality-control flags and later revisions.
1. Original publication: https://ieeexplore.ieee.org/document/8900532
2. BigEarthNet project and downloads: https://bigearth.net/
3. BigEarthNet v2.0 record: https://zenodo.org/records/10891137
4. Follow-up publications: https://ieeexplore.ieee.org/document/11242834 , https://ieeexplore.ieee.org/document/9552024