Open reuse of geospatial foundation-model embeddings for Earth observation
SuperCode
AI-assisted optimisation of scientific software for sustainable computing
Open reuse of geospatial foundation-model embeddings for Earth observation
Madeline C. Lisaius, Srinivasan Keshav, Anil Madhavapeddy
CMS Collaboration; INFN; Scuola Normale Superiore; CERN
United Kingdom
Earth Sciences
Research institute & Higher education institution
Open
High-resource
TESSERA is a pixel-wise foundation model for Earth observation that converts a full year of Sentinel-1 radar and Sentinel-2 optical observations into compact 128-dimensional embeddings at 10 m resolution. It uses self-supervised learning based on Barlow Twins, sparse temporal sampling and regularisation so that the representation remains stable despite clouds, irregular revisit times and missing observations. The pre-trained AI model offers a paradigm shift for working in geospatial: instead of each research team processing raw satellite time series and training a large model from scratch, researchers can retrieve precomputed TESSERA embeddings through the GeoTessera libraries and Zarr protocol and train a comparatively small classifier or regression head using local labelled data. The embeddings have been evaluated for classification, segmentation and regression tasks including land-cover and crop mapping, canopy height and biomass-related applications. The key output is therefore a reusable, analysis-ready geospatial embedding product that can support multiple environmental research questions from the same underlying satellite observations, as well as an open model.
The practice depends on large-scale, multisensor Earth-observation data, centralised model training, and an open distribution layer that separates the expensive pre-training stage from downstream pre-computed embeddings, inference, and scientific use. TESSERA combines Sentinel-1 and Sentinel-2 time series and was trained with self-supervised methods that do not require task labels at pretraining. The project used substantial research-computing infrastructure, including UKRI-supported systems and the DAWN and Isambard AI supercomputers, while the resulting embeddings are compressed for use relative to input data. Global annual embeddings are released at 10 m resolution, with open model weights and source code. GeoTessera provides a documented open-source HTTP and Zarr interface for retrieving embeddings by place and year, and verifies downloaded files with end-to-end checksums to prevent corrupted data entering local analyses. The workflow also provides example notebooks, evaluation tooling and a user discussion forum. This architecture allows researchers without access to large GPU systems to work with pretrained representations on ordinary computing environments while retaining the option to reproduce or extend the underlying model when sufficient resources are available. Downstream task evaluation and application insights were developed through collaborations across domain experts in environmental sciences from the University of Cambridge and beyond.
Tessera demonstrates that reusable geospatial embeddings can maintain strong scientific performance while reducing the labelled data and task-specific computation required for downstream Earth-observation analysis. In the core evaluation, TESSERA closely matched or outperformed task-specific models and other foundation models across diverse classification, segmentation and regression tasks, often using only a small head. Subsequent applications have tested the same embeddings across environmental mapping problems, providing evidence that the representation transfers beyond a single benchmark. The main productivity gain is therefore a change in where computational effort is spent: costly representation learning is performed centrally, while downstream researchers can retrieve compact embeddings and focus their resources on local labels, validation and scientific interpretation. Openness and transferability are strong: model weights, code, annual global embeddings and the GeoTessera access library are public, and the library includes integrity checks and documented interfaces for reproducible retrieval. The practice is also explicitly resource-conscious: downstream reuse can avoid repeated processing of long satellite time series and repeated training of large models. The principal limitation is that users inherit the assumptions and biases of the pretrained representation, so task-specific validation remains necessary before scientific interpretation.
TESSERA project and GeoTessera access: https://geotessera.org/
Feng, Z., Atzberger, C., Jaffer, S., Knezevic, J., Sormunen, S., Young, R., Lisaius, M. C., Immitzer, M., Jackson, T., Ball, J., Coomes, D. A., Madhavapeddy, A., Blake, A., & Keshav, S. (2025). TESSERA: Temporal embeddings of surface spectra for Earth representation and analysis. arXiv: https://arxiv.org/abs/2506.20380
Feng, Z., Jaffer, S., Shokar, I., Knezevic, J., Ball, J., Sousa, P., Elvers, M., Lisaius, M., Atzberger, C., Young, R., Naik, A., Robinson, N., Coomes, D., Madhavapeddy, A., & Keshav, S. (2026). TESSERA v2: Scaling pixel-wise Earth foundation models: https://arxiv.org/abs/2607.03949
TESSERA code: https://github.com/ucam-eo/tessera
GeoTessera library: https://github.com/ucam-eo/geotessera (https://doi.org/10.5281/zenodo.22127732)