An initiative of the European Commission

TESSERA / GeoTessera

Open reuse of geospatial foundation-model embeddings for Earth observation

Details

Practice contact

Madeline C. Lisaius, Srinivasan Keshav, Anil Madhavapeddy

Organisation(s)

CMS Collaboration; INFN; Scuola Normale Superiore; CERN

Country

United Kingdom

Scientific Domain

Earth Sciences

Context

Institutional Type

Research institute & Higher education institution

Data Governance

Open

Resource Conditions

High-resource

AI Use

TESSERA is a pixel-wise foundation model for Earth observation that converts a full year of Sentinel-1 radar and Sentinel-2 optical observations into compact 128-dimensional embeddings at 10 m resolution. It uses self-supervised learning based on Barlow Twins, sparse temporal sampling and regularisation so that the representation remains stable despite clouds, irregular revisit times and missing observations. The pre-trained AI model offers a paradigm shift for working in geospatial: instead of each research team processing raw satellite time series and training a large model from scratch, researchers can retrieve precomputed TESSERA embeddings through the GeoTessera libraries and Zarr protocol and train a comparatively small classifier or regression head using local labelled data. The embeddings have been evaluated for classification, segmentation and regression tasks including land-cover and crop mapping, canopy height and biomass-related applications. The key output is therefore a reusable, analysis-ready geospatial embedding product that can support multiple environmental research questions from the same underlying satellite observations, as well as an open model.

Enabling Conditions

The practice depends on large-scale, multisensor Earth-observation data, centralised model training, and an open distribution layer that separates the expensive pre-training stage from downstream pre-computed embeddings, inference, and scientific use. TESSERA combines Sentinel-1 and Sentinel-2 time series and was trained with self-supervised methods that do not require task labels at pretraining. The project used substantial research-computing infrastructure, including UKRI-supported systems and the DAWN and Isambard AI supercomputers, while the resulting embeddings are compressed for use relative to input data. Global annual embeddings are released at 10 m resolution, with open model weights and source code. GeoTessera provides a documented open-source HTTP and Zarr interface for retrieving embeddings by place and year, and verifies downloaded files with end-to-end checksums to prevent corrupted data entering local analyses. The workflow also provides example notebooks, evaluation tooling and a user discussion forum. This architecture allows researchers without access to large GPU systems to work with pretrained representations on ordinary computing environments while retaining the option to reproduce or extend the underlying model when sufficient resources are available. Downstream task evaluation and application insights were developed through collaborations across domain experts in environmental sciences from the University of Cambridge and beyond.

Outcomes

Tessera demonstrates that reusable geospatial embeddings can maintain strong scientific performance while reducing the labelled data and task-specific computation required for downstream Earth-observation analysis. In the core evaluation, TESSERA closely matched or outperformed task-specific models and other foundation models across diverse classification, segmentation and regression tasks, often using only a small head. Subsequent applications have tested the same embeddings across environmental mapping problems, providing evidence that the representation transfers beyond a single benchmark. The main productivity gain is therefore a change in where computational effort is spent: costly representation learning is performed centrally, while downstream researchers can retrieve compact embeddings and focus their resources on local labels, validation and scientific interpretation. Openness and transferability are strong: model weights, code, annual global embeddings and the GeoTessera access library are public, and the library includes integrity checks and documented interfaces for reproducible retrieval. The practice is also explicitly resource-conscious: downstream reuse can avoid repeated processing of long satellite time series and repeated training of large models. The principal limitation is that users inherit the assumptions and biases of the pretrained representation, so task-specific validation remains necessary before scientific interpretation.

Sources

TESSERA project and GeoTessera access: https://geotessera.org/
Feng, Z., Atzberger, C., Jaffer, S., Knezevic, J., Sormunen, S., Young, R., Lisaius, M. C., Immitzer, M., Jackson, T., Ball, J., Coomes, D. A., Madhavapeddy, A., Blake, A., & Keshav, S. (2025). TESSERA: Temporal embeddings of surface spectra for Earth representation and analysis. arXiv: https://arxiv.org/abs/2506.20380

Feng, Z., Jaffer, S., Shokar, I., Knezevic, J., Ball, J., Sousa, P., Elvers, M., Lisaius, M., Atzberger, C., Young, R., Naik, A., Robinson, N., Coomes, D., Madhavapeddy, A., & Keshav, S. (2026). TESSERA v2: Scaling pixel-wise Earth foundation models: https://arxiv.org/abs/2607.03949
TESSERA code: https://github.com/ucam-eo/tessera
GeoTessera library: https://github.com/ucam-eo/geotessera (https://doi.org/10.5281/zenodo.22127732)

Similar Good Practices

  • Good Practices
  • AI Science Community
  • AI Research

SuperCode

AI-assisted optimisation of scientific software for sustainable computing

  • Good Practices
  • AI Science Community
  • AI Research

WorldCereal

Open, retrainable crop-mapping workflows with emerging foundation-model integration

  • Good Practices
  • AI Science Community
  • AI Research

YieldSAT

A multimodal benchmark dataset for field and subfield crop yield prediction

  • Good Practices
  • AI Science Community
  • AI Research

TESSERA / GeoTessera

Open reuse of geospatial foundation-model embeddings for Earth observation

  • Good Practices
  • AI Science Community
  • AI Research

Semantic workflows for atomistic simulations

Toward Knowledge-Based Workflows: A Semantic Approach to Atomistic Simulations for Mechanical and Thermodynamic Properties

  • Good Practices
  • AI Science Community
  • AI Research

Bonding Analysis Database and Machine Learning Framework

  • Good Practices
  • AI Science Community
  • AI Research

Ontology-Aligned Structuring and Reuse of Multimodal Materials Data and Workflows Toward Automatic Reproduction

  • Good Practices
  • AI Science Community
  • AI Research

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

polySCOUT

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

pyMarAI

Human-in-the-Loop Deep-Learning Toolchain for Tumor Spheroid Delineation

  • Good Practices
  • AI Science Community
  • AI Research

Segmentation of various organelles of microalgae in free-living cell and symbiotic forms in large 3D electron microscopy images

Atlas of microalgae in plankton symbioses revealed by 3D electron microscopy

  • Good Practices
  • AI Science Community
  • AI Research

Implicit neural image field for biological microscopy image compression