An initiative of the European Commission

Past Predictive Modeling

Validated machine-learning reconstruction of incomplete historical economic time series

Details

Practice contact

Meredith M. Paker, Judy Z. Stephenson, Patrick Wallis

Organisation(s)

University of Oxford; University College London; London School of Economics and Political Science

Country

United Kingdom

Scientific Domain

Social Sciences & Humanities

Context

Institutional Type

Research institute & Higher education institution

Data Governance

Managed access

Resource Conditions

Moderate-resource

AI Use

Past Predictive Modeling (PPM) addresses a methodological problem in quantitative historical research: long-run economic series often contain substantial temporal and geographical gaps, while conventional regression approaches may impose a single parametric relationship across long periods or become unstable when observations are sparse. PPM uses supervised machine learning to reconstruct historical quantities from incomplete observations. In the main application, researchers use individual English wage observations from the thirteenth to twentieth centuries to estimate annual and regional wage levels. AI enters after historical observations have been assembled and prepared. Temporally ordered data are separated into training, validation and testing periods, followed by an expanding-window “walk-backward” procedure in which models trained on later observations predict progressively earlier years. Several predictive algorithms are evaluated, with LightGBM gradient-boosted decision trees selected on out-of-sample performance. The resulting estimates are aggregated into historical wage series used in analyses of inequality, regional development and productivity.

Enabling Conditions

The practice depends on detailed historical microdata containing both the outcome to be reconstructed and attributes informative for prediction. The English application draws on a large collection of individual wage observations spanning several centuries, with variables such as occupation, location, wage type, sex, season and payment conditions available as predictors. Uneven temporal coverage makes sufficient observations in the training period an important condition for reliable use of the method. The research combines economic-history knowledge with predictive-modeling expertise and uses a standard Python machine-learning stack. Implementation is organised around temporal sample separation: hyperparameters and model choice are determined using validation data, while an untouched testing period is reserved for out-of-sample evaluation. Gradient-boosted decision trees are compared with regularised linear models, neural networks and other ML approaches before LightGBM is selected according to prediction error. Bootstrap resampling provides uncertainty estimates and permutation feature importance is used to inspect predictor contributions. Public code, derived English wage series and a second application to Japanese servant-wage data support reuse. The Japanese case uses public source data; exact reproduction of the English application still requires access to source microdata that cannot be redistributed.

Outcomes

The evaluation shows measurable gains over the conventional historical-regression benchmark. In the revised English application, the LightGBM implementation reduces bootstrap standard errors by about 60% on average and improves out-of-sample predictive accuracy by more than 26%. A second application to Japanese agricultural-servant wages reports an approximately 40% reduction in average bootstrap standard errors, providing evidence that the procedure can transfer beyond the original dataset. The reconstructed estimates also contribute to substantive historical analysis: they support consistent long-run regional wage series and population-weighted national estimates, modify some conclusions about the development of the skill premium after the Black Death, and reproduce the principal productivity-growth break around 1600 when substituted into an existing productivity analysis while recovering additional evidence of earlier productivity change. Trustworthiness rests on methodological checks built into the workflow: held-out temporal testing, comparison with alternative models and linear regression, bootstrap uncertainty, feature-importance analysis, external robustness testing, and simulations. Code and derived outputs are public, and the Japanese application is directly reproducible from public data. Reproducibility remains partial for the principal English case because the underlying source microdata cannot be redistributed.

Sources

LSE Research Online record: https://researchonline.lse.ac.uk/id/eprint/128852/

Implementation code and derived series: https://github.com/merpaker/PastPredictiveModeling/

Japanese robustness-data source described in the repository: Kumon, Y. The Labor Intensive Path: Servant Wage Dataset, ICPSR. https://doi.org/10.3886/E147081V1

Similar Good Practices

  • Good Practices
  • AI Science Community
  • AI Research

SuperCode

AI-assisted optimisation of scientific software for sustainable computing

  • Good Practices
  • AI Science Community
  • AI Research

WorldCereal

Open, retrainable crop-mapping workflows with emerging foundation-model integration

  • Good Practices
  • AI Science Community
  • AI Research

YieldSAT

A multimodal benchmark dataset for field and subfield crop yield prediction

  • Good Practices
  • AI Science Community
  • AI Research

TESSERA / GeoTessera

Open reuse of geospatial foundation-model embeddings for Earth observation

  • Good Practices
  • AI Science Community
  • AI Research

Semantic workflows for atomistic simulations

Toward Knowledge-Based Workflows: A Semantic Approach to Atomistic Simulations for Mechanical and Thermodynamic Properties

  • Good Practices
  • AI Science Community
  • AI Research

Bonding Analysis Database and Machine Learning Framework

  • Good Practices
  • AI Science Community
  • AI Research

Ontology-Aligned Structuring and Reuse of Multimodal Materials Data and Workflows Toward Automatic Reproduction

  • Good Practices
  • AI Science Community
  • AI Research

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

polySCOUT

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

pyMarAI

Human-in-the-Loop Deep-Learning Toolchain for Tumor Spheroid Delineation

  • Good Practices
  • AI Science Community
  • AI Research

Segmentation of various organelles of microalgae in free-living cell and symbiotic forms in large 3D electron microscopy images

Atlas of microalgae in plankton symbioses revealed by 3D electron microscopy

  • Good Practices
  • AI Science Community
  • AI Research

Implicit neural image field for biological microscopy image compression