Predictive models on sensitive social data
SuperCode
AI-assisted optimisation of scientific software for sustainable computing
Predictive models on sensitive social data
Elizaveta Sivak, Gert Stulp
University of Groningen; ODISSEI; Eyra; Centerdata
Netherlands
Social Sciences & Humanities
Research institute & Higher education institution
Restricted / sensitive
High-resource
PreFer data challenge tested how well machine-learning and statistical models can predict whether people aged 18–45 would have a child within the next three years. Teams worked on a common target using longitudinal LISS survey data and population-scale Dutch register data. The design allowed theory-driven and data-driven approaches to be compared on the same unseen outcomes. For the downloadable LISS track, participants submitted the trained model, preprocessing and training code, and a method description through an open-source submission platform. The platform first checked that the workflow executed on example data and then evaluated it on a protected holdout set. For the register-data tracks, modelling took place inside the secure Statistics Netherlands (CBS) Remote Access environment. The key output was therefore not only a set of fertility predictions, but a reproducible benchmark of what different modelling strategies and data sources can predict under common evaluation conditions.
The challenge brought together demography, survey methodology, data science and software engineering. Its design depended on two complementary data resources: a LISS dataset with thousands of survey variables, including subjective measures such as fertility intentions, and Dutch administrative registers covering life-course information for the full population. LISS respondents could be linked to CBS records inside the secure environment; the reported linkage succeeded for about 90% of LISS participants for whom linkage was attempted.
Evaluation was separated from model development. Training and holdout data were split at household level to prevent leakage between related individuals, and the register-data holdout was divided into intermediate and final leaderboard sets to limit overfitting to repeated submissions. Reproducibility was built into submission: teams provided trained models, preprocessing and training scripts, and methodological documentation. ODISSEI covered CBS data-access costs for selected teams, while CBS vetting and disclosure rules governed access and release of code from the restricted environment.
More than 150 people participated and submitted over 70 models, ranging from traditional machine-learning approaches to newer computational methods. Despite the scale and richness of the data, predictive performance remained modest. Survey-based models performed slightly better than register-based models, and combining the strongest register prediction with the strongest survey model produced only a small improvement. The result challenges the idea that weak prediction of social outcomes is mainly a consequence of small social-science samples: even full-population registers and extensive longitudinal survey data left individual fertility highly uncertain.
PreFer also produced a reusable model-evaluation framework. Submitted workflows included model-training and preprocessing code, and project scripts were prepared for release where data-governance conditions allowed. The practice shows how predictive AI can be evaluated with protected holdout data, executable workflows and explicit privacy controls, so that comparisons between models contribute evidence about the scientific problem itself. The same design can be adapted to other sensitive social-data prediction tasks.
Sivak, E. et al. (2024). Combining the strengths of Dutch survey and register data in a data challenge to predict fertility (PreFer). Journal of Computational Social Science, 7, 1403–1431. https://doi.org/10.1007/s42001-024-00275-6
Sivak, E. & Stulp, G. (2026). Are births predictable with linked survey and register data? Evidence from the Predicting Fertility data challenge. International Journal of Population Data Science, 11(5), 3486. https://doi.org/10.23889/ijpds.v11i5.3486