An initiative of the European Commission

SoBigData

Responsible computational social mining as a shared infrastructure service

Details

Practice contact

Roberto Trasarti

Organisation(s)

National Research Council (CNR) — coordinator; University of Pisa; Scuola Normale Superiore; Sant’Anna School of Advanced Studies; and partner institutions across Italy, the Netherlands, Germany, Spain, Greece and the United Kingdom

Country

Italy

Scientific Domain

Social Sciences & Humanities

Context

Institutional Type

Research infrastructure

Data Governance

Managed access

Resource Conditions

High-resource

AI Use

Studying social phenomena through large datasets requires researchers to consider several factors simultaneously: data, analytical methods, computing capacity and appropriate privacy and data protection arrangements. If these are assembled individually for each project, it can create substantial barriers for researchers who are not data scientists. SoBigData treats this process as an infrastructure service. Its virtual research environments combine catalogued datasets and methods with computational services and environments, such as Jupyter, RStudio, and Galaxy. Researchers can import or select analytical methods, prototype analyses in notebooks, or construct and execute workflows using shared computing resources. The Social Mining Analytics Engine supports the repeated and scheduled execution of analyses and automatically records provenance information. Researchers retain control of the research design and interpretation, while the infrastructure provides an environment in which computational social science analyses can be developed, run, modified and repeated.

Enabling Conditions

The infrastructure was built across successive funding rounds of the European Union’s Horizon 2020 and Horizon Europe programmes and Italy’s National Recovery and Resilience Plan by a multidisciplinary consortium combining computer science, social science, legal and ethical expertise. It is currently maintained by five full-time senior researchers and technicians, supported by around ten further developers and other staff. Its design treats computational capacity and responsible data use as shared infrastructural concerns, and has been shaped by legislation, ethics approval, data governance and organisational policy. User demand has influenced not only individual components but the overall organisation of services, providing environments suited to users’ needs. All software, platforms, tools and repositories are open source. The core infrastructure currently operates with 14,772 CPU cores and federates with three external nodes, adding another 1,950 cores. The platform is equipped with 24 high-performance NVIDIA A100 GPUs, which are dedicated to accelerating machine learning, deep learning, and data-intensive computational tasks; it offers 3,572 TB of storage and 102 TB of RAM, with an additional 2.5 TB from the three federated nodes. Datasets come from institutional holdings and from data generated for the purpose using ad hoc data procurement tools.

A structured integration process turns research outputs into reusable infrastructure resources. Partners use short micro-projects, typically lasting one to six months, to develop or integrate a dataset or analytical method. Methods can enter the infrastructure at three levels: description in the common catalogue; integration into JupyterHub for live experiment prototyping; or integration into the Method Engine, where users can execute them through an interface on shared computing resources. To integrate datasets, models and tools, the infrastructure requires metadata schemas, submission requirements and model documentation from contributors. This allows methods developed by individual research groups to become discoverable and, at higher integration levels, executable by other users.

Collaboration is organised through thematic Research Spaces, while researchers are encouraged to form teams spanning several spaces. The Gateway and virtual research environments connect users, who can create shared working groups, contact each other and forward queries to the infrastructure team. Ethical and legal expertise is organised through the Board of Operational Ethics and Legality (BOEL), which provides guidance on legal, ethical and societal questions arising in SoBigData research, including methodological approaches, standards and policies. Training and documentation further support researchers in using the infrastructure.

Outcomes

SoBigData reduces the technical and organisational barriers to computational social science research by providing access to data, analytical methods and computing resources via shared infrastructure. Methods developed by individual research groups can be documented in the shared catalogue and, at a higher level of integration, made available for experiment prototyping or execution by other researchers. This promotes the reuse of existing computational work and reduces the need for each project to assemble its own analytical environment. Persistent identifiers, metadata, standard formats and implementation guidance are requested from contributors and provided to users, and SoBigData components can be used across other infrastructures.

The infrastructure also facilitates more transparent and reproducible computational workflows. The Social Mining Analytics Engine automatically records provenance information and supports repeated and scheduled execution. Integrated workflow environments allow analytical steps and their outputs to be combined and revisited.

Responsible use is supported by an ongoing ethical and legal service. Users are first requested to take a MOOC on ethics and legislation, and the ethics board evaluates requests to use the infrastructure. Researchers can seek guidance on methodological, legal, and societal issues, and cases requiring formal ethical review can be identified and referred accordingly. This embeds responsible data considerations into the infrastructure available to researchers, providing a model that can be transferred to other research infrastructures that combine AI methods, sensitive data, and shared computational services.

Sources

Official Website: http://www.sobigdata.eu

Research Infrastructure Catalogue & Workspace: https://sobigdata.d4science.org/

Similar Good Practices

  • Good Practices
  • AI Science Community
  • AI Research

SuperCode

AI-assisted optimisation of scientific software for sustainable computing

  • Good Practices
  • AI Science Community
  • AI Research

WorldCereal

Open, retrainable crop-mapping workflows with emerging foundation-model integration

  • Good Practices
  • AI Science Community
  • AI Research

YieldSAT

A multimodal benchmark dataset for field and subfield crop yield prediction

  • Good Practices
  • AI Science Community
  • AI Research

TESSERA / GeoTessera

Open reuse of geospatial foundation-model embeddings for Earth observation

  • Good Practices
  • AI Science Community
  • AI Research

Semantic workflows for atomistic simulations

Toward Knowledge-Based Workflows: A Semantic Approach to Atomistic Simulations for Mechanical and Thermodynamic Properties

  • Good Practices
  • AI Science Community
  • AI Research

Bonding Analysis Database and Machine Learning Framework

  • Good Practices
  • AI Science Community
  • AI Research

Ontology-Aligned Structuring and Reuse of Multimodal Materials Data and Workflows Toward Automatic Reproduction

  • Good Practices
  • AI Science Community
  • AI Research

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

polySCOUT

Deep learning-enhanced physical modelling for tape-casting slurry microstructures of solid oxide cell substrates

  • Good Practices
  • AI Science Community
  • AI Research

pyMarAI

Human-in-the-Loop Deep-Learning Toolchain for Tumor Spheroid Delineation

  • Good Practices
  • AI Science Community
  • AI Research

Segmentation of various organelles of microalgae in free-living cell and symbiotic forms in large 3D electron microscopy images

Atlas of microalgae in plankton symbioses revealed by 3D electron microscopy

  • Good Practices
  • AI Science Community
  • AI Research

Implicit neural image field for biological microscopy image compression