Responsible computational social mining as a shared infrastructure service
SuperCode
AI-assisted optimisation of scientific software for sustainable computing
Responsible computational social mining as a shared infrastructure service
Roberto Trasarti
National Research Council (CNR) — coordinator; University of Pisa; Scuola Normale Superiore; Sant’Anna School of Advanced Studies; and partner institutions across Italy, the Netherlands, Germany, Spain, Greece and the United Kingdom
Italy
Social Sciences & Humanities
Research infrastructure
Managed access
High-resource
Studying social phenomena through large datasets requires researchers to consider several factors simultaneously: data, analytical methods, computing capacity and appropriate privacy and data protection arrangements. If these are assembled individually for each project, it can create substantial barriers for researchers who are not data scientists. SoBigData treats this process as an infrastructure service. Its virtual research environments combine catalogued datasets and methods with computational services and environments, such as Jupyter, RStudio, and Galaxy. Researchers can import or select analytical methods, prototype analyses in notebooks, or construct and execute workflows using shared computing resources. The Social Mining Analytics Engine supports the repeated and scheduled execution of analyses and automatically records provenance information. Researchers retain control of the research design and interpretation, while the infrastructure provides an environment in which computational social science analyses can be developed, run, modified and repeated.
The infrastructure was built across successive funding rounds of the European Union’s Horizon 2020 and Horizon Europe programmes and Italy’s National Recovery and Resilience Plan by a multidisciplinary consortium combining computer science, social science, legal and ethical expertise. It is currently maintained by five full-time senior researchers and technicians, supported by around ten further developers and other staff. Its design treats computational capacity and responsible data use as shared infrastructural concerns, and has been shaped by legislation, ethics approval, data governance and organisational policy. User demand has influenced not only individual components but the overall organisation of services, providing environments suited to users’ needs. All software, platforms, tools and repositories are open source. The core infrastructure currently operates with 14,772 CPU cores and federates with three external nodes, adding another 1,950 cores. The platform is equipped with 24 high-performance NVIDIA A100 GPUs, which are dedicated to accelerating machine learning, deep learning, and data-intensive computational tasks; it offers 3,572 TB of storage and 102 TB of RAM, with an additional 2.5 TB from the three federated nodes. Datasets come from institutional holdings and from data generated for the purpose using ad hoc data procurement tools.
A structured integration process turns research outputs into reusable infrastructure resources. Partners use short micro-projects, typically lasting one to six months, to develop or integrate a dataset or analytical method. Methods can enter the infrastructure at three levels: description in the common catalogue; integration into JupyterHub for live experiment prototyping; or integration into the Method Engine, where users can execute them through an interface on shared computing resources. To integrate datasets, models and tools, the infrastructure requires metadata schemas, submission requirements and model documentation from contributors. This allows methods developed by individual research groups to become discoverable and, at higher integration levels, executable by other users.
Collaboration is organised through thematic Research Spaces, while researchers are encouraged to form teams spanning several spaces. The Gateway and virtual research environments connect users, who can create shared working groups, contact each other and forward queries to the infrastructure team. Ethical and legal expertise is organised through the Board of Operational Ethics and Legality (BOEL), which provides guidance on legal, ethical and societal questions arising in SoBigData research, including methodological approaches, standards and policies. Training and documentation further support researchers in using the infrastructure.
SoBigData reduces the technical and organisational barriers to computational social science research by providing access to data, analytical methods and computing resources via shared infrastructure. Methods developed by individual research groups can be documented in the shared catalogue and, at a higher level of integration, made available for experiment prototyping or execution by other researchers. This promotes the reuse of existing computational work and reduces the need for each project to assemble its own analytical environment. Persistent identifiers, metadata, standard formats and implementation guidance are requested from contributors and provided to users, and SoBigData components can be used across other infrastructures.
The infrastructure also facilitates more transparent and reproducible computational workflows. The Social Mining Analytics Engine automatically records provenance information and supports repeated and scheduled execution. Integrated workflow environments allow analytical steps and their outputs to be combined and revisited.
Responsible use is supported by an ongoing ethical and legal service. Users are first requested to take a MOOC on ethics and legislation, and the ethics board evaluates requests to use the infrastructure. Researchers can seek guidance on methodological, legal, and societal issues, and cases requiring formal ethical review can be identified and referred accordingly. This embeds responsible data considerations into the infrastructure available to researchers, providing a model that can be transferred to other research infrastructures that combine AI methods, sensitive data, and shared computational services.
Official Website: http://www.sobigdata.eu
Research Infrastructure Catalogue & Workspace: https://sobigdata.d4science.org/