Automatic Information Extraction System for MACro-organisms
SuperCode
AI-assisted optimisation of scientific software for sustainable computing
Automatic Information Extraction System for MACro-organisms
Perrine Paul-Gilloteaux
CNRS, France BioImaging, ANERIS project
France
Life Sciences
Research institute & Higher education institution; Research infrastructure
Open
High
Underwater photographs from fixed cameras and from divers with smartphones now arrive faster than anyone can identify what is in them, and the images are hard: crowded sea-floor backgrounds, blur, poor visibility. AIES-MAC processes a single image in stages. It first computes measures of the image’s own quality — sharpness, contrast, colourfulness — so that later steps can be adapted to the material. A neural network then draws boxes around organisms and proposes a species name for each. A second network converts each box into an outline of the organism. From that outline the system computes descriptive measurements: size and shape, the contour’s geometry, and the texture of the animal’s surface. These measurements, not the network’s internal reasoning, are the intended product — the system is built to add quantities a biologist can read and check to observation databases, rather than to deliver species labels on trust.
The system is developed by a French research infrastructure and imaging group within a European marine-observation consortium, drawing images and test cases from partners running citizen-science platforms, fixed underwater cameras and intertidal survey work. Existing published models are used, sourced from marine research institutes and community model repositories, and the documentation tables each one with the region and species it was trained on and a link to its weights. When no model was available publicly, like In the case of citizen-science platform, a specific model was trained and will be made available publicly once validated. The accompanying warning is explicit: each is tied to a fixed species list and a particular kind of image, and will fail on species or imagery outside it. Outlining is handled by a general-purpose segmentation model, steered by the boxes the detector produced, which removes the need for anyone to trace organisms by hand for training. The choice to derive characteristics by classical measurement rather than learned representation is deliberate, so that what enters a biodiversity database is a set of quantities with a stated mathematical definition.
The work was developped on a research cluster provided by the European grid, with a high-end graphics card; it was also deployed on the final infrastructure driving the fixed underwater camera respecting real-time constraints of acquisition. The code is released publicly as documented notebooks and a python library with a persistent identifier, following the project’s data management plan.
The practice has produced a released workflow. Its first version is public, documented step by step, with a persistent identifier, and the models it depends on are named and traceable to the institutes and repositories that trained them, each with its geographic and species coverage recorded. What the workflow delivers into a biodiversity database is a table of defined measurements, one row per detected organism, alongside the proposed species name. Because those measurements have stated mathematical definitions, they can be inspected, plotted and disputed by a biologist who did not build the system: worked examples show them separating several fish species and, applied to a camera site over one day, registering a genuine drop in visibility. The same quantities are intended to flag specimens sitting apart from others carrying the same label, which offers a route to catching misclassifications and to noticing species missing from a model’s reference list.
https://aneris.eu/technology/AIES-MAC