Description
The reproducibility, interoperability, and long-term utility of molecular dynamics (MD) simulations are frequently hindered by heterogeneous output formats, disparate metadata vocabularies, and the lack of automated, user-friendly deposition pipelines. To address these challenges, we present an integrated, open-source suite of data tools within BioSimDR (Biomolecular Simulation Data Resources) comprising: biosim-schema, biosim-extractor, and biosimdb-interface, each designed to streamline the curation, standardisation, and publishing of biomolecular MD datasets.
At the core of this ecosystem is biosim-schema, a community-driven, LinkML-based semantic framework that defines terms for molecular composition, force fields, and simulation settings while mapping directly to terms used in MD software. Leveraging this schema, biosim-extractor automatically parses outputs from major MD engines (including GROMACS and Amber), normalises physical units, extracts key chemical descriptors (e.g., SMILES, InChIKeys, sequences), and integrates with the AiiDA framework to preserve provenance. To facilitate user deposition, biosimdb-interface provides a containerised web GUI that allows researchers to validate and seamlessly publish curated datasets directly to the Invenio-based BioSimDB data-collection hosted by PSDI (Physical Sciences Data Infrastructure).
Crucially, the decoupled, schema-first architecture of this toolset serves as a highly generalisable blueprint for other computational disciplines. By simply substituting the biomolecular terms within the schema for domain-specific definitions, such as crystal lattices, density functional theory (DFT) settings, or mechanical properties, the downstream extractor and interface tools can be readily adapted to fields like materials chemistry and solid-state physics. This work demonstrates a scalable pathway toward establishing FAIR (Findable, Accessible, Interoperable, and Reusable) data infrastructures across diverse scientific communities.