PSDI & Royce Materials Data Summit

Europe/London
2B.020 (Nancy Rothwell Building, University of Manchester, Manchester, UK)

2B.020

Nancy Rothwell Building, University of Manchester, Manchester, UK

Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
Description

Date: 29th-31st July 2026 
Location: Nancy Rothwell Building, University of Manchester, Manchester, UK 
Organised by: The Henry Royce Institute & the Physical Sciences Data Infrastructure

Overview

The Henry Royce Institute and the Physical Sciences Data Infrastructure (PSDI) are pleased to announce the PSDI & Royce Materials Data Summit, which will be held from 29th to 31st July 2026 at the University of Manchester. This event aims to bring together professionals working on digitalisation in materials science, giving them the opportunity to share their work, learn about recent developments, and build collaborations. Topics in scope for the event include, but are not limited to, the following (in the context of materials): 

  • Data standards, metadata quality, ontologies, and semantic interoperability 

  • Digital research tools, including workflow automation frameworks and electronic laboratory notebooks 

  • Community databases and curated data collections 

  • Data-driven applications of AI  

  • Autonomous laboratories and digital twins 

 

We welcome participants from a wide range of backgrounds, e.g. experimentalists, computational scientists, industry, software engineers, data engineers, data stewards.  

 

Format 

The event will be split into two parts:

  1. The 29th July will be dedicated to hands‑on training sessions where participants will be given the chance to try bleeding-edge digital tools. This training day will run from 1100-1845.
  2. The 30th and 31st July will have the format of a traditional conference, featuring invited talks from leaders in the field and a poster exhibition. Participants attending this part of the event will be invited to submit abstracts for the poster exhibition. The conference will run from 0900-1800 on 30th July, and 0900-1730 on 31st July. Moreover, we hope to provide a optional conference dinner on the evening of the 30th July (1925-2145 to be confirmed), and tours of the host institution's laboratories during the conference days. 

 

Participants can register for one or both parts of the event (see below for link). At registration participants can specify their interest in, e.g. the conference dinner and lab tour. Note that this event will be free to attend, but requires registration. The exception is the optional conference dinner. We anticipate the cost of the dinner will be £29.70 for 2 courses and £35.20 for 3 courses, but this is not final yet.

Note that this event is in-person only for participants. 

A detailed timetable for the event (both parts) will be published at PSDI & Royce Materials Data Summit (29-31 July 2026): Timetable · STFC Indico.

 

Participate

 

Motivation 

This event builds on the success of the 2025 PSDI Materials Community Workshop, hosted at the Royce Institute, which attracted strong interest from researchers across the UK and Europe. That workshop became a catalyst for ongoing cross‑community dialogue on digital approaches in materials research. 

While numerous specialised meetings exist - for example, on computational simulation, data curation, AI, semantic interoperability, or experimental data analysis - these events typically focus on their own technical domain rather than on materials as a unifying theme. As a result, communication across the materials community remains fragmented, slowing the development of coherent, multi‑aspect digitalisation strategies.

 

Conversely, broad materials‑research conferences (such as the UK’s Materials Research Exchange) include digitalisation as one topic among many, but their generalist audiences limit the depth and continuity of discussions around digital technologies. 

 

‘PSDI & Royce Materials Data Summit’ aims to close this gap by offering a dedicated and inclusive forum for a holistic, multi‑faceted conversation about materials digitalisation. 

 

Further information 

For specific enquiries please contact: 

 

Surveys
Friday 31/07 Laboratory Tours
    • Registration Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Room 1 Tutorials 2A.034 (Nancy Rothwell Building)

      2A.034

      Nancy Rothwell Building

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
      • 1
        Data analysis workflows, reproducibility and sharing with RO-Crate in Materials Galaxy 2A.034

        2A.034

        This tutorial introduces the Research Crate Object (RO-Crate), Galaxy and reproducible workflows in Materials Science. We begin with RO-Crate and applications in research, narrowing the scope to Materials. Moving to the practical element of the tutorial, we demonstrate usage of RO-Crates by reproducible workflows on the Galaxy platform with Materials Galaxy where participants will create, run, share, and inspect workflows in Materials Galaxy.

        Instructors

        • eScience Lab, the University of Manchester
        • Materials Galaxy, Cardiff University; Science and Technology Facilities Council (STFC)

        Requirements

        To follow along with this session you will need to bring your own device, preferably a laptop.

        Reading and setup

        Reading

        1. The following presentation will provide an introduction to the Galaxy user interface and motivation for using Galaxy
        2. This material provides a brief introduction to working with RO-Crates in Galaxy, please note this material forms the basis of one of the practical parts of the tutorial

        Setup

        In advance to the workshop participants should:

        1. Register for an account on Materials Galaxy. The following tutorial can be used as a guide to signing up, however, please note that the Galaxy Platform you are signing up to is Materials Galaxy
        2. Download the required tutorial materials from Zenodo
        3. Verify they can access the tutorial pages for each section (the materials will be used to instruct during the tutorial): Importing and Running a workflow, Creating an RO-Crate, Importing an RO-Crate

        Timetable

        12:30 - 12:50: RO-Crate

        • Introduction to Research Object Crate (RO-Crate)
        • RO-Crate applications in Materials Science

        12:50 - 13:45: Materials Galaxy

        • Importing a workflow and dataset, running the workflow, and exporting the workflow run as an RO-Crate in Materials Galaxy
        • Retrieving and inspecting the RO-Crate in Materials Galaxy
        • Running the RO-Crate and comparing the results to demonstrate reproducibility

        13:45 - 14:00: Break

        14:00 - 14:45: Materials Galaxy

        • Providing guidelines on annotation and publishing to Zenodo/Invenio
          for long‑term preservation
        • Accessing and inspecting workflows in WorkflowHub
        • CDIF-4-XAS
      • 2
        Automated nanoparticle segmentation of TEM images and generation of synthetic databases with TEMPOS 2A.034

        2A.034

        Nancy Rothwell Building

        Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL

        Motivation

        Quantitative analysis of nanoparticle populations from transmission electron micrographs remains a bottleneck in materials characterisation: manual segmentation is slow, subjective, and poorly scalable to the large datasets demanded by modern synthesis workflows. TEMPOS (Transmission Electron Microscopy Particle Outline Segmentation) addresses this through an instance-segmentation pipeline built on Mask R-CNN, enabling automated detection and morphometric analysis of nanoparticles across a wide range of imaging conditions and material systems. A distinguishing feature of the package is its synthetic data engine, which generates physically plausible TEM micrographs with ground-truth annotations, providing an effective route to training data in domains where experimentally labelled images are scarce. This tutorial offers a hands-on introduction to both capabilities. Participants will install TEMPOS, run inference on reference and personal datasets, extract per-particle descriptors, equivalent diameter, aspect ratio, spatial distribution, and export results for downstream analysis. The second half of the session covers synthetic database generation: configuring substrate and particle parameters, sampling across imaging conditions, and using the resulting data to fine-tune the model for a target material system. No prior machine learning experience is assumed; the tutorial is aimed at electron microscopists who wish to automate and standardise nanoparticle analysis in their own workflows.

        Tutorial

        This hands-on tutorial introduces TEMPOS, an open-source platform for automated nanoparticle segmentation and morphometric analysis in TEM images. Participants will learn how to install and run TEMPOS, apply the pre-trained Mask R-CNN model to their own micrographs, interpret per-particle outputs (size, shape descriptors, spatial distributions), and fine-tune the model on a small custom dataset. The session combines short conceptual segments with guided live demonstrations and self-paced exercises, and is designed to leave attendees with a working installation, sample notebooks, and a clear path to integrating TEMPOS into their own analysis workflows.

        Learning outcomes

        By the end of the tutorial, participants will be able to:

        1. Explain, at a conceptual level, how instance-segmentation networks (Mask R-CNN) differ from semantic segmentation and classical thresholding, and why this matters for touching or overlapping nanoparticles.
        2. Install and configure TEMPOS in a reproducible Python environment (conda or container) and verify the installation against a reference dataset.
        3. Run inference on TEM micrographs using the pre-trained model and adjust key parameters (confidence threshold, NMS, tile size) for their own image conditions.
        4. Interpret TEMPOS output: extract per-particle masks, equivalent diameters, aspect ratios, and population-level distributions, and export results to standard formats (CSV, JSON, OME-TIFF).
        5. Annotate a small dataset and fine-tune the model on a domain-specific particle type, evaluating performance with appropriate metrics (mAP, IoU).
        6. Identify common failure modes (low contrast, beam damage artefacts, agglomeration) and apply targeted mitigations.
        7. Locate documentation, the GitHub repository, and community resources for ongoing use.

        Tutorial Structure

        Please note that breaks are included in this 3-hour tutorial

        Part 1 — Context and theory

        • Motivation: bottlenecks in manual nanoparticle analysis, reproducibility, the autonomous-TEM context.
        • Instance segmentation in a nutshell: bounding boxes, masks, Mask R-CNN architecture without the maths.
        • What TEMPOS is and what it isn't: scope, supported imaging modes, current limitations.

        Part 2 — Installation and first run

        • Walk-through of installation via pip/conda and the containerised release; verifying the environment.
        • Guided inference on a provided reference micrograph; reading the output structure.

        Part 3 — Hands-on inference and analysis

        • Participants run TEMPOS on a provided dataset (or, optionally, their own)
        • Parameter sweeps: confidence threshold, tile size, overlap; effects on recall and precision.
        • Extracting morphometrics, plotting size distributions, exporting for downstream analysis.

        Part 4 — Fine-tuning on custom data

        • Annotation workflow with a lightweight labelling tool; what makes a good training set.
        • Launching a fine-tuning run on a small pre-annotated dataset; monitoring training.
        • Evaluating the fine-tuned model and comparing against the baseline.

        Part 5 — Failure modes, integration, and Q&A

        • Common pitfalls and how to diagnose them; integration sketches (Jupyter, scripted pipelines, autonomous loops).
        • Resources, roadmap, and Q&A.
        Speaker: Andrew Stewart (University College London)
    • Room 2 Tutorials 2B.002 (Nancy Rothwell Building)

      2B.002

      Nancy Rothwell Building

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
      • 3
        Querying remote materials data sources

        Schedule TBA

      • 4
        How to use CrystaLLM-pi to predict crystal structures

        Full title: How to use CrystaLLM-pi to predict crystal structures with desired functional properties, or matching experimental data and adapt the model to your own data

        In this tutorial, we will cover how the user can load up and make predictions with pre-packaged deep learning models to generate materials structure with a target property or matching a measured experimental data signal on their own devices. Additionally, we will show how to format your own data in order to adapt the model so it can make structure predictions conditioned on a desired property.

        Learning outcomes

        • Increased understanding on the inside workings of a specialized transformer model that can generate crystal structures conditioned on functional properties
        • Gain the skills to load pre-trained models and generate materials with desired attributes on your own device
        • Understand how to format structure and property data so it can be used for training a transformer
        • Learn how to adapt a toy model to this dataset in order to understand how it could be applied to your data

        Audience

        Computational materials scientists

        Pre-requisite background

        GitHub usage to access and download software; Familiarity with Python and code execution in computational notebooks, e.g. Jupyter Notebook; Python environment basic understanding.

        Pre-requisite setup

        A Hugging-Face API key. A markdown file with instructions on how to do this is provided.

        Tutorial outline

        Background
        - Crystallographic Information File
        - Transformers/LLMs for materials generation
        - Conditioning on Functional Properties
        Basic usage
        - Loading and generating desired materials using pre-trained publicly available CrystaLLM models
        Fine-tuning
        - Make a data frame containing CIFs and properties of interest in correct ML training format
        - How to link up your dataset to CrystaLLM to fine-tune it yourself
        - How to choose the correct parameters to train your model
        - Train your model with a toy example
        - Load a state of the art trained model & generate structures with it

    • Room 3 Tutorials 2B.003 (Nancy Rothwell Building)

      2B.003

      Nancy Rothwell Building

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
      • 5
        From Data Silos to Linked Research Workflows: Hands-on with OpenSemanticLab and OO-LD

        This hands-on tutorial introduces OpenSemanticLab (OSL) as an integrated OpenSource platform for research data management, combining ELN, LIMS, and workflow support. Participants will learn how to model data using Object-Oriented Linked Data (OO-LD) and build simple linked data applications using Python and LLM integration.

        This tutorial is split in 3 parts.

        • Part 1: A single Platform as ELN, LIMS, Workflow-, Project- and Terminology Management & more - Using OpenSemanticLab (OSL) for Research Data Management
        • Part 2: A custom schema in 5 minutes - Adapting OSL to your need by understanding Object-Oriented Linked Data (OO-LD)
        • Part 3: A linked research data app in 50 lines of code - Development with oold-python and its LLM integration

        Audience

        Researchers, data stewards, and R&D professionals with an interest in research data management and digitalisation in materials science; basic programming experience is helpful but not strictly required.

        Pre-requisites

        No prior knowledge of semantic technologies is required.
        For part 3 basic familiarity with programming concepts (preferably Python) is beneficial

        Setup

        For part 3 a laptop with a modern web browser and Python (>=3.11, uv package manager) installed or access to a JupyterHub Notebook

    • Registration Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • 14:30
      Break
    • Registration & morning coffee 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Conference Opening 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Keynote talk: Keynote talk 1 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
      • 6
        Machine learning for materials science: A bittersweet lesson
        Speaker: Keith Butler
    • 11:00
      Break 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Invited talks: Thursday session 1 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
      • 7
        Introduction to the Physical Sciences Data Infrastructure
        Speaker: Brian Matthews
      • 8
        Introduction to segmenting TEM images via machine learning and synthetic data

        Supervised segmentation of nanoparticles in TEM images is held back by annotation: hand-labelling is slow, subjective, and hard to reproduce, the very opposite of FAIR. TEMPOS (Transmission Electron Microscopy Pipeline for Object Segmentation) inverts the problem. Rather than annotate experimental images, it generates physically informed synthetic micrographs whose ground truth is known by construction, trains a Mask R-CNN model purely on that data, and segments real, unseen images, validated on gold nanoparticles and Co₃O₄ nanocrystals The simulation parameters serve as exact, machine-readable metadata, and results are published in FAIR-compliant form using Datasette, alongside a community database for shared datasets.

        Beyond the method, this presentation reflects on the process. TEMPOS grew from a side project, begun around 2020–21, into a containerised and openly released tool over roughly five years, a candid case study in open research: what it takes to build reproducible, AI-ready datasets; why "software contribution" proved a truer framing than "novel method"; and how much of the work is social, sharing data across groups, engaging with the open-source community, and learning from the makers of the tools one depends upon.

        Speaker: Andrew Stewart (University College London)
      • 9
        Large Language Model Agents for Materials Research Workflows

        Large language models (LLMs) are emerging as a new interface between researchers, scientific data, and computational tools. In materials science, they offer opportunities to simplify access to complex workflows, accelerate data-driven research, and support inverse materials design. However, the reliability and scientific utility of LLMs depend critically on the availability of standardized workflows and well-structured research data.

        In this contribution, we present LangSim1, an LLM-based interface for materials simulation workflows built on the pyiron workflow framework2. Rather than generating simulation code directly, LLM agents interact with validated scientific workflows, enabling robust execution of simulations and automated analysis of materials properties. Beyond forward simulations, these workflows can be combined with statistical and machine-learning models to identify candidate materials that satisfy target property requirements.

        A key enabler for this approach is the standardization of workflows through the Python Workflow Definition (PWD)3, an interoperable workflow representation that supports workflow exchange between pyiron, jobflow, and AiiDA. By separating scientific intent from implementation details, PWD provides a structured interface between workflows, research data, and AI agents, improving reproducibility, interoperability, and reuse.

        Our results highlight how standardized workflows can serve as a foundation for AI-assisted materials research. By linking research data to the workflows that generated it and exposing these workflows through natural-language interfaces, LLM agents can help researchers access, combine, and automate computational tools while maintaining transparency and reproducibility. This provides a pathway towards integrating AI agents with broader materials research infrastructures, autonomous laboratories, and digital twins.

    • Poster lightning talks: Thursday lightning talks 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • 13:25
      Lunch 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Poster Session 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • 13:55
      Advanced Metals Processing Facilities Tour 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Invited talks: Thursday session 2 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
      • 10
        TBA
        Speaker: Christoph Eberl
      • 11
        A Whole New Lab: Exploring ELNs and Digital Transformation

        Electronic Lab Notebooks (ELNs) are becoming an increasingly important part of the digital research landscape. They have evolved from basic digital versions of paper notebooks to fully fledged systems that (if implemented properly) can help researchers improve collaboration, reproducibility, and FAIR (Findable, Accessible, Interoperable, and Reusable) data practices. Yet the introduction of any new technology doesn't guarantee a magical transformation. Successfully introducing ELNs requires more than technology alone; it involves embedding good data practices, supporting cultural change, and carefully integrating tools into everyday research workflows.Drawing on experiences from the University of Southampton on our ELN deployments, FAIR data initiatives, and our wider work on digital laboratories, this talk will explore both the challenges and opportunities of creating more connected research environments. From implementing ELNs in academic settings and supporting FAIR data capture, to emerging developments in smart laboratories, connected systems, automation, and voice-enabled technologies. Join me on a magic carpet ride to explore the potential of the modern digital laboratory.

    • 16:40
      Break 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Forum: Thursday forum 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Conference Dinner
    • Registration & morning coffee 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Welcome to day 2 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Keynote talk: Keynote 2 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
      • 12
        People are infrastructure: the strategic case for open simulation science.

        Scientists and policymakers have learned how to make long-term strategic investments in singular physical infrastructures. Particle accelerators, fusion reactors and space telescopes are conceived, built and operated through coordinated programmes spanning decades. Such investments are essential to scientific health, whether they are curiosity-driven or their societal returns lie far in the future.

        We have not been systematic in another class of research infrastructures: the shared computational capabilities on which much of modern science depends. I will focus on electronic-structure theory and simulation, whose scientific reach is unusually broad and measurable. In Nature’s 2014 census of the most-cited papers ever, 12 of the top 100 across all fields of science, medicine, and engineering were on electronic-structure simulations, including two of the top ten; in the 2025 update, one more paper moved in the top ten. Citation counts are not measures of societal value, but they provide striking evidence of the field’s pervasive and enabling role.

        Electronic-structure methods combine theoretical depth with societal impact. They empower materials discovery and characterisation across academia, national laboratories and industry, and provide data, physical priors and validation for AI-enabled science. Computational tools, data and simulation were central to the 2011 US Materials Genome Initiative and are embedded in the 2025 Genesis Mission.

        Their economic model differs fundamentally from that of a physical facility. Open software, curated data and shared workflows can be replicated worldwide at the flick of a switch, with great synergies for everyone. Yet creating, validating, maintaining and extending these capabilities requires sustained human expertise. The underfunded infrastructure is therefore not hardware, but people: the physics, software, and data scientists working on purposes and timescales broader than a conventional grant.

        The UK has developed notable exceptions, but worldwide support for the development of scientific software and the theories that underpin it, for the production and curation of data, and for verification-and-validation efforts remains fragmented, project-based and short-lived, despite its genuinely modest costs and truly exceptional leverage and impact. My simple and simply bewildering question is: why?

    • Invited talks: Friday session 1 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
      • 13
        (StC) Digital transformation at Royce
        Speaker: Stuart Kitney
    • 11:20
      Break 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Invited talks: Friday session 2 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
      • 14
        TBA
        Speaker: Robert Quarshie
      • 15
        TBA
        Speaker: Jesper Friis
    • Poster lightning talks: Friday lightning talks 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • 13:20
      Lunch 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Poster Session 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • 14:00
      Photon Science Institute Laboratory Tours 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Invited talks: Friday session 3 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
      • 16
        Multimodal AI for Materials Design

        Artificial intelligence (AI) is accelerating materials prediction and design by enabling efficient exploration of chemical and structural spaces, with particular promise for novel materials discovery. However, novelty in materials discovery encompasses chemical plausibility, structural distinctiveness, property relevance and experimental realisability, making AI-driven novelty claims difficult to substantiate. We introduce a materials property hierarchy, from intrinsic, composition-determined properties to extrinsic, processing-dependent performance, to clarify deployment constraints and distinguish structural, physical and deployment novelty. This framework motivates an evidence-based view of multimodal materials data spanning chemical composition, microstructure, processing, and testing and characterisation, showing that current evidence remains concentrated in composition and idealised structure while heterogeneous, under-represented and weakly integrated modalities limit support for physical and deployment novelty. It also highlights the limitations of benchmarks based mainly on computational labels and proxy novelty criteria. Community-wide standards for data collection, modality alignment and evidence synthesis are needed to support multimodal data construction, process-aware multimodal modelling, feasibility-first generative modelling and deployment-aware benchmarking, so that generative and multimodal AI can design experimentally realisable materials with defensible scientific and practical novelty.

      • 17
        ICAT: metadata cataloguing for scientific facilities

        ICAT is a flexible solution for managing scientific metadata and data from a wide variety of domains following the FAIR data principles. In addition to the core service providing relational models for scientific and administrative metadata, there are a number of additional components that extend the functionality to provide web-based user interfaces, DOI minting and landing pages, plugin-based data retrieval, and more.

        The open-source project is maintained by an international collaboration with members from the facilities and organisations running ICAT: STFC (supporting Diamond and ISIS), ESRF, HZB, ALBA, Sirius, and SESAME. The software has a proven history of operating at scale to meet the needs of these organisations, with the Diamond Light Source storing records for 6 billion files amounting to 90PB of data. It also allows facilities to build customisations on top of the common core functionality for their specific use cases, such as ESRF's Human Organ Atlas.

        This talk will provide an overview of the ICAT project and software components, current and future developments, and how it can be used to build curated, high quality metadata collections for scientific experiments.

        Speaker: Patrick Austin (STFC)
    • 16:00
      Break 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL
    • Forum: Friday forum & closing remarks 2B.020

      2B.020

      Nancy Rothwell Building, University of Manchester, Manchester, UK

      Nancy Rothwell Building, The University of Manchester, Oxford Road, Manchester, M13 9PL