Seminar: How Data Complexity Shapes Machine Learning Performance

→ Europe/London
Description

Does the structure of your data affect how well your models perform and how stable your feature selection is?

This seminar presents a case study in software defect prediction, exploring how structural characteristics of tabular datasets influence both algorithm performance and feature selection stability.

What to expect:

  • Data complexity metrics tested across multiple algorithms
  • How dimensionality, overlap and network structure affect predictive stability and effectiveness
  • Practical guidance on choosing models and feature selection strategies for your dataset
  • Steps towards more robust, interpretable defect prediction

How to join:

Sign up via the link below and you will receive a link to attend online.

Adam Featherstone
Registration
Seminar: How Data Complexity Shapes Machine Learning Performance
    • 13:00 → 14:00
      Welcome, talk and discussion

      Abstract: We explore the role of data complexity focusing on how structural characteristics of tabular datasets influence both machine learning algorithm performance and feature selection stability. Using a set of complexity metrics and multiple algorithms, this work reveals that certain data traits, such as dimensionality, overlap, and network structure, can significantly affect predictive stability and effectiveness of the classification algorithms. Knowing the complexity of the data can offer practical insights for selecting appropriate models and feature selection strategies based on dataset properties, contributing to more robust and interpretable defect prediction systems.

      Bio: Daniel Rodriguez is currently an associate professor at the Computer Science Department of the University of Alcala, Madrid, Spain. Previously, he was a lecturer at the University of Reading, UK. Daniel earned his degree in Computer Science at the University of the Basque Country (EHU) and PhD degree at the University of Reading, UK. His research interests include data mining and software engineering in general and the application of data and optimisation techniques to Software Engineering in particular.