Skip to main navigation Skip to search Skip to main content

Protocol to benchmark and evaluate the status of imbalance measure using correlation, data complexity, and ablation analyses

Research output: Contribution to journalArticleAcademicpeer-review

Abstract

Machine learning struggles with imbalanced data. Although several mitigation approaches exist, their application depends on the extent of imbalance. To determine the latter, a protocol was developed. Across 428 synthetic and 70 real datasets, 8 imbalance measures were benchmarked and evaluated using multiple classifiers, metrics, and correlation coefficients. The coding environment, data preparation, and correlation and complexity analyses are described. These are complemented by procedures for an ablation study of the most efficient measure: SIMBA (status of imbalance). For complete details on the use and execution of this protocol, please refer to Pivin-Bachler et al.

Original languageEnglish
Article number104500
Number of pages21
JournalSTAR Protocols
Volume7
Issue number2
DOIs
Publication statusPublished - 19 Jun 2026

Bibliographical note

Publisher Copyright:
© 2026 The Author(s).

Funding

The authors thank the Honda Research Institute in Japan for funding this research.

Keywords

  • Artificial Intelligence (AI)
  • Evaluation
  • bioinformatics
  • complexity
  • computer science
  • data science
  • health sciences
  • machine learning
  • statistics

Fingerprint

Dive into the research topics of 'Protocol to benchmark and evaluate the status of imbalance measure using correlation, data complexity, and ablation analyses'. Together they form a unique fingerprint.

Cite this