Independent thesis Advanced level (degree of Master (Two Years)), 80 credits / 120 HE credits
Microbiome data analysis faces multiple challenges, including compositional constraints, high sparsity, and high dimensionality. In longitudinal studies, these challenges combine with temporal dependencies, further complicating the analytical process. Additionally, in clinical research, disease samples are typically far fewer than healthy controls, leading to severe class imbalance problems that reduce model recognition capabilities for minority classes. This study, using postpartum depression prediction as an application scenario, developed and evaluated a systematic framework aimed at addressing class imbalance in time-series microbiome data through oversampling techniques.
We conducted a systematic comparative analysis of microbiome data from the BASIC prospective study at Uppsala University Hospital, evaluating different zero-value replacement strategies, feature selection methods, data transformation techniques, and oversampling algorithms. Results showed that row-level zero-value replacement, feature selection based on early time point data and labels, centered log-ratio transformation after feature selection, and oversampling with conditional generative models collectively constituted the most effective data processing pathway. This framework increased the recognition rate of minority class samples from a baseline of near zero to over 0.60, significantly enhancing model performance on imbalanced datasets.
The research also revealed that data completeness is crucial for model performance; when missing data was introduced, predictive performance declined significantly even with the complete processing workflow applied. Furthermore, our results indicated that traditional deep learning generative models like conditional Generative Adversarial Networks and conditional Variational Autoencoders struggle to effectively learn distributions from small samples, while statistical models such as conditional Gaussian Mixture Models and conditional Dirichlet distributions perform better with limited samples.
This study provides a viable solution for addressing class imbalance in time-series microbiome data. The developed framework is not only applicable to postpartum depression prediction but can also be extended to other research areas involving time-series microbiome data, such as mental health, allergic diseases, and metabolic disorders, laying the foundation for the application of microbiome analysis in clinical prediction and early intervention.
2025. , p. 61