Logo: to the web site of Uppsala University

uu.sePublications from Uppsala University
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Oversampling Longitudinal Compositional Data for Classification of Microbiome Samples
Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Computational Science.
2025 (English)Independent thesis Advanced level (degree of Master (Two Years)), 80 credits / 120 HE creditsStudent thesis
Abstract [en]

Microbiome data analysis faces multiple challenges, including compositional constraints, high sparsity, and high dimensionality. In longitudinal studies, these challenges combine with temporal dependencies, further complicating the analytical process. Additionally, in clinical research, disease samples are typically far fewer than healthy controls, leading to severe class imbalance problems that reduce model recognition capabilities for minority classes. This study, using postpartum depression prediction as an application scenario, developed and evaluated a systematic framework aimed at addressing class imbalance in time-series microbiome data through oversampling techniques.

We conducted a systematic comparative analysis of microbiome data from the BASIC prospective study at Uppsala University Hospital, evaluating different zero-value replacement strategies, feature selection methods, data transformation techniques, and oversampling algorithms. Results showed that row-level zero-value replacement, feature selection based on early time point data and labels, centered log-ratio transformation after feature selection, and oversampling with conditional generative models collectively constituted the most effective data processing pathway. This framework increased the recognition rate of minority class samples from a baseline of near zero to over 0.60, significantly enhancing model performance on imbalanced datasets.

The research also revealed that data completeness is crucial for model performance; when missing data was introduced, predictive performance declined significantly even with the complete processing workflow applied. Furthermore, our results indicated that traditional deep learning generative models like conditional Generative Adversarial Networks and conditional Variational Autoencoders struggle to effectively learn distributions from small samples, while statistical models such as conditional Gaussian Mixture Models and conditional Dirichlet distributions perform better with limited samples.

This study provides a viable solution for addressing class imbalance in time-series microbiome data. The developed framework is not only applicable to postpartum depression prediction but can also be extended to other research areas involving time-series microbiome data, such as mental health, allergic diseases, and metabolic disorders, laying the foundation for the application of microbiome analysis in clinical prediction and early intervention.

Place, publisher, year, edition, pages
2025. , p. 61
Series
IT ; mTBV 25 005
Series
Master's Programme in Computational Scienc
National Category
Microbiology in the Medical Area
Identifiers
URN: urn:nbn:se:uu:diva-557502OAI: oai:DiVA.org:uu-557502DiVA, id: diva2:1973870
Educational program
Master Programme in Computational Science
Presentation
2025-05-27, 09:15 (English)
Supervisors
Examiners
Available from: 2025-06-23 Created: 2025-06-20 Last updated: 2025-06-23Bibliographically approved

Open Access in DiVA

fulltext(1299 kB)270 downloads
File information
File name FULLTEXT01.pdfFile size 1299 kBChecksum SHA-512
ca830f62f3a8f109615a9dfa108fe7b20dadd5193b33e8843c73bfe9683383908056ad870d26aeb2edf705724fa14f45842959f8a6cf530b8dc19cba1e2e4128
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Linjing, Shen
By organisation
Computational Science
Microbiology in the Medical Area

Search outside of DiVA

GoogleGoogle Scholar
Total: 270 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 294 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf