Logo: to the web site of Uppsala University

uu.sePublikasjoner fra Uppsala universitet
Endre søk
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Oversampling Longitudinal Compositional Data for Classification of Microbiome Samples
Uppsala universitet, Teknisk-naturvetenskapliga vetenskapsområdet, Matematisk-datavetenskapliga sektionen, Institutionen för informationsteknologi, Tillämpad beräkningsvetenskap.
2025 (engelsk)Independent thesis Advanced level (degree of Master (Two Years)), 80 poäng / 120 hpOppgave
Abstract [en]

Microbiome data analysis faces multiple challenges, including compositional constraints, high sparsity, and high dimensionality. In longitudinal studies, these challenges combine with temporal dependencies, further complicating the analytical process. Additionally, in clinical research, disease samples are typically far fewer than healthy controls, leading to severe class imbalance problems that reduce model recognition capabilities for minority classes. This study, using postpartum depression prediction as an application scenario, developed and evaluated a systematic framework aimed at addressing class imbalance in time-series microbiome data through oversampling techniques.

We conducted a systematic comparative analysis of microbiome data from the BASIC prospective study at Uppsala University Hospital, evaluating different zero-value replacement strategies, feature selection methods, data transformation techniques, and oversampling algorithms. Results showed that row-level zero-value replacement, feature selection based on early time point data and labels, centered log-ratio transformation after feature selection, and oversampling with conditional generative models collectively constituted the most effective data processing pathway. This framework increased the recognition rate of minority class samples from a baseline of near zero to over 0.60, significantly enhancing model performance on imbalanced datasets.

The research also revealed that data completeness is crucial for model performance; when missing data was introduced, predictive performance declined significantly even with the complete processing workflow applied. Furthermore, our results indicated that traditional deep learning generative models like conditional Generative Adversarial Networks and conditional Variational Autoencoders struggle to effectively learn distributions from small samples, while statistical models such as conditional Gaussian Mixture Models and conditional Dirichlet distributions perform better with limited samples.

This study provides a viable solution for addressing class imbalance in time-series microbiome data. The developed framework is not only applicable to postpartum depression prediction but can also be extended to other research areas involving time-series microbiome data, such as mental health, allergic diseases, and metabolic disorders, laying the foundation for the application of microbiome analysis in clinical prediction and early intervention.

sted, utgiver, år, opplag, sider
2025. , s. 61
Serie
IT ; mTBV 25 005
Serie
Master's Programme in Computational Scienc
HSV kategori
Identifikatorer
URN: urn:nbn:se:uu:diva-557502OAI: oai:DiVA.org:uu-557502DiVA, id: diva2:1973870
Utdanningsprogram
Master Programme in Computational Science
Presentation
2025-05-27, 09:15 (engelsk)
Veileder
Examiner
Tilgjengelig fra: 2025-06-23 Laget: 2025-06-20 Sist oppdatert: 2025-06-23bibliografisk kontrollert

Open Access i DiVA

fulltext(1299 kB)273 nedlastinger
Filinformasjon
Fil FULLTEXT01.pdfFilstørrelse 1299 kBChecksum SHA-512
ca830f62f3a8f109615a9dfa108fe7b20dadd5193b33e8843c73bfe9683383908056ad870d26aeb2edf705724fa14f45842959f8a6cf530b8dc19cba1e2e4128
Type fulltextMimetype application/pdf

Søk i DiVA

Av forfatter/redaktør
Linjing, Shen
Av organisasjonen

Søk utenfor DiVA

GoogleGoogle Scholar
Totalt: 273 nedlastinger
Antall nedlastinger er summen av alle nedlastinger av alle fulltekster. Det kan for eksempel være tidligere versjoner som er ikke lenger tilgjengelige

urn-nbn

Altmetric

urn-nbn
Totalt: 309 treff
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf