Logotyp: till Uppsala universitets webbplats

uu.sePublikationer från Uppsala universitet
Ändra sökning
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Two Optimization Problems in Genetics: Multi-dimensional QTL Analysis and Haplotype Inference
Uppsala universitet, Teknisk-naturvetenskapliga vetenskapsområdet, Matematisk-datavetenskapliga sektionen, Institutionen för informationsteknologi, Avdelningen för beräkningsvetenskap. Uppsala universitet, Teknisk-naturvetenskapliga vetenskapsområdet, Matematisk-datavetenskapliga sektionen, Institutionen för informationsteknologi, Tillämpad beräkningsvetenskap.
2012 (Engelska)Doktorsavhandling, sammanläggning (Övrigt vetenskapligt)
Abstract [en]

The existence of new technologies, implemented in efficient platforms and workflows has made massive genotyping available to all fields of biology and medicine. Genetic analyses are no longer dominated by experimental work in laboratories, but rather the interpretation of the resulting data. When billions of data points representing thousands of individuals are available, efficient computational tools are required. The focus of this thesis is on developing models, methods and implementations for such tools.

The first theme of the thesis is multi-dimensional scans for quantitative trait loci (QTL) in experimental crosses. By mating individuals from different lines, it is possible to gather data that can be used to pinpoint the genetic variation that influences specific traits to specific genome loci. However, it is natural to expect multiple genes influencing a single trait to interact. The thesis discusses model structure and model selection, giving new insight regarding under what conditions orthogonal models can be devised. The thesis also presents a new optimization method for efficiently and accurately locating QTL, and performing the permuted data searches needed for significance testing. This method has been implemented in a software package that can seamlessly perform the searches on grid computing infrastructures.

The other theme in the thesis is the development of adapted optimization schemes for using hidden Markov models in tracing allele inheritance pathways, and specifically inferring haplotypes. The advances presented form the basis for more accurate and non-biased line origin probabilities in experimental crosses, especially multi-generational ones. We show that the new tools are able to reconstruct haplotypes and even genotypes in founder individuals and offspring alike, based on only unordered offspring genotypes. The tools can also handle larger populations than competing methods, resolving inheritance pathways and phase in much larger and more complex populations. Finally, the methods presented are also applicable to datasets where individual relationships are not known, which is frequently the case in human genetics studies. One immediate application for this would be improved accuracy for imputation of SNP markers within genome-wide association studies (GWAS).

Ort, förlag, år, upplaga, sidor
Uppsala: Acta Universitatis Upsaliensis, 2012. , s. 57
Serie
Digital Comprehensive Summaries of Uppsala Dissertations from the Faculty of Science and Technology, ISSN 1651-6214 ; 973
Nyckelord [en]
quantitative trait loci, genome-wide association studies, hidden Markov models, numerical optimization, linkage analysis, haplotype inference, genotype imputation, high performance computing
Nationell ämneskategori
Beräkningsmatematik Sannolikhetsteori och statistik Bioinformatik och beräkningsbiologi Genetik och genomik Bioinformatik (beräkningsbiologi) Programvaruteknik
Identifikatorer
URN: urn:nbn:se:uu:diva-180920ISBN: 978-91-554-8473-6 (tryckt)OAI: oai:DiVA.org:uu-180920DiVA, id: diva2:552121
Disputation
2012-10-26, Room 2446, Polacksbacken, Lägerhyddsvägen 2D, Uppsala, 13:15 (Engelska)
Opponent
Handledare
Projekt
eSSENCETillgänglig från: 2012-10-04 Skapad: 2012-09-13 Senast uppdaterad: 2025-02-05Bibliografiskt granskad
Delarbeten
1. Coherent estimates of genetic effects with missing information
Öppna denna publikation i ny flik eller fönster >>Coherent estimates of genetic effects with missing information
2012 (Engelska)Ingår i: Open Journal of Genetics, ISSN 2162-4453, E-ISSN 2162-4461, Vol. 2, s. 31-38Artikel i tidskrift (Refereegranskat) Published
Nyckelord
genetic effects, missing genotypes, orthogonal estimation, QTL analysis
Nationell ämneskategori
Bioinformatik och beräkningsbiologi Genetik och genomik Sannolikhetsteori och statistik Beräkningsmatematik
Identifikatorer
urn:nbn:se:uu:diva-180915 (URN)10.4236/ojgen.2012.21003 (DOI)
Projekt
eSSENCE
Tillgänglig från: 2012-03-02 Skapad: 2012-09-12 Senast uppdaterad: 2025-02-05Bibliografiskt granskad
2. Fast and accurate detection of multiple quantitative trait loci
Öppna denna publikation i ny flik eller fönster >>Fast and accurate detection of multiple quantitative trait loci
2013 (Engelska)Ingår i: Journal of Computational Biology, ISSN 1066-5277, E-ISSN 1557-8666, Vol. 20, s. 687-702Artikel i tidskrift (Refereegranskat) Published
Nationell ämneskategori
Bioinformatik och beräkningsbiologi Genetik och genomik Beräkningsmatematik Sannolikhetsteori och statistik
Identifikatorer
urn:nbn:se:uu:diva-180916 (URN)10.1089/cmb.2012.0242 (DOI)000323822000006 ()
Projekt
eSSENCE
Tillgänglig från: 2013-08-06 Skapad: 2012-09-13 Senast uppdaterad: 2025-02-05Bibliografiskt granskad
3. A Grid-Enabled Problem Solving Environment for QTL Analysis in R
Öppna denna publikation i ny flik eller fönster >>A Grid-Enabled Problem Solving Environment for QTL Analysis in R
Visa övriga...
2010 (Engelska)Ingår i: Proc. 2nd International Conference on Bioinformatics and Computational Biology, Cary, NC: ISCA , 2010, s. 202-209Konferensbidrag, Publicerat paper (Refereegranskat)
Ort, förlag, år, upplaga, sidor
Cary, NC: ISCA, 2010
Nationell ämneskategori
Programvaruteknik Genetik och genomik
Identifikatorer
urn:nbn:se:uu:diva-111594 (URN)978-1-880843-76-5 (ISBN)
Projekt
eSSENCE
Tillgänglig från: 2010-01-12 Skapad: 2009-12-17 Senast uppdaterad: 2025-02-01Bibliografiskt granskad
4. cnF2freq: Efficient determination of genotype and haplotype probabilities in outbred populations using Markov models
Öppna denna publikation i ny flik eller fönster >>cnF2freq: Efficient determination of genotype and haplotype probabilities in outbred populations using Markov models
2009 (Engelska)Ingår i: Bioinformatics and Computational Biology, Berlin: Springer-Verlag , 2009, s. 307-319Konferensbidrag, Publicerat paper (Refereegranskat)
Ort, förlag, år, upplaga, sidor
Berlin: Springer-Verlag, 2009
Serie
Lecture Notes in Computer Science ; 5462
Nationell ämneskategori
Beräkningsmatematik Genetik och genomik
Identifikatorer
urn:nbn:se:uu:diva-103916 (URN)10.1007/978-3-642-00727-9_29 (DOI)000265785800029 ()978-3-642-00726-2 (ISBN)
Tillgänglig från: 2009-05-25 Skapad: 2009-05-25 Senast uppdaterad: 2025-02-01Bibliografiskt granskad
5. An improved method for estimating chromosomal line origin in QTL analysis of crosses between outbred lines
Öppna denna publikation i ny flik eller fönster >>An improved method for estimating chromosomal line origin in QTL analysis of crosses between outbred lines
2011 (Engelska)Ingår i: G3: Genes, Genomes, Genetics, E-ISSN 2160-1836, Vol. 1, s. 57-64Artikel i tidskrift (Refereegranskat) Published
Nationell ämneskategori
Beräkningsmatematik Genetik och genomik
Identifikatorer
urn:nbn:se:uu:diva-156197 (URN)10.1534/g3.111.000109 (DOI)000312405400007 ()
Projekt
eSSENCE
Tillgänglig från: 2011-06-01 Skapad: 2011-07-15 Senast uppdaterad: 2025-02-01Bibliografiskt granskad
6. MAPfastR: Quantitative trait loci mapping in outbred line crosses
Öppna denna publikation i ny flik eller fönster >>MAPfastR: Quantitative trait loci mapping in outbred line crosses
Visa övriga...
2013 (Engelska)Ingår i: G3: Genes, Genomes, Genetics, E-ISSN 2160-1836, Vol. 3, s. 2147-2149Artikel i tidskrift (Refereegranskat) Published
Nationell ämneskategori
Beräkningsmatematik Genetik och genomik Bioinformatik och beräkningsbiologi
Identifikatorer
urn:nbn:se:uu:diva-180917 (URN)10.1534/g3.113.008623 (DOI)000328334500005 ()
Projekt
eSSENCE
Tillgänglig från: 2013-10-11 Skapad: 2012-09-13 Senast uppdaterad: 2025-02-05Bibliografiskt granskad
7. Haplotype inference based on hidden Markov models in the QTL–MAS 2010 multigenerational dataset
Öppna denna publikation i ny flik eller fönster >>Haplotype inference based on hidden Markov models in the QTL–MAS 2010 multigenerational dataset
2011 (Engelska)Ingår i: Proc. 14th European Workshop on QTL Mapping and Marker Assisted Selection, London: BioMed Central , 2011, s. S10:1-7Konferensbidrag, Publicerat paper (Refereegranskat)
Ort, förlag, år, upplaga, sidor
London: BioMed Central, 2011
Serie
BMC Proceedings, ISSN 1753-6561 ; 5:3
Nationell ämneskategori
Beräkningsmatematik Genetik och genomik
Identifikatorer
urn:nbn:se:uu:diva-153449 (URN)10.1186/1753-6561-5-S3-S10 (DOI)
Projekt
eSSENCE
Tillgänglig från: 2010-05-17 Skapad: 2011-05-12 Senast uppdaterad: 2025-02-01Bibliografiskt granskad
8. Inferring haplotypes and parental genotypes in larger full sib-ships and other pedigrees with missing or erroneous genotype data
Öppna denna publikation i ny flik eller fönster >>Inferring haplotypes and parental genotypes in larger full sib-ships and other pedigrees with missing or erroneous genotype data
2012 (Engelska)Ingår i: BMC Genetics, E-ISSN 1471-2156, Vol. 13, s. 85:1-13Artikel i tidskrift (Refereegranskat) Published
Nyckelord
haplotyping, phasing, genotype inference, nuclear family data, hidden Markov models
Nationell ämneskategori
Sannolikhetsteori och statistik Beräkningsmatematik Genetik och genomik Bioinformatik och beräkningsbiologi
Identifikatorer
urn:nbn:se:uu:diva-182488 (URN)10.1186/1471-2156-13-85 (DOI)000314354600001 ()
Projekt
eSSENCE
Tillgänglig från: 2012-10-10 Skapad: 2012-10-10 Senast uppdaterad: 2025-02-05Bibliografiskt granskad
9. Breakdown of methods for phasing and imputation in the presence of double genotype sharing
Öppna denna publikation i ny flik eller fönster >>Breakdown of methods for phasing and imputation in the presence of double genotype sharing
2012 (Engelska)Rapport (Övrigt vetenskapligt)
Abstract [en]

In genome-wide association studies, results have been improved through imputation of a denser marker set based on reference haplotypes and phasing of the genotype data. To better handle very large sets of reference haplotypes, pre-phasing with only study individuals has been suggested. We present a possible problem which is aggravated when pre-phasing strategies are used, and suggest a modification avoiding these issues with application to the MaCH tool.

We evaluate the effectiveness of our remedy to a subset of Hapmap data, comparing the original version of MaCH and our modified approach. Improvements are demonstrated on the original data (phase switch error rate decresasing by 10%), but the differences are more pronounced in cases where the data is augmented to represent the presence of closely related individuals, especially when siblings are present (30% reduction in switch error rate in the presence of children, 47% reduction in the presence of siblings). When introducing siblings, the switch error rate in results from the unmodified version of MaCH increases significantly compared to the original data.

The main conclusions of this investigation is that existing statistical methods for phasing and imputation of unrelated individuals might give subpar quality results if a subset of study individuals nonetheless are related. As the populations collected for general genome-wide association studies grow in size, including relatives might become more common. If a general GWAS framework for unrelated individuals would be employed on datasets where sub-populations originally collected as familial case-control sets are included, caution should also be taken regarding the quality of haplotypes.

Our modification to MaCH is available on request and straightforward to implement. We hope that this mode, if found to be of use, could be integrated as an option in future standard distributions of MaCH.

Serie
Technical report / Department of Information Technology, Uppsala University, ISSN 1404-3203 ; 2012-027
Nationell ämneskategori
Sannolikhetsteori och statistik Beräkningsmatematik Genetik och genomik Bioinformatik och beräkningsbiologi
Identifikatorer
urn:nbn:se:uu:diva-181598 (URN)
Projekt
eSSENCE
Tillgänglig från: 2012-09-25 Skapad: 2012-09-26 Senast uppdaterad: 2025-02-05Bibliografiskt granskad

Open Access i DiVA

fulltext(543 kB)3669 nedladdningar
Filinformation
Filnamn FULLTEXT01.pdfFilstorlek 543 kBChecksumma SHA-512
73fe6b655910eb28f1ea5073c001a0d4db5459d856df6190f5226a82ce123dfc13c68d44681eb2248d7842bb5ba1f6ffba423b9597bdd4708228b74123174fb0
Typ fulltextMimetyp application/pdf

Person

Nettelblad, Carl

Sök vidare i DiVA

Av författaren/redaktören
Nettelblad, Carl
Av organisationen
Avdelningen för beräkningsvetenskapTillämpad beräkningsvetenskap
BeräkningsmatematikSannolikhetsteori och statistikBioinformatik och beräkningsbiologiGenetik och genomikBioinformatik (beräkningsbiologi)Programvaruteknik

Sök vidare utanför DiVA

GoogleGoogle Scholar
Totalt: 3671 nedladdningar
Antalet nedladdningar är summan av nedladdningar för alla fulltexter. Det kan inkludera t.ex tidigare versioner som nu inte längre är tillgängliga.

isbn
urn-nbn

Altmetricpoäng

isbn
urn-nbn
Totalt: 2674 träffar
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf