Logo: to the web site of Uppsala University

uu.sePublications from Uppsala University
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Breakdown of methods for phasing and imputation in the presence of double genotype sharing
Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Division of Scientific Computing. Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Computational Science.
2012 (English)Report (Other academic)
Abstract [en]

In genome-wide association studies, results have been improved through imputation of a denser marker set based on reference haplotypes and phasing of the genotype data. To better handle very large sets of reference haplotypes, pre-phasing with only study individuals has been suggested. We present a possible problem which is aggravated when pre-phasing strategies are used, and suggest a modification avoiding these issues with application to the MaCH tool.

We evaluate the effectiveness of our remedy to a subset of Hapmap data, comparing the original version of MaCH and our modified approach. Improvements are demonstrated on the original data (phase switch error rate decresasing by 10%), but the differences are more pronounced in cases where the data is augmented to represent the presence of closely related individuals, especially when siblings are present (30% reduction in switch error rate in the presence of children, 47% reduction in the presence of siblings). When introducing siblings, the switch error rate in results from the unmodified version of MaCH increases significantly compared to the original data.

The main conclusions of this investigation is that existing statistical methods for phasing and imputation of unrelated individuals might give subpar quality results if a subset of study individuals nonetheless are related. As the populations collected for general genome-wide association studies grow in size, including relatives might become more common. If a general GWAS framework for unrelated individuals would be employed on datasets where sub-populations originally collected as familial case-control sets are included, caution should also be taken regarding the quality of haplotypes.

Our modification to MaCH is available on request and straightforward to implement. We hope that this mode, if found to be of use, could be integrated as an option in future standard distributions of MaCH.

Place, publisher, year, edition, pages
2012.
Series
Technical report / Department of Information Technology, Uppsala University, ISSN 1404-3203 ; 2012-027
National Category
Probability Theory and Statistics Computational Mathematics Genetics and Genomics Bioinformatics and Computational Biology
Identifiers
URN: urn:nbn:se:uu:diva-181598OAI: oai:DiVA.org:uu-181598DiVA, id: diva2:557045
Projects
eSSENCEAvailable from: 2012-09-25 Created: 2012-09-26 Last updated: 2025-02-05Bibliographically approved
In thesis
1. Two Optimization Problems in Genetics: Multi-dimensional QTL Analysis and Haplotype Inference
Open this publication in new window or tab >>Two Optimization Problems in Genetics: Multi-dimensional QTL Analysis and Haplotype Inference
2012 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

The existence of new technologies, implemented in efficient platforms and workflows has made massive genotyping available to all fields of biology and medicine. Genetic analyses are no longer dominated by experimental work in laboratories, but rather the interpretation of the resulting data. When billions of data points representing thousands of individuals are available, efficient computational tools are required. The focus of this thesis is on developing models, methods and implementations for such tools.

The first theme of the thesis is multi-dimensional scans for quantitative trait loci (QTL) in experimental crosses. By mating individuals from different lines, it is possible to gather data that can be used to pinpoint the genetic variation that influences specific traits to specific genome loci. However, it is natural to expect multiple genes influencing a single trait to interact. The thesis discusses model structure and model selection, giving new insight regarding under what conditions orthogonal models can be devised. The thesis also presents a new optimization method for efficiently and accurately locating QTL, and performing the permuted data searches needed for significance testing. This method has been implemented in a software package that can seamlessly perform the searches on grid computing infrastructures.

The other theme in the thesis is the development of adapted optimization schemes for using hidden Markov models in tracing allele inheritance pathways, and specifically inferring haplotypes. The advances presented form the basis for more accurate and non-biased line origin probabilities in experimental crosses, especially multi-generational ones. We show that the new tools are able to reconstruct haplotypes and even genotypes in founder individuals and offspring alike, based on only unordered offspring genotypes. The tools can also handle larger populations than competing methods, resolving inheritance pathways and phase in much larger and more complex populations. Finally, the methods presented are also applicable to datasets where individual relationships are not known, which is frequently the case in human genetics studies. One immediate application for this would be improved accuracy for imputation of SNP markers within genome-wide association studies (GWAS).

Place, publisher, year, edition, pages
Uppsala: Acta Universitatis Upsaliensis, 2012. p. 57
Series
Digital Comprehensive Summaries of Uppsala Dissertations from the Faculty of Science and Technology, ISSN 1651-6214 ; 973
Keywords
quantitative trait loci, genome-wide association studies, hidden Markov models, numerical optimization, linkage analysis, haplotype inference, genotype imputation, high performance computing
National Category
Computational Mathematics Probability Theory and Statistics Bioinformatics and Computational Biology Genetics and Genomics Bioinformatics (Computational Biology) Software Engineering
Identifiers
urn:nbn:se:uu:diva-180920 (URN)978-91-554-8473-6 (ISBN)
Public defence
2012-10-26, Room 2446, Polacksbacken, Lägerhyddsvägen 2D, Uppsala, 13:15 (English)
Opponent
Supervisors
Projects
eSSENCE
Available from: 2012-10-04 Created: 2012-09-13 Last updated: 2025-02-05Bibliographically approved

Open Access in DiVA

fulltext(150 kB)110 downloads
File information
File name FULLTEXT01.pdfFile size 150 kBChecksum SHA-512
b7b868348fff53b844b1fa9406d22e334f8c07c50273c9fdb30ca086f5eef4622e0544eb93c7778f067134da3b2e6f79fea0c70bb72f35deee545fe44d7925c3
Type fulltextMimetype application/pdf

Authority records

Nettelblad, Carl

Search in DiVA

By author/editor
Nettelblad, Carl
By organisation
Division of Scientific ComputingComputational Science
Probability Theory and StatisticsComputational MathematicsGenetics and GenomicsBioinformatics and Computational Biology

Search outside of DiVA

GoogleGoogle Scholar
Total: 110 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 623 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf