Logo: to the web site of Uppsala University

uu.sePublications from Uppsala University
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Automated Removal of Non-homologous Sequence Stretches with PREQUAL
Uppsala University, Disciplinary Domain of Science and Technology, Biology, Department of Organismal Biology, Systematic Biology. Univ Gottingen, Inst Microbiol & Genet, Dept Appl Bioinformat, Gottingen, Germany; Department of Biodiversity and Evolutionary Biology, Museo Nacional de Ciencias Naturales, Madrid, Spain.ORCID iD: 0000-0002-3628-1137
Uppsala University, Disciplinary Domain of Science and Technology, Biology, Department of Organismal Biology, Systematic Biology. Uppsala University, Science for Life Laboratory, SciLifeLab.ORCID iD: 0000-0002-8248-8462
Uppsala University, Disciplinary Domain of Science and Technology, Biology, Department of Ecology and Genetics, Evolutionary Biology.
2021 (English)In: Multiple Sequence Alignment: Methods and Protocols / [ed] Kazutaka Katoh, Humana Press, 2021, p. 147-162Chapter in book (Refereed)
Abstract [en]

Large-scale multigene datasets used in phylogenomics and comparative genomics often contain sequence errors inherited from source genomes and transcriptomes. These errors typically manifest as stretches of non-homologous characters and derive from sequencing, assembly, and/or annotation errors. The lack of automatic tools to detect and remove sequence errors leads to the propagation of these errors in large-scale datasets. PREQUAL is a command line tool that identifies and masks regions with non-homologous adjacent characters in sets of unaligned homologous sequences. PREQUAL uses a full probabilistic approach based on pair hidden Markov models. On the front end, PREQUAL is user-friendly and simple to use while also allowing full customization to adjust filtering sensitivity. It is primarily aimed at amino acid sequences but can handle protein-coding nucleotide sequences. PREQUAL is computationally efficient and shows high sensitivity and accuracy. In this chapter, we briefly introduce the motivation for PREQUAL and its underlying methodology, followed by a description of basic and advanced usage, and conclude with some notes and recommendations. PREQUAL fills an important gap in the current bioinformatics tool kit for phylogenomics, contributing toward increased accuracy and reproducibility in future studies.

Place, publisher, year, edition, pages
Humana Press, 2021. p. 147-162
Series
Methods in Molecular Biology, ISSN 1064-3745, E-ISSN 1940-6029 ; 2231
Keywords [en]
Filtering, Genomics, HMM, Homology, Phylogenomics, Sequence analysis
National Category
Molecular Biology Bioinformatics (Computational Biology)
Identifiers
URN: urn:nbn:se:uu:diva-556188DOI: 10.1007/978-1-0716-1036-7_10ISI: 000697781700011PubMedID: 33289892Scopus ID: 2-s2.0-85097583339ISBN: 9781071610367 (electronic)ISBN: 9781071610350 (print)OAI: oai:DiVA.org:uu-556188DiVA, id: diva2:1957717
Funder
Carl Tryggers foundation Available from: 2025-05-12 Created: 2025-05-12 Last updated: 2025-05-12Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textPubMedScopus

Authority records

Irisarri, IkerBurki, FabienWhelan, Simon

Search in DiVA

By author/editor
Irisarri, IkerBurki, FabienWhelan, Simon
By organisation
Systematic BiologyScience for Life Laboratory, SciLifeLabEvolutionary Biology
Molecular BiologyBioinformatics (Computational Biology)

Search outside of DiVA

GoogleGoogle Scholar

doi
pubmed
isbn
urn-nbn

Altmetric score

doi
pubmed
isbn
urn-nbn
Total: 55 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf