Research Protocol for Phylogenetic tree
Materials Required
Background
A phylogenetic tree is a hypothesis of evolutionary relationships among homologous sequences or taxa, inferred from shared sequence variation using distance-based, maximum-likelihood, Bayesian, or related statistical methods[1][2][3].
Modern phylogenetic inference depends on a curated sequence set, a reliable multiple sequence alignment, an appropriate substitution model, tree-search strategy, and branch-support assessment[3][4][5][6].
Phylogenetic trees are used to study species relationships, gene-family evolution, pathogen transmission, protein diversification, orthology/paralogy, and evolutionary origin of functional traits[3][7].
Unresolved issues include alignment uncertainty, model misspecification, recombination, horizontal gene transfer, long-branch attraction, insufficient phylogenetic signal, and disagreement between gene trees and species trees[8][9][10].
MCE has not independently verified the accuracy of these methods. They are for reference only.
Project Analysis
• Next, align sequences using an appropriate multiple sequence alignment method, inspect conserved regions, identify uncertain alignment columns, and avoid overinterpreting highly gapped or poorly aligned regions[8][10][11].
• Then, select a substitution model or partition scheme using model-selection tools and infer trees using maximum likelihood, Bayesian inference, or both, depending on dataset size, computational resources, and the need for posterior uncertainty estimates[4][5][6][14].
• After tree inference, assess clade support using bootstrap-based and/or Bayesian measures, compare alternative topologies when needed, and check whether support values are consistent with biological interpretation[6][12][13].
• Finally, validate major conclusions by repeating analyses with alternative alignments, alternative models, trimmed and untrimmed datasets, and different outgroup choices when those factors could affect the result[8][9][10].
Phased Objectives
Objective 1: Define the phylogenetic question and sequence set.
• Research approach: define whether the goal is species-tree inference, gene-tree inference, strain typing, orthology analysis, viral evolution, or protein-family evolution.• Experimental model: homologous DNA, RNA, or protein sequences.
• Experimental groups: ingroup taxa, outgroup taxa, reference sequences, duplicate or paralog candidates, and excluded low-quality sequences.
• Key techniques: database sequence retrieval, BLAST search, orthology screening, domain annotation, and metadata curation.
• Detection indices: taxon coverage, sequence length, missing data, percent identity, outgroup suitability, and evidence of paralogy or recombination.
• Expected results: a curated sequence set that matches the biological question.
• Interpretation: incorrect taxon sampling or inclusion of non-homologous sequences will compromise the tree before inference begins[3][7][8].
Objective 2: Generate and evaluate the alignment.
• Research approach: align curated homologs and assess alignment uncertainty before tree inference.• Experimental model: curated FASTA sequence set.
• Experimental groups: unaligned sequences, aligned sequences, confidence-scored alignment, and sensitivity alignments from alternative aligners.
• Key techniques: multiple sequence alignment, manual inspection of conserved motifs, GUIDANCE/TCS-style confidence assessment, and conservative trimming if justified.
• Detection indices: gap distribution, conserved-site alignment, low-confidence regions, retained alignment length, and alignment-method concordance.
• Expected results: conserved homologous regions should align consistently, while highly gapped or divergent regions may remain uncertain.
• Interpretation: tree conclusions depending on low-confidence alignment regions should be treated cautiously[8][10][11].
Objective 3: Infer phylogenetic trees using appropriate models.
• Research approach: infer trees using model-based methods and compare with simpler methods when useful.• Experimental model: final alignment.
• Experimental groups: neighbor-joining tree, maximum-likelihood tree, Bayesian tree, partitioned model tree, and alternative-model tree.
• Key techniques: model selection, maximum-likelihood inference, Bayesian inference, partition analysis, and tree topology comparison.
• Detection indices: best-fit substitution model, log-likelihood, topology, branch length, posterior probability, bootstrap support, and convergence diagnostics.
• Expected results: well-supported relationships should remain stable across reasonable model choices.
• Interpretation: relationships that change with method or model should be reported as uncertain[1][2][4][5][6].
Objective 4: Validate support and biological interpretation.
• Research approach: quantify branch support and test whether the inferred tree supports the biological hypothesis.• Experimental model: inferred tree set.
• Experimental groups: original tree, bootstrap trees, Bayesian posterior tree sample, alternative topology trees, and constrained trees if testing hypotheses.
• Key techniques: nonparametric bootstrap, ultrafast bootstrap, approximate likelihood-ratio testing, Bayesian posterior probability analysis, and topology tests.
• Detection indices: bootstrap percentage, posterior probability, concordance among methods, tree distance, and support for predefined clades.
• Expected results: robust clades should have high support and appear across inference approaches.
• Interpretation: weakly supported or method-dependent clades should not be used as strong evidence for evolutionary conclusions[6][12][13].
Critical Points
Objective 1
• Produce a clean and biologically justified dataset; if paralogs, contaminants, or unsuitable outgroups remain, the resulting tree may answer the wrong evolutionary question[3][7][8].Objective 2
• Produce an alignment with clearly defined high-confidence and uncertain regions; if alignment uncertainty is high, downstream phylogenetic conclusions should be limited to stable regions or tested across alternative alignments[8][10][11].Objective 3
• Produce one or more inferred trees with documented models and reproducible settings; stable topology across methods supports the hypothesis, while strong method dependence suggests insufficient signal or model problems[4][5][6].Objective 4
• Identify which branches are strongly supported and which remain unresolved; strong support for hypothesis-relevant clades supports the research hypothesis, while weak support or conflicting topologies refute or weaken it[12][13].Troubleshooting
1: poor multiple sequence alignment can generate false phylogenetic relationships.
Alternative: compare aligners, score alignment confidence, and repeat tree inference after conservative treatment of low-confidence regions[8][10][11].2: model misspecification can bias topology and branch lengths.
Alternative: use model-selection procedures and compare results under alternative plausible models or partition schemes[6][14].3: long-branch attraction can group rapidly evolving taxa incorrectly.
Alternative: improve taxon sampling, remove unstable sequences only when justified, compare model-based methods, and test alternative topologies[9][13].4: recombination or horizontal gene transfer can make a single tree inappropriate for the full sequence.
Alternative: screen for recombination, analyze non-recombinant regions separately, or use gene-tree/species-tree-aware approaches when gene histories differ[7][15].5: bootstrap support and Bayesian posterior probabilities are not interchangeable.
Alternative: report the support method used and interpret bootstrap values and posterior probabilities separately[12].References:
- [1]. Saitou N, et al. The neighbor-joining method: a new method for reconstructing phylogenetic trees. Mol Biol Evol. 1987;4(4):406-425. [Content Brief]
- [2]. Felsenstein J. Evolutionary trees from DNA sequences: a maximum likelihood approach. J Mol Evol. 1981;17(6):368-376. [Content Brief]
- [3]. King KM, et al. Building viral phylogenetic trees using a maximum likelihood approach. Curr Protoc Microbiol. 2018;51(1):e63. [Content Brief]
- [4]. Guindon S, et al. A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood. Syst Biol. 2003;52(5):696-704. [Content Brief]
- [5]. Stamatakis A. RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics. 2014;30(9):1312-1313. [Content Brief]
- [6]. Nguyen LT, et al. IQ-TREE: a fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies. Mol Biol Evol. 2015;32(1):268-274. [Content Brief]
- [7]. Morel B, et al. GeneRax: a tool for species-tree-aware maximum likelihood-based gene family tree inference under gene duplication, transfer, and loss. Mol Biol Evol. 2020;37(9):2763-2774. [Content Brief]
- [8]. Chatzou M, et al. Multiple sequence alignment modeling: methods and applications. Brief Bioinform. 2016;17(6):1009-1023.
- [9]. Hossain ASMM, et al. Evidence of statistical inconsistency of phylogenetic methods in the presence of multiple sequence alignment uncertainty. Genome Biol Evol. 2015;7(8):2102-2116. [Content Brief]
- [10]. Tan G, et al. Current methods for automated filtering of multiple sequence alignments frequently worsen single-gene phylogenetic inference. Syst Biol. 2015;64(5):778-791. [Content Brief]
- [11]. Sela I, et al. GUIDANCE2: accurate detection of unreliable alignment regions accounting for the uncertainty of multiple parameters. Nucleic Acids Res. 2015;43(W1):W7-W14. [Content Brief]
- [12]. Douady CJ, et al. Comparison of Bayesian and maximum likelihood bootstrap measures of phylogenetic reliability. Mol Biol Evol. 2003;20(2):248-254. [Content Brief]
- [13]. Felsenstein J. Confidence limits on phylogenies: an approach using the bootstrap. Evolution. 1985;39(4):783-791. [Content Brief]
- [14]. Kalyaanamoorthy S, et al. ModelFinder: fast model selection for accurate phylogenetic estimates. Nat Methods. 2017;14(6):587-589. [Content Brief]
- [15]. Martin DP, et al. RDP4: detection and analysis of recombination patterns in virus genomes. Virus Evol. 2015;1(1):vev003. [Content Brief]