Complete Tutorial for Protein Phylogenetic Analysis workflow
Published on August 13, 2026

Introduction
Got a batch of homologous sequences and want to build a phylogenetic tree but have no idea where to start? Too many tools, too many complicated parameters, or your computer just can't handle it? Protein phylogenetic analysis is a fundamental core technique in protein engineering and bioinformatics research. Whether it's tracing the evolutionary history of enzyme families, predicting protein functions, or revealing relationships between species, you can't do without it.
This article systematically breaks down the complete analysis steps, common pitfalls, and optimization tips. It also introduces how the online phylogenetic analysis module built into the MatwingsVenus™ (XiaoWu™) platform from Shanghai Matwings Technology enables one-click full-process analysis.
I. What is protein phylogenetic analysis? Why is it important?
Protein phylogenetic analysis is a method of inferring evolutionary relationships between proteins by comparing their sequence similarities and differences and representing them visually as a phylogenetic tree. The branching structure of the tree reflects the closeness of the relationships: the closer the branches, the closer the common ancestor; the longer the branches, the greater the sequence divergence.
For researchers in protein engineering, synthetic biology, and functional genomics, its core value lies in three aspects:
• Family classification and function prediction. Phylogenetic trees provide evolutionary clues for inferring orthologs and paralogs. Combined with species tree reconciliation, this can further confirm the results and provide evolutionary evidence for function prediction, helping to avoid annotation errors.
• Conserved site identification and guidance for directed evolution. By analyzing conservation, you can locate highly conserved functional or structural sites, understand sequence diversity within a family, and provide evolutionary support for designing mutants.
• New function inference. By analyzing gene duplication events and positively selected sites, you can infer the evolutionary drivers of functional divergence and discover protein subfamilies with novel functions.
According to the UniProt Official Release 2026_02 (released in June 2026), the UniProtKB/TrEMBL automatically annotated sub-database contains approximately 149 million protein sequences (note: in 2026, UniProt adjusted the reference proteome strategy for TrEMBL, and non-reference proteome sequences were moved to UniParc, so the number of TrEMBL entries has decreased significantly compared to previous versions). Meanwhile, the manually curated Swiss-Prot sub-database contains about 575,000 entries. Each curated Swiss-Prot entry includes family- and evolution-related annotations, which are the database-level representation of phylogenetic analysis results.
II. The Five-Step Process of Protein Phylogenetic Analysis

Diagram of the Five-Step Standard Process
A standard workflow for protein phylogenetic analysis includes five core steps, and the quality of each step directly affects the reliability of the final phylogenetic tree.
Step 1: How to prepare a high-quality sequence dataset?
Download homologous sequences from databases like UniProt or NCBI, or use sequences from lab experiments, standardize the format to FASTA, and remove redundant and low-quality sequences. The key here is to ensure the completeness and accuracy of the sequence set—missing key branch sequences or including contaminated sequences can cause errors in tree topology.
[Common Pitfall] Focusing only on the number of sequences while ignoring quality. Sequences that are too short, poorly annotated, or contaminated across species can seriously interfere with alignment and tree construction results. Less is more—this is the first principle of sequence preparation.
Step 2: How to choose tools for multiple sequence alignment and improve quality?
Align all sequences and identify homologous sites; this is the basis of phylogenetic analysis. Common tools include MAFFT, MUSCLE, and Clustal Omega. After alignment, trimming is often needed—to remove regions with too many gaps or poor alignment quality—commonly using trimAl or Gblocks. Misaligned sites can introduce 'noise,' leading to biased tree topology.
[Common Pitfall] Constructing trees directly from untrimmed alignments. In practice, regions with more than 50% gaps often lack reliable homologous information, so trimming before downstream analysis is recommended.
Step 3: How to choose the evolutionary model?
Different protein families evolve differently, with large variations in amino acid substitution rates, heterogeneity among sites, and proportions of invariant sites. Choosing the right amino acid substitution model (like JTT, LG, WAG, etc.) is essential for building an accurate phylogenetic tree. Common tools include ModelFinder and ProtTest.
[Common Pitfall] Randomly picking a model and starting tree construction. Choosing the wrong model can lead to biased branch length estimates or even incorrect topology, so spending a few minutes running a model selection tool is well worth it.
Step 4: Which method to use for tree construction?
Based on the alignment and evolutionary model, construct a phylogenetic tree using Maximum Likelihood (ML), Bayesian Inference (BI), or Neighbor-Joining (NJ). The ML method is currently the most popular due to its balance of accuracy and efficiency, with representative tools like IQ-TREE 3 and RAxML-NG. IQ-TREE 3, as the latest generation of phylogenetic software, has significant improvements in model selection, computational efficiency, and result interpretability.
【Common Misconceptions】 The number of bootstrap replicates in bootstrap analysis isn’t better the more you do. 1,000 bootstrap replicates is the usual choice, but for preliminary exploratory analysis, 100–500 replicates are enough to get a rough idea of the topology; only the final version for publication needs higher replicate numbers.
Step 5: What should you pay attention to when visualizing phylogenetic trees?
Visualize tree files in Newick format, adjust branch colors, labels, outgroup positions, and add multidimensional information like functional annotation, domains, or expression levels, ultimately producing a publication-ready phylogenetic tree figure. Common tools include iTOL, FigTree, and ggtree.
【Common Misconceptions】 Visualization is not the same as "beautifying." The core of phylogenetic tree visualization is to clearly convey biological information—like which branches show functional divergence, which sites are under selection, or whether the outgroup is reasonable—far more important than flashy color schemes.
III. Four Main Pain Points of Traditional Protein Phylogenetic Analysis
Although the methodology is quite mature, running a full protein phylogenetic analysis workflow is still challenging.
Pain Point 1: Long toolchains and high learning cost. From sequence alignment to phylogenetic tree visualization, you need to master multiple independent tools, each with different installation methods, parameter systems, and input/output formats. Beginners often take weeks just to get the workflow running.
Pain Point 2: Time-consuming calculations and insufficient local computing power. Multiple sequence alignment and phylogenetic tree construction are computationally intensive tasks. Constructing an ML tree for hundreds of sequences may take days on an ordinary desktop; Bayesian analysis or high-replicate bootstrap calculations can grow exponentially.
Pain Point 3: Difficult parameter selection, hard-to-assess results. Which alignment tool to use? Which substitution model to choose? How many bootstrap replicates? There’s no standard answer; it heavily depends on experience. Beginners often rely on trial-and-error, ending up with unstable phylogenetic trees and little understanding of how to evaluate tree reliability.
Pain Point 4: Time-consuming visualization, editing the figure is harder than building it. Constructing a phylogenetic tree might take only a day, but tweaking it to be publication-ready can take a week—branch thickness, label positions, annotation layers, color schemes, every detail needs repeated fine-tuning.

Traditional Multi-Tool Process vs MatwingsVenus™
IV. MatwingsVenus™: One-Stop Online Phylogenetic Analysis
Facing the pain points mentioned above, an ideal online tool for protein phylogenetic analysis should answer four issues: long tool chains → one-stop integration, insufficient computing power → cloud-based elastic computing, difficulty choosing parameters → smart parameter recommendations, visualization takes too much time → automated chart generation.
A quick comparison: traditional workflows require five tools, several days of learning, and are limited by local computing power; while the MatwingsVenus™ online platform offers one-stop cloud analysis, smart parameter recommendations, and automated visualization, producing usable phylogenetic trees in just a few hours for a typical dataset.
Shanghai Matwings Technology’s MatwingsVenus™ conversational protein R&D assistant was designed with this approach in mind. The platform has built-in phylogenetic analysis modules, so researchers don’t need to install software or use the command line—they just upload sequences or describe their needs in natural language, and the platform handles the full process from sequence alignment to tree beautification.
High-performance cloud computing, fully deployment-free. Based on a cloud-native architecture, it integrates mainstream multiple sequence alignment tools and phylogenetic tree construction algorithms, supporting alignment strategies like MAFFT and MUSCLE, and providing efficient maximum likelihood tree building. Typical datasets (hundreds of sequences of medium length) can be completed in minutes; for thousands of sequences, depending on model complexity and bootstrap settings, results usually come in hours—completely eliminating local computing limitations.
Smart parameter recommendations, lowering the usage barrier. The platform automatically recommends the most suitable alignment tools, evolutionary models, and tree-building methods based on sequence number, length, and family characteristics. Users can also manually adjust parameters according to research needs. Every analysis step comes with methodological explanations and parameter descriptions, so researchers not only get results but also understand how they were generated.
Automated visualization, producing ready-to-use charts directly. Once a tree is built, the platform automatically generates interactive visualization pages, allowing online adjustment of topology, outgroup selection, branch colors, and label styles, with one-click addition of species info, functional annotation, domain distribution, and other layered annotations. All charts support high-resolution vector export, ready for journal submission.
Conversational interaction, doing phylogenetic analysis in natural language. Unlike traditional tools, MatwingsVenus™ uses a conversational intelligent agent. You can say something like, 'Help me build a maximum likelihood tree with these sequences using the LG G4 model and 1000 bootstrap replicates,' or ask, 'Which branches in this tree have low bootstrap support?' The agent automatically schedules the tool chain and returns results with a natural language interpretation.

MatwingsVenus™(晓鹜™)
V. Application Value of Three Major Research Scenarios
Protein Family Functional Evolution Research: Construct evolutionary trees for all homologous sequences of the target family, and analyze functional divergence patterns of different branches in combination with functional annotations. Identify conserved sites and rapidly evolving sites to provide evolutionary evidence for subsequent experimental verification. The MatwingsVenus™ (Xiaowu™) platform can directly link evolutionary analysis results to downstream functional site prediction and directed evolution design modules, forming a complete research workflow.
Metagenomic Novel Enzyme Discovery: Perform phylogenetic analysis on candidate enzyme sequences predicted from metagenomic data. If a sequence forms an independent evolutionary branch and is distant from known functional enzymes, it is considered a potential new subfamily, prioritizing experimental activity verification and significantly improving screening efficiency. The MatwingsVenus™ (Xiaowu™) platform supports clustering and evolutionary analysis of large batches of sequences in one go.
Guidance for Directed Evolution Modification: Conduct phylogenetic analysis on natural homologous sequences to identify evolutionarily conserved and rapidly evolving sites. Conserved sites are usually crucial for maintaining protein structure and core functions, while rapidly evolving sites often relate to functional specificity and are ideal targets for directed evolution modification.
VI. Frequently Asked Questions in Research | 5 Common Phylogenetic Analysis Questions
Q1: What should I do if the confidence of evolutionary tree branches is low?
It’s most likely due to a messy dataset or excessive alignment noise. On the MatwingsVenus™ (Xiaowu™) platform, you can first use UniRef90 clustering to remove redundancy, perform precise alignment to lock conserved domains, and automatically trim low-quality regions, significantly improving bootstrap support for branches.
Q2: Can phylogenetic analysis be done with low-homology sequences?
Traditional alignment tools have limited sensitivity for distant sequences with less than 30% identity. MatwingsVenus™ (Xiaowu™), leveraging a self-developed protein language model for deep alignment, can capture hidden evolutionary signals in low-homology sequences, supporting clustering and phylogenetic/functional inference even for distant families or artificially designed proteins.
Q3: How can evolutionary tree figures meet SCI paper requirements?
The key is using proper models and methods with clear visualization. MatwingsVenus™ (Xiaowu™) uses mainstream academic standard evolutionary models and tree-building algorithms throughout the workflow, outputs high-resolution vector images, supports multi-level custom annotations, and the format meets the requirements of major biology journals, ready for use in the main text or supplementary material.
Q4: What should I do if large-scale sequence phylogenetic analysis can't run?
Running thousands of sequences on a regular computer does take too long. MatwingsVenus™ (XiaoWu™) distributed cloud computing has no hardware barriers. After uploading a batch of sequences, the platform can automatically perform clustering to reduce the dataset. Thousands of sequences can be aligned, trees built, and visualized within a few hours.
Q5: How to choose outgroups for a phylogenetic tree?
Outgroups should be distantly related homologous sequences that have some relationship to the target sequences but clearly belong to a different branch. They are used to determine the root of the tree. The MatwingsVenus™ (XiaoWu™) platform comes with the UniProt classification database and can automatically recommend suitable distant homolog outgroup sequences, so there's no need to manually search and download—they can be set as roots with one click.
Conclusion: Let evolutionary analysis return to the biological questions themselves
From manual alignment to automated workflows, from neighbor-joining methods to maximum likelihood and Bayesian approaches, the protein phylogenetic analysis workflow has evolved toward being "more accurate, faster, and easier to use." But as tools get more complex, many researchers spend a lot of time on operations rather than focusing on the biological questions they actually want to answer.
Shanghai Matwings Technology's MatwingsVenus™ (XiaoWu™) platform is changing that. It positions phylogenetic analysis as a core part of the protein research and development chain, connecting upward to billions of sequences databases and intelligent homolog searches, and downward to functional site prediction, directed evolution design, and automated experimental validation, forming a complete closed loop: "evolutionary analysis → functional interpretation → experimental verification."
Only when researchers no longer have to worry about setting up environments, adjusting parameters, or organizing data can phylogenetic analysis truly realize its value as a powerful tool in evolutionary research.