Predicting the Combined Effects of Protein Mutations: Cracking the Epistasis
Published on August 18, 2026

Introduction
In protein engineering and enzyme engineering research, the performance improvement from a single-point mutation is often limited. True performance jumps usually come from the synergy of multiple mutation sites. However, the combination space of amino acids is astronomical—saturated combinations for just 10 sites amount to 20¹⁰ possibilities, far beyond the capacity of experimental screening. Predicting the combinatorial effects of protein mutations is the core computational method to solve this problem: by modeling the epistasis between mutation sites, it identifies synergistic mutations that can actually lead to functional leaps from a massive number of combinations, upgrading the traditional trial-and-error 'random library and high-throughput screening' approach into a more targeted and rational design. This article systematically breaks down the technical principles, main challenges, and platform solutions for combinatorial effect prediction, and introduces how Shanghai Matwings Technology's MatwingsVenus™ (Xiaowu™) platform drives directed evolution with AI.
1. What is Protein Mutation Combinatorial Effect Prediction?
The function of a protein is determined by both its amino acid sequence and its three-dimensional structure. Predicting the effect of a single-point mutation evaluates how replacing a single amino acid affects the phenotype, which is relatively mature. In contrast, predicting the combinatorial effects of protein mutations focuses on the interactions between sites when multiple mutations are introduced simultaneously—what we call epistasis: the functional impact of one mutation depends on the amino acid state at one or more other sites.
These nonlinear effects are everywhere in protein engineering. Two single-point mutations that individually slightly increase activity could, when combined, lead to several or even tens of times performance improvement (positive epistasis/synergy), or they might cancel each other out or even have negative effects (negative epistasis/antagonism). Traditional methods can barely predict this kind of epistasis, relying instead on large-scale random library and high-throughput screening, which is expensive, time-consuming, and has a low success rate.
From a prediction perspective, combinatorial effect prediction covers the main indicators of interest in protein engineering: changes in stability (ΔΔG, the difference in folding free energy between mutant and wild type, with positive values usually indicating decreased stability and negative values indicating increased stability), enzyme activity and catalytic efficiency, binding affinity, expression and solubility, optimum temperature/pH, and other physicochemical properties. Among these, the nonlinear gains from multiple-point combinations are key to achieving orders-of-magnitude performance breakthroughs in directed evolution.
The core value of combinatorial effect prediction lies in freeing researchers from the 'needle-in-a-haystack' random screening—using computational methods to narrow the search space in advance, lock onto high-value synergistic mutation combinations, and then invest in experimental validation, significantly boosting the efficiency and success rate of directed evolution.
2. Technology Evolution: From Single-Point Scoring to Combinatorial Design

The Evolution of Three Generations of Technology
The prediction of protein mutation effects has gone through three generations of methodological evolution, with continuous improvements in prediction accuracy and applicability. The research focus is also shifting from single-point mutations to combinatorial effects.
First Generation: Statistical methods based on conservation and physico-chemical properties. Early methods evaluated the evolutionary conservation of sites through multiple sequence alignment (MSA) or inferred mutation effects using amino acid physico-chemical properties. Representative tools include conservation scores based on homology comparison, functional impact predictions using BLOSUM substitution matrices, and statistical learning tools for protein stability prediction like CUPSAT and I-Mutant. These methods are fast to compute and easy to interpret but cannot capture long-range dependencies or local structural environment changes, and their ability to predict combinatorial mutations is almost zero.
Second Generation: Computational methods based on 3D structures and physical force fields. With advances in structural determination, physics- or statistical potential-based methods like FoldX and Rosetta have been used to calculate mutation stability changes (ΔΔG) with clear physical meaning. However, these methods are computationally expensive, heavily depend on high-quality experimental structures, and accumulate significant errors for multi-point combinatorial mutations, limiting their practical use.
Third Generation: Deep learning methods based on protein language models. In recent years, protein language models (pLMs) using the Transformer architecture have completely changed the field. By self-supervised pretraining on tens of billions of sequences, pLMs learn evolutionary information and contextual sequence representations, capturing long-range dependencies via the self-attention mechanism. On standard test sets like ProteinGym (covering 217 deep mutational scanning experiments and around 2.7 million mutants), pLM-based methods have significantly outperformed traditional methods across tasks such as enzyme activity, molecular binding, expression levels, and stability in terms of Spearman correlation. More importantly, these methods can make zero-shot predictions using only sequence input, breaking the reliance on experimental structures and labeled data.
Since 2026, research has further focused on multi-point combinatorial mutations and higher-order effect prediction. Several studies published in top journals like Science and PNAS show that frameworks combining protein language models with higher-order effect modeling achieved up to 10-fold performance improvements with just a single round of machine learning-guided evolution; methods reconstructing high-resolution fitness landscapes using co-mutation information from directed evolution trajectories have identified MEK1 variants with over 1000-fold drug resistance. These advances signify that protein mutation effect prediction is moving from 'single-point scoring' to 'combinatorial design,' and from 'predicting effects' to 'creating functions.'
3. the four core challenges of combinatorial effect prediction.

Four Major Challenges
Despite rapid advances in algorithms, predicting the combined effects of protein mutations still faces multiple bottlenecks in practical research applications.
Challenge 1: Hard-to-model epistatic effects lead to insufficient combo prediction accuracy. Predictions for single-point mutations are relatively mature, but multi-point combinations involve extremely complex non-linear epistatic effects, which make traditional methods' accuracy drop sharply. In protein engineering, the real performance jumps often come from the synergy of multiple mutation sites — and this "core need" is exactly the "core weakness" of current methods.
Challenge 2: Exploding sequence space means search scales far exceed computational and experimental capacity. With 20 amino acids × dozens of mutable sites, the combination space grows exponentially. Even computational screening can't exhaust all possibilities, and experimental validation can only scratch the surface.
Challenge 3: Predictions are hard to validate, and iteration cycles are long. Candidate combinations from computational predictions still need wet-lab verification. Going from predicted results to mutant construction, expression, purification, and functional testing can take weeks or even months per iteration, and the barrier between computation and experiments seriously slows down R&D.
Challenge 4: High entry threshold makes it tough for experiment-focused researchers. Most advanced pLM combination prediction methods require command-line operations, GPU setup, and installing deep learning dependencies, which is very hard for enzyme engineers with an experimental background to get started.
4. MatwingsVenus™ (Xiaowu™): An integrated platform for combinatorial mutation design and directed evolution

MatwingsVenus™(晓鹜™)
Facing the challenges mentioned above, Shanghai Matwings Technology has independently developed MatwingsVenus™ (Xiaowu™), a conversational protein R&D intelligent agent, which deeply integrates the ability to predict the effects of protein mutation combinations into the full chain of directed evolution design. The platform is centered on self-developed Venus series protein large models, trained on a dataset of tens of billions of protein sequences, covering complete prediction and design capabilities from single-point scans to multi-point combination optimization.
Driven by self-developed large models for multi-dimensional unified design. The platform, based on the Venus general protein model and a specialized model for directed evolution, supports multi-dimensional mutation effect analysis and optimization design for stability, activity, expression, solubility, and more, allowing researchers to avoid switching between multiple tools and formats. Several proteins designed by the platform have already reached industrialization, such as PET-degrading enzymes and high-activity alkaline phosphatase.
Synergistic optimization of multi-point combined mutations. To tackle the challenge of synergistic effects in multi-point combined mutations, the platform leverages the sequence-function mapping capabilities of large models to help identify combinations of mutation sites with positive synergy. Through multi-round iterative optimization strategies, it continuously approaches the target function, breaking the bottlenecks of traditional combinatorial design and providing more precise candidate solutions for directed evolution. In the Sweet protein Monellin modification project, through multiple rounds of 'predict-design-validate' iterations, several variants increased sweetness more than tenfold compared to the wild type, while maintaining heat resistance around 75°C — a performance leap that is a typical result of multi-point synergistic mutations.
Conversational interaction lowers the usage barrier. Users only need to describe their needs in natural language — for example, 'Scan candidate sites on this sequence and design multi-point mutation combinations to improve thermal stability' — and the intelligent agent automatically schedules the model computations, returning a list of candidate mutations with confidence scores and functional interpretations. No command-line operations or GPU setups are needed, making it easy for experimental researchers to get started quickly.
From prediction to experimental closed-loop. More transformatively, the platform connects mutation effect predictions directly with automated wet-lab experiments. High-value predicted mutation combinations can be instantly converted into experimental plans and executed in an automated shared lab for gene synthesis, protein expression, purification, and functional testing. Experimental results are automatically fed back to guide the next design round, forming a 'predict-design-validate-iterate' closed loop. The platform integrates over 200 protein design tools and more than 30 expert-tuned skill packages around the task goal with the intelligent agent at the center, automatically organizing the workflow and truly bridging computation with experimentation.
5. Three Major Research Application Scenarios
Modification of Industrial Enzyme Thermal Stability and Tolerance. In industrial enzyme applications, thermal stability and pH tolerance are core engineering indicators, and performance leaps often require multi-point cooperative mutations. The platform can systematically scan the sequence space of target enzymes to screen for combinations of mutations that enhance stability, and further improve industrial properties like temperature and alkali tolerance through multiple rounds of optimization. In official cases, the platform has successfully delivered design projects for highly stable industrial enzymes, such as PET-degrading enzymes.
Combined Protein Affinity Optimization. In the development of receptor-binding proteins or functional ligands, affinity maturation is a key step, usually requiring coordinated optimization at multiple interface sites. The platform can predict how combinations of interface residue mutations affect binding affinity and quickly screen for cooperative mutation combinations that improve affinity. In de novo design projects targeting immune-regulatory receptors, the platform has successfully generated dozens of entirely new binding molecules with in vitro cellular blocking activity, completing the full process from design to functional validation.
Modification of Enzyme Activity and Substrate Specificity. For modifying enzymatic catalytic functions, single-point mutations yield limited improvements, while combination mutations often lead to order-of-magnitude breakthroughs. The platform can predict how combinations of mutations in catalytic pockets and surrounding residues affect enzyme activity and substrate selectivity, and, combined with structural information, design variants with higher activity or better substrate profiles, providing efficient design solutions for synthetic biology and biocatalysis applications.
Conclusion: From Single-Point Trial-and-Error to Combination Design Paradigm Shift
From conservation scoring to physical force field calculations, and now to protein language models, the evolution of predicting protein mutation combination effects has always focused on one goal: in the vast sequence combination space, more accurately finding those cooperative mutations that truly drive functional improvements. Today, we can already make fairly reliable predictions for single-point mutations; tomorrow, combination effect modeling will turn directed evolution from 'luck-based' to 'methodical,' moving from single-point optimization to systematic design.
The Shanghai Matwings Technology MatwingsVenus™ (Xiaowu™) platform is driving this transformation. It uses protein mutation combination effect prediction as a core entry point, linking upwards to billions of sequence databases and large protein models, and downwards to directed evolution design, automated experimental validation, and wet-dry loop iteration. This upgrades the traditional 'random mutation + high-throughput screening' trial-and-error approach to a rational design paradigm of 'intelligent prediction + precise validation.'
The next breakthrough in protein engineering will belong to researchers who can deeply integrate AI combination mutation prediction with experimental validation. MatwingsVenus™ (Xiaowu™) is precisely providing a unified capability foundation for such research.