Back to list

Protein optimization: from trying it out to calculating it

Published on July 7, 2026

Protein optimization: from trying it out to calculating it

Nature has evolved hundreds of millions of proteins, each a product of billions of years of evolution. But when you need an enzyme that works continuously for three months in an 80°C reactor without deactivation, an antibody that locks onto a specific antigen on the surface of cancer cells with picmolar affinity, or an industrial protein that shuttles between strong acids and bases without collapse—nature falls silent.

 

Protein optimization was created precisely to solve this "natural shortage" problem. From enhancing the affinity of antibody drugs, to improving the heat resistance of industrial enzymes, and further enhancing the functionality of food proteins—this technology is moving from basic laboratory research to industrial value creation.

 

1. What exactly is protein optimization supplied?

 

The term "protein optimization" covers a wide range of scenarios, but ultimately all converge on several core dimensions of "boundary pushing further."

 

Stability—helping proteins "live longer." This is the most common optimization requirement in industrial applications. Natural enzymes often become inactive within minutes or even seconds in high temperatures, extreme acids and bases, or organic solvents. The optimization goal is to raise the melting temperature (Tm), pushing the enzyme's working window from "hours" to "weeks" or even "months." In the detergent industry, an optimized protease needs to remain active at room temperature for several months in an alkaline solution containing surfactants; In the field of biofuels, cellulase needs to be continuously hydrolyzed at 50–60°C for several days without becoming inactivated.

 

Affinity—making the "lock" and the "key" more compatible. This is the core dimension of therapeutic antibody optimization. The lead molecule obtained from natural antibodies or in vitro screening typically has KD in the range of micromolar to low nanomolar (10⁻⁶–10⁻⁹ M). The optimization aims to push this figure toward the Pemoore class (10⁻¹² M) or even the Femole class (10⁻¹⁵ M)—about a 1000-fold increase from Namole to Pemoore level, and up to a million times to Femoore level. Higher affinity means lower dosages, less off-target toxicity, and better clinical efficacy.

 

Expressiveness—making "manufacturing" more economical. A perfectly functioning protein cannot be industrially produced if its expression in cells is only at the milligram per liter level. Optimizing codons, improving folding efficiency, reducing aggregation and degradation—these are all hurdles that must be overcome on the "manufacturing side."

 

Specificity—making the "front sight" more accurate. The substrate selectivity of enzymes and the cross-reactivity of antibodies directly determine whether the product can be used safely. A hydrolase that degrades PET plastics preferably has high selectivity for PET ester bonds and minimal activity for other polyester substrates—in mixed waste plastic recycling scenarios, high specificity means higher product purity and lower downstream separation costs.

 

Solubility and aggregation tendency — making 'processing' more stable. High-concentration formulations are the mainstream trend for antibody drugs (subcutaneous injections usually require over 100–150 mg/mL), but proteins tend to aggregate at high concentrations, which can affect effectiveness and potentially trigger immunogenicity. Optimizing surface charge distribution and introducing solubility-promoting mutations can help proteins stay monomeric even at high concentrations.


2. Traditional protein optimization: a long, drawn-out 'needle-in-a-haystack' process.

 

Traditional protein optimisation

Traditional protein optimisation

 

The essence of protein optimization is to find variants in a huge sequence space that are "better than the current one." The scale of the problem determines the upper limit of any method.


A protein made up of 300 amino acids, with 20 natural amino acids, has possible sequence combinations of 20 to the power of 300—far exceeding the total number of atoms in the observable universe. You can't test them one by one. So traditional methods have taken two routes.


Rational Design: Based on structural information and physicochemical knowledge, humans judge which site mutations might improve performance. For example, replacing a polar residue buried in a hydrophobic core with a hydrophobic one may increase thermal stability. The problem with this route is that our understanding of the "sequence-structure-function" relationship in proteins is still incomplete. Often, a mutation that looks "theoretically effective" turns out to have no effect or even a negative impact when actually tested.


Directed Evolution: Mimicking natural selection—generate a large variant library with random mutations, then apply selective pressure to pick the best performers, and move them into the next round of mutation and screening. The 2018 Nobel Prize in Chemistry was awarded for this technology. The advantage of directed evolution is that it doesn’t require understanding "why it's good" beforehand—as long as the screening throughput is large enough, the "better" molecules will naturally emerge.


But here’s a mathematical problem that is often ignored: if a protein has 300 amino acids and you randomly introduce a single-point mutation, the possible number of replacements is 300 × 19 = 5700—which is still within screening reach. If you introduce three-point mutations at the same time (which is common in engineering practice), the combination count jumps to about 30 billion—far beyond the throughput of any screening platform. And most performance improvements require the synergistic effect of multiple mutations; the cumulative effect of single-point mutations is often far below that of combined mutations.


3. AI Reshaping Protein Optimization: From Experience-Driven to Data-Driven

 

Faced with the combinatorial explosion dilemma of traditional methods, AI offers a completely different path.


A 2023 review published in *Cell Systems* hits right at the heart of this dilemma: unsupervised machine learning algorithms, trained on massive protein sequence data, can predict antibody variant libraries with 'natural-like' intrinsic properties (such as high stability), significantly reducing the amount of downstream experimental screening. Supervised algorithms, on the other hand, can directly predict variants with multiple target traits by training on deep sequencing data from enriched screening, without additional screening. In short, AI isn’t 'trying'—it’s 'reasoning.' It predicts which mutation combinations are most likely to work based on amino acid sequences, so subsequent experiments only focus on a small set of the most promising candidates.


This 'guided search' works on three levels.


Level one: functional site mapping. Before making any mutations, AI first 'reads' the protein— which residues are in the active site? Which ones maintain the fold? Which positions are highly conserved evolutionarily (indicating they are critical for function)? Which positions are highly variable across species (indicating they might be tweakable 'knobs')? A high-resolution map of functional sites helps engineers avoid those 'break-on-contact' danger zones and focus mutations on safe areas with potential.


Level two: single-point scanning and ranking. After identifying the mutable regions, AI virtually evaluates every possible amino acid replacement at each candidate site—will this mutation improve or damage stability? Enhance or impair catalytic efficiency? Then it outputs a ranked list of mutations based on predicted effects. This step narrows the screening from '19 possibilities per site' to 'a few dozen most worth testing.'


Level three: combinatorial optimization and epistasis modeling. This is the core level where AI surpasses traditional methods. Protein mutations exhibit 'epistasis'—two individually 'beneficial' mutations together might not add up, and could even cancel or harm each other. Traditional methods cannot predict these non-additive effects, but AI models, trained on massive mutation datasets, can capture synergistic or antagonistic relationships between residues, recommending the optimal 'mutation combos' rather than isolated single-point changes.


In this technological paradigm shift, domestic companies have already started productizing and platformizing AI protein optimization capabilities—Shanghai Matwings Technology is one of the pioneers.

 

4. MatwingsVenus™ (Xiaowu ™): Turning protein optimization into a "conversational wet and dry loop"

 

As AI-driven protein optimization moves from academic papers to industrial applications, a key question is: can this capability be "packaged" into a platform service that can be directly accessed by R&D teams across different industries?

This is precisely the core positioning of Shanghai Matwings Technology's MatwingsVenus™ platform ™. In April 2026, Matwings Technology officially launched this conversational protein R&D agent—users optimize goals through natural language input, the system automatically breaks down tasks, schedules structural prediction, functional locus analysis, mutation scanning, and screening, completing the entire process from analysis to design to recommendation. The platform integrates over 200 protein design tools and supports retrieval of tens of billions of real label protein data.

Even more groundbreaking is its "wet and dry closed-loop" model: after the agent completes its design, it can automatically import candidate sequences into the experimental orchestration process, driving the robot to complete sample preparation and functional testing, and then the experimental results are fed back to the next round of AI design—forming a closed-loop iteration of "design → verification→ redesign." In practical projects, this model has been validated: in a certain immune regulatory receptor target project, dozens of novel binding molecules with cell-blocking activity were successfully obtained; In the modification of the sweet protein Monellin, after multiple rounds of iteration, a candidate variant with sweetness increased more than tenfold and heat resistance maintained at about 75°C.

If traditional protein optimization is a "hands-on workshop driven by expert experience and massive experiments," then full-chain AI platforms are transforming it into an engineered process driven by "computational prediction and precise validation."


5. The commercial future of protein optimization

 

The Commercial Future of Protein Optimisation

 The Commercial Future of Protein Optimisation

 

Protein optimization is shifting from a 'science' that relies on experience and luck to an 'engineering' that is predictable and efficient. The deep involvement of AI technology not only greatly lowers the threshold and cost of protein design but, more importantly, upgrades protein product development from 'slow trial and error' to 'efficient and precise design.'


Looking ahead, AI-driven protein optimization will unlock greater commercial value in the following areas:


Green biomanufacturing. As the 'dual carbon' goals advance, more and more traditional chemical processes are being replaced by biocatalysis. Plastic-degrading enzymes (like PET hydrolases), biofuel enzymes (like cellulases), and enzymes for synthesizing bio-based material monomers all need intensive optimization to be economically viable in industrial scenarios. AI protein optimization is significantly shortening the cycle of turning these from 'interesting in the lab' to 'useful in industry.'


Personalized biomedicine. From affinity maturation of antibody drugs to designing synthetic receptors in cell therapy, protein optimization is becoming a fundamental technological support for precision medicine. Platform-based design capabilities make it theoretically possible to 'customize protein drugs for specific patient groups.'


Food and agricultural innovation. Alternative proteins, sweet proteins, enzymes for food additives—the demand from consumers is pushing the food industry to find healthier and more sustainable solutions, and protein optimization is an indispensable technological lever in this process.


For companies, improving protein optimization capabilities means faster product launch cycles, lower R&D costs, higher success rates, and stronger market competition barriers. Whether it’s biopharmaceutical companies seeking differentiated innovation, synthetic biology firms exploring new markets, or traditional manufacturers achieving green transformation through enzyme replacement—protein optimization is becoming a core technology engine driving business growth.


Platforms like MatwingsVenus™ (Xiaowu™) are making this once out-of-reach high-end technology accessible through a 'conversational wet-dry loop' model. As the Matwings Technology team says, 'The value of technology lies not in breakthroughs in the lab, but in solving real-world problems.'


And the story of protein optimization is just beginning.