Back to list

Ways to Boost Enzyme Solubility: A Full Breakdown of 7 Core Strategies from Lab to Mass Production

Published on August 23, 2026

Ways to Boost Enzyme Solubility: A Full Breakdown of 7 Core Strategies from Lab to Mass Production

Most people working in enzyme engineering know that frustrating feeling: cloning works, induction is done, but when you finally break the cells open—the target protein mostly ends up in inclusion bodies, with only a tiny amount in the supernatant. This doesn’t necessarily mean your technique is bad; enzyme solubility issues are widely recognized as one of the biggest roadblocks in protein engineering. From basic lab research to industrial biotech mass production, improving enzyme solubility has always been a core issue you can’t avoid.


The methods for regulating protein solubility are very extensive, ranging from optimizing the expression host, modifying the protein sequence, adjusting fermentation conditions, to AI-assisted design nowadays, covering biology, chemistry, and physics. In today’s article, we combine cutting-edge progress in the field with hands-on experience to systematically summarize ways to improve enzyme solubility into a practical seven-step strategy framework.


Why is improving enzyme solubility so important?

For heterologous proteins expressed in E. coli cytoplasm, about 30%–60% form inclusion bodies to varying degrees. The proportion is even higher for eukaryotic enzymes, membrane proteins, and multi-subunit proteins. Inclusion bodies aren’t unusable—you can use denaturation and refolding—but for structurally complex industrial enzymes with multiple domains, the refolding rate is often just a few percent. Of course, simple small proteins (like lysozyme and RNase) can have refolding rates of 30%–60%; in industry, some products (like insulin and certain recombinant proteases) still take the inclusion body route because high expression and resistance to protease degradation can make the refolding process cost-effective once optimized.


But for the vast majority of enzyme engineering projects, soluble expression is still the preferred choice, because it directly affects:

· Experiment timeline: highly soluble enzymes can be purified in one step

· Active enzyme yield: correctly folded proteins usually have complete catalytic activity

· Industrial cost: poor solubility may significantly reduce the amount of usable enzyme

· Downstream applications: crystallography, enzymatic processing, and protein drug development all rely on solubility


Improving enzyme solubility isn’t just a 'nice-to-have' trick—it’s a key skill that determines whether a project can smoothly progress and whether a product can be cost-effectively implemented.


7 Strategies to Improve Enzyme Solubility.

7 Strategies to Improve Enzyme Solubility


Method 1 to Increase Enzyme Solubility: Rational Choice of Expression System and Host Strain

Many people immediately go for BL21(DE3), and when it doesn’t work, they start doubting themselves. Actually, choosing the right host is the first step and also the cheapest optimization path.


In E. coli systems: for enzymes with many rare codons, switching to the Rosetta series often has a noticeable effect (but note that codon optimization isn’t better the more you do it—natural rare codons can play a role in co-translational folding, and blindly replacing them may actually reduce folding quality). For proteins that degrade easily, consider BL21 Star (more stable mRNA) or the Origami series (more oxidative cytoplasm, good for disulfide bond formation).


Auto-induction is often overlooked but very practical: it uses the natural lactose metabolism regulation, so expression starts automatically once the culture reaches a certain density. Induction is milder and more uniform, and in many cases soluble protein yields are higher, which is especially good for batch screening.


Periplasmic expression is another efficient route—direct the target protein to the periplasm, where the oxidative environment and the Dsb system help proper disulfide bond formation. Common signal peptides include pelB, OmpA, MalE, etc. Just adding a short sequence to the N-terminus can surprisingly improve disulfide-dependent enzymes.


Cross-system choice: for eukaryotic enzymes, Pichia pastoris is preferred (has post-translational modifications, and secreted expression avoids cell lysis). Mammalian cells (HEK293, CHO) are costly, but for complex human enzymes or glycosylation-dependent enzymes, they are often a safer bet.


A core principle: higher expression isn’t always better. If a strong promoter drives protein synthesis too fast, the folding machinery can’t keep up, and inclusion bodies form more easily. Lowering the temperature (16–25°C), reducing IPTG concentration, or shortening induction time can surprisingly improve solubility.


Method 2 to Improve Enzyme Solubility: Co-expression with Molecular Chaperones and Folding Enzymes

Molecular chaperones are like 'folding assistants' in the cell. If the endogenous ones aren't enough, you just add exogenous ones.

Common co-expression systems:

· DnaK-DnaJ-GrpE: initial folding of nascent peptide chains

· GroEL-GroES: the classic 'chaperonin barrel,' more effective for large molecules or enzymes with complex folding paths

· Trigger Factor: binds the ribosome exit, assisting folding during translation

· DsbA/DsbC: involved in disulfide bond formation


Practical experience: combination strategies usually work better (e.g., GroEL/GroES plus DnaK/DnaJ/GrpE five-component system). But there are three 'pitfalls' to watch out for:

First, more chaperones isn’t always better — chaperones themselves consume expression resources, and overexpression can actually steal resources from the target protein, leading to lower overall expression.

Second, solubility ≠ correct folding. Studies show that chaperones increase the protein’s 'solubility,' but don't necessarily improve 'conformational quality' or bioactivity — sometimes chaperones just keep misfolded protein soluble but inactive. So always check both solubility and activity after co-expression.

Third, GroEL overexpression might depend on substrate size — the GroEL-GroES cavity naturally fits proteins around 30–60 kDa with complex folding paths (like the TIM barrel family); small proteins (<20 kDa) usually fold quickly on their own and typically don’t need GroEL, so overexpressing GroEL offers limited help for these proteins.

Molecular chaperones aren’t a magic key. If the sequence inherently folds poorly, you need to troubleshoot the protein itself.


Method 3 to Increase Enzyme Solubility: Fusion Tag Strategy — Works Well but Use Wisely

Fusion tags are one of the most complained-about yet commonly used methods to boost enzyme solubility.

Commonly used solubility tags:

· MBP: Known as the 'king of solubility enhancement', most effective but large (~42 kDa)

· GST: Classic tag, moderate solubility improvement, easy to purify

· SUMO: Pretty good solubility effect, efficiently cleavable by proteases, leaves natural N-terminus

· Trx: Small (~12 kDa), works wonders for certain enzymes

· NusA, GB1, etc.: Each has its own suitable scenarios


Besides the classic tags, newer fusion partners like SlyD (chaperone activity protein), Fh8 (small tag), Tsf, RpoS have emerged in recent years. Some can match or even surpass MBP in solubility effect for specific proteins. Short disordered peptide tags are also a new direction — short, charged-rich disordered peptides can significantly improve solubility with minimal interference on the target protein structure.


Regarding the mechanism: It was initially thought to be a 'chaperone-like effect', but later studies found that the surface charge properties of the tag are also crucial — highly net negatively charged (acidic) tags generally show broader solubility enhancement, mainly by reducing non-specific aggregation via intermolecular electrostatic repulsion. However, note that highly positively charged tags may cause precipitation due to binding negatively charged biomolecules like nucleic acids, so choose based on the protein properties.


Two common pitfalls:

1. Tag-dependent solubility: if it’s soluble only with the tag and precipitates once removed — indicates the enzyme itself has folding issues and the tag is just a 'crutch'.

2. Tags may affect enzyme activity: large tags might alter active site conformation or substrate channels, so always check activity after tagging.


Suggestion: Try MBP first to see if solubility improves. If MBP doesn’t help, it indicates serious folding issues that need sequence-level modification.


Method 4 to Increase Enzyme Solubility: Rational Design — Site-specific modification based on sequence and structure

Rational Design.

Rational Design

This is where the real depth of protein engineering really shows.


Core idea: Protein solubility largely depends on the degree of exposed hydrophobic surfaces and charge distribution. If there are large hydrophobic patches or uneven charge distribution on the surface, proteins tend to aggregate and precipitate.


Practical strategies:

· Surface hydrophobic residue mutation: Mutate surface-exposed hydrophobic residues (Leu, Ile, Val, Phe, etc.) to charged residues (Asp, Glu, Lys, Arg) or polar residues. The key is 'surface-exposed'—don’t touch internal hydrophobic residues, or you might disrupt the overall fold.

· Surface charge engineering: Don’t just increase net charge; pay attention to the distribution of charge patches. Even with the same net charge, proteins with evenly distributed charges are more soluble than those with concentrated charges. Eliminating large surface positive/negative charge patches often significantly reduces aggregation tendency.

· Disulfide bond design: Introduce disulfide bonds to rigidify the structure and reduce unfolding and aggregation.

· Truncation constructs: Remove naturally disordered regions (IDRs) at the N/C termini, which are often the 'culprits' of aggregation.

· pH-dependent optimization: Use tools like CamSol to predict solubility curves at different pH values, find optimal buffer conditions, or alter the pH-solubility curve through mutations.


How to know which residues are on the surface? Look at structures in the PDB, or predict with AlphaFold if no structure is available. As for which sites to mutate to improve solubility—that’s where AI tools really shine.


Method 5 to improve enzyme solubility: Directed evolution—the logic of natural selection

Rational design is limited by our understanding of folding mechanisms, while directed evolution takes a different approach: random mutation followed by selection for solubility.


Classic screening systems:

· Folding reporter systems: Fuse the target enzyme with GFP or a resistance gene; if the reporter protein is active, the upstream protein is folded correctly. But note that the GFP system can give false positives—sometimes the target protein precipitates but GFP folds on its own and fluoresces.

· Phage display/yeast surface display: Proteins displayed on the surface are usually properly folded and soluble; after multiple rounds of enrichment, you get positive mutants.

· Ribosome display/mRNA display: In vitro systems with much larger library capacity (10¹²~10¹⁴), not limited by cell transformation.

· Flow cytometry sorting: Combine with fluorescent labeling for high-throughput single-cell screening.


Recently, continuous directed evolution (e.g., PACE) has greatly accelerated the evolution process—traditional rounds take days to weeks, while continuous evolution can complete dozens of rounds in a single day. However, applying continuous evolution directly for solubility evolution is still rare, mainly due to the bottleneck of high-throughput solubility readouts.


The advantage of directed evolution is that it doesn’t rely on prior knowledge and often finds sites that rational design wouldn’t think of. But the downside is obvious: establishing a screening system is often the most time-consuming step, and many enzymes don’t have suitable high-throughput screening methods.


Hence, the combination strategy of 'rational design + directed evolution' is becoming mainstream—first, use computational tools to narrow the mutation range, then construct focused libraries for screening, improving efficiency by several orders of magnitude.


Method 6 to Improve Enzyme Solubility: Optimizing Culture Conditions and Fermentation

Earlier we talked about 'modifying the protein,' but this one is about 'modifying the environment.' Sometimes the protein itself isn’t the problem; it’s the harsh expression conditions.


Temperature: Lowering the temperature is one of the most cost-effective and impactful optimization methods. At low temperatures, protein synthesis is slower, giving the folding machinery more time to work, and thermal inactivation is reduced.


Media and Additives: LB isn’t the only option. TB has richer nutrients and higher cell density, and sometimes solubility is actually better. Adding chemical chaperones to the media is a zero-cost trial-and-error approach:

· Osmoprotectants (sorbitol, betaine)

· Chemical chaperones (5%–10% glycerol, 0.5 M arginine, low concentrations of DMSO)

· Precursors of cofactors for coenzyme-dependent enzymes

· Metal ions for metalloproteins


Induction Conditions: IPTG concentrations from 0.1 mM to 1 mM are worth trying, while lactose induction is generally gentler and more sustained. The timing of induction matters too—OD₆₀₀ = 0.4–0.6 is standard, but sometimes inducing at a higher OD (e.g., OD = 1.0) works better because the cells are more mature and the folding systems more complete, potentially boosting solubility. Cold-shock expression systems (cspA promoter) are also an option.


Oxygen and pH: Low oxygen can lead to abnormal cell metabolism, reducing protein expression quality. pH indirectly affects protein expression by influencing cell metabolism.


Scale-Up for Industry: Conditions that work in a shake flask often need reoptimization when scaling up to a fermenter. Fed-batch fermentation, which controls the rate of carbon source addition to avoid acetate accumulation, usually yields much higher amounts of soluble protein than batch culture.


Method 7 to Improve Enzyme Solubility: AI-Assisted Design — A Highly Efficient Solution Today

The explosion of AI protein design in the past two years has significantly boosted the efficiency of methods to enhance enzyme solubility. A clear trend in the field is that protein solubility regulation is moving from traditional trial-and-error approaches toward computation-driven rational design.


AI solubility prediction tools have gone through several generations:

· First generation (empirical/physical models): Wilkinson-Harrison model, CamSol — based on physico-chemical properties. CamSol remains a classic tool for pH-dependent solubility prediction.

· Second generation (traditional machine learning): SOLpro, Protein-Sol, SoluProt — based on manually designed sequence features.

· Third generation (deep learning): DeepSol, DeepSoluE — using CNNs/LSTMs to learn solubility features directly from sequences.

· Fourth generation (multi-modal fusion): Pro4S in 2025 combined protein language models, structural features, and surface descriptors; SurfSol published in 2026 used E(3)-equivariant graph neural networks to directly represent protein surface properties, combined with ESM-2 sequence embeddings, achieving surface-based multi-modal prediction.

· Fifth generation (from prediction to design): AI no longer just "predicts after the fact" but directly designs highly soluble functional proteins from scratch. Generative models like RFdiffusion and protein language models can already generate completely new protein scaffolds with high solubility.


Core values of AI in enzyme solubility optimization:

· Accurate prediction of mutation effects: success rates far higher than random mutations or purely empirical design

· High-throughput virtual screening: reduces experimental work by more than an order of magnitude

· Multi-objective optimization: balances solubility, activity, stability, and other coupled indicators

· De novo design of soluble enzymes: directly generates functional enzyme scaffolds with high solubility


Many research groups know AI is useful but find it hard to put into practice — not enough computing power, complex toolchains, lack of cross-disciplinary talent. This is one reason why the MatwingsVenus™ (Xiaowu™) agent is valuable.


MatwingsVenus™ (Xiaowu™): An AI research assistant that makes methods to improve enzyme solubility more efficient and practical.

MatwingsVenus™ protein agent

MatwingsVenus™

MatwingsVenus™ (Xiaowu™) is an AI agent platform designed for researchers in protein engineering, enzyme engineering, and synthetic biology. It integrates cutting-edge protein computational toolchains into a unified workflow, so experimenters don’t need to learn programming, set up computing resources, or switch between a dozen different software—they can directly use advanced AI-assisted design.


So, what can MatwingsVenus™ (Xiaowu™) do specifically for improving enzyme solubility?

· Input an enzyme sequence or PDB structure and predict solubility hotspot sites with one click, providing prioritized mutation suggestions.

· Support the entire workflow for rational design and directed evolution: homology modeling, mutation scanning, stability prediction, library design, and result analysis.

· Integrate various specialized tools such as protein structure prediction, molecular docking, functional site prediction, and affinity maturation, covering the complete enzyme engineering process from sequence to function.

· Built-in professional literature and patent search to quickly understand the research status of target enzymes and known modification strategies, avoiding wasted effort.


Many times, it’s not that you’re not working hard—it’s that the tools aren’t convenient enough. While others are blindly testing a single site, you can run the full screening with MatwingsVenus™ (Xiaowu™) in a week and get positive results in a month—that’s the difference in efficiency.


A final note: There’s no universal solution for improving enzyme solubility, but there is a systematic methodology.

No single method can guarantee 100% resolution of all enzyme solubility issues. Every enzyme is unique; other people’s experience can only serve as a reference, not a copy.

It’s worth noting that the idea of synergistic modifications is getting more attention: single methods have limited effect, but combining multiple strategies can act on different molecular levels (hydrophobicity, electrostatics, hydrogen bonding, etc.), often achieving a “1+1>2” effect and reducing side effects from single approaches.


But improving enzyme solubility does have a systematic approach. A proven practical workflow:

1. Zero-cost optimizations first: codon optimization, host switching, lower temperature, media additives—most don’t require extra cloning and can be tested in a week.

2. Tags and co-expression: MBP/SUMO tags, co-expression with molecular chaperones, with preliminary results in two weeks.

3. Sequence modification: AI tools predict hotspot sites, followed by small-scale site-directed mutation validation.

4. Directed evolution: focused libraries guided by rational design, high-throughput screening.

5. Collaborative thinking throughout: consider multiple strategies from the start, rather than trial-and-error linearly.


Following this workflow, the solubility of most enzymes can be improved to varying degrees.


In research, don’t get stuck on a single method. Keep your mind open and use the tools well; many problems that seem unsolvable can have simple solutions if you look at them from a different angle.