Methods for Modifying Enzyme Substrate Specificity: From 'Lock and Key' to Smart Design
Published on August 19, 2026

Enzymes are nature's most efficient catalysts, and their defining characteristic lies in substrate specificity—natural enzymes typically exhibit higher catalytic efficiency toward specific substrates. Early research used the "lock-and-key model" to describe this recognition relationship, while modern structural biology has confirmed that enzymes undergo induced-fit conformational changes upon substrate binding. This specificity ensures precise metabolic regulation in biological systems, but in industrial applications, we often need to engineer enzymes to catalyze non-natural substrates, expand substrate scope, or enhance preference for specific substrates.
How can we engineer enzyme substrate specificity—modifying this "lock" to accommodate more "keys"? This is not only a core challenge in protein engineering but also a critical technical bottleneck for translating biocatalysis into industrial practice. This article systematically reviews the mainstream approaches to enzyme substrate specificity engineering, from traditional site-directed mutagenesis to AI-driven precision design, presenting the complete technological evolution of this field.
I. The Underlying Logic of Substrate Specificity Engineering

binding pocket recognition
The recognition between enzyme and substrate occurs within the active-site binding pocket. The pocket's shape, charge distribution, hydrophobicity, and hydrogen-bond network collectively determine which substrates can enter and be catalyzed. Therefore, the essence of substrate specificity engineering is to reshape the molecular environment of the binding pocket to better accommodate the target substrate.
This can be achieved through several approaches: enlarging the pocket to accommodate larger substrates, adjusting pocket polarity to accommodate different chemical groups, or remodeling the hydrogen-bond network to stabilize specific transition states. Different engineering strategies correspond to different methodological approaches.

Mechanism of enzyme substrate-specificity remodelling
II. Classical Methods: Site-Directed Mutagenesis and Saturation Mutagenesis
Site-directed mutagenesis is the most fundamental approach to substrate specificity engineering. Based on understanding of enzyme structure and catalytic mechanism, researchers select specific residues in the active center or binding pocket for substitution, thereby altering the pocket's physicochemical properties. A classic example comes from glutamate dehydrogenase research: through site-directed saturation mutagenesis, two key pocket residues, V143 and A145, were identified. The A145G/V143G double mutant achieved a specific activity increase from 0.14 U/mg to 137.42 U/mg against a model substrate—an improvement of nearly 1000-fold.
Saturation mutagenesis extends this concept further—instead of replacing a single amino acid, one or multiple sites are simultaneously mutated to all 20 natural amino acids, creating a "compact but comprehensive" mutant library for screening. In polyphenol oxidase substrate scope expansion research, researchers first established sequence-function relationships through rational analysis, then used saturation mutagenesis to systematically replace key residues, successfully expanding the enzyme's substrate range. Additionally, fusion expression of the target protein with streptococcal protein G (SpG) can simultaneously improve both expression yield and specific activity.

Three classical technical routes for specificity engineering
III. Directed Evolution: Structure-Independent "Darwinian" Engineering
When three-dimensional structural information is lacking, directed evolution is the most powerful tool. It mimics natural selection—random mutagenesis, recombination, and screening—to rapidly evolve enzymes with novel properties in the laboratory.
In L-threonine aldolase Cβ-stereoselectivity engineering, researchers combined directed evolution with high-throughput screening to improve the wild-type de value from 0.98% to 71.9%. In cytochrome P450BM3 engineering, directed evolution-derived mutants achieved hydroxylation catalysis of non-natural substrates.
The core advantage of directed evolution is its independence from structural information, but its bottleneck is equally clear: the limited coverage of random mutant libraries, with screening throughput often constraining efficiency.
Q: Can substrate specificity engineering be done without the enzyme's 3D structure?
Absolutely. Directed evolution is the preferred approach—it does not require structural information and can discover beneficial mutations through random mutagenesis combined with high-throughput screening. Additionally, protein language model-based AI methods can also perform functional prediction and mutation design directly from sequence, without requiring structural information either.
IV. Computational-Aided Rational Design: Structure-Driven Precision Engineering
With advances in structural biology and computational power, computational-aided rational design is becoming a mainstream approach.
The Rosetta-driven evolution framework is a typical success story. In engineering the substrate specificity of meso-diaminopimelate dehydrogenase (DAPDH), researchers used Rosetta design to guide the construction of "compact but comprehensive" mutant libraries, requiring only 1-2 rounds of evolution and screening of 100–1000 transformants to complete the engineering. Two DAPDH variants achieved activities against model substrates from undetectable levels to 29.5 U/mg and 5.1 U/mg, respectively, with activity toward benzoylformic acid nearly 10-fold higher than previously reported best mutants.
The NAC4ED high-throughput computational platform represents another technological direction. Based on the "near-attack conformation" design strategy, it achieves high-throughput systematic computation of enzyme mutants through automated protein model construction, molecular dynamics simulation, and active conformation population analysis. In validation with 40 mutations, the prediction accuracy reached 92.5%, while the computational time per mutant was only 1/764 of the experimental approach.
Physics-guided computational tools are also advancing rapidly. SubTuner can automatically design enzymes to catalyze specified non-natural substrates, demonstrating superior capability in accelerating the discovery of functionally enhanced mutants across multiple rounds of experiments.
Q: With structural information available, which approach is most efficient?
Computational-aided rational design is the most efficient path. The Rosetta-driven evolution framework requires only 1-2 rounds of experimentation and screening of 100–1000 transformants to complete the engineering—far fewer than the screening scale required for directed evolution.
V. AI and Machine Learning: The Paradigm Shift from "Prediction" to "Generation"

AI-driven paradigm shift for substrate specificity design
AI and machine learning are fundamentally reshaping the substrate specificity engineering paradigm—moving from "how to engineer" toward "what to engineer."
Protein language model-assisted directed evolution is a typical path for AI-empowered substrate specificity engineering. In cyclodextrin glucanotransferase engineering, researchers used the Pro-PRIME protein language model to simultaneously optimize three catalytic properties based on very limited beneficial mutation data. Among 68 variants screened, the optimal mutant achieved a 12-fold improvement in transglycosylation/hydrolysis ratio, with yield increasing from 63% to 98%. In transaminase engineering, researchers integrated machine learning, rational design, and directed evolution, improving the wild-type conversion rate from only 4% against 20 mM substrate to 95% in the mutant, with a melting temperature reaching 77.6°C.
AI-driven autonomous enzyme engineering platforms go even further. A 2025 study published in Nature Communications reported a general-purpose autonomous enzyme engineering platform integrating machine learning, large language models, and biofoundry automation, capable of completing the entire enzyme engineering workflow without human intervention. In proof-of-concept validation, the platform improved the substrate preference of Arabidopsis halomethyltransferase by 90-fold and ethyltransferase activity by 16-fold, requiring construction and characterization of fewer than 500 variants within 4 weeks.
AI prediction models for enzyme substrate specificity have also achieved breakthroughs. The EZSpecificity model, published in Nature in October 2025, combines cross-attention mechanisms with SE(3)-equivariant graph neural networks, achieving 91.7% accuracy in top-pairing predictions, helping researchers rapidly identify optimal enzyme-substrate combinations.
AI-guided enzyme specificity engineering in metabolic engineering applications. Researchers developed the EKAD method integrating the ESM-2 protein language model, the DLKcat catalytic rate prediction model, AlphaFold2 structure modeling, and molecular docking calculations. Through just a single round of screening, they identified a single-point mutation (V365E) in fumarate hydratase (FHase) that significantly enhanced enzyme specificity toward succinate synthesis. Building on this, the research team further reconstructed the reductive tricarboxylic acid (rTCA) cycle pathway in Kluyveromyces marxianus and optimized fermentation through ORP control strategies, ultimately achieving succinate production of 103.9 g/L—a more than 1000-fold improvement over the wild-type strain. This case comprehensively covers the entire metabolic engineering chain from AI-driven enzyme specificity engineering to metabolic pathway reconstruction to fermentation process optimization.
Q: How much initial data does AI-assisted engineering require?
Protein language model-assisted approaches require only a small amount (10–50) of beneficial mutation data to initiate; AI autonomous engineering platforms can start from zero and accumulate data through active learning in iterative cycles; deep learning prediction models require large pre-training datasets, but users only need to input the target sequence to obtain prediction results.
VI. Industrial Implementation: Matwings Technology and MatwingsVenus™ (Xiaowu™)
The industrialization of AI protein design technologies is accelerating. Matwings Technology's independently developed conversational protein R&D agent MatwingsVenus™ (Xiaowu™), integrates billion-scale real-labeled protein data retrieval, over 200 protein design tools, and more than 30 domain-expert-tuned skills. In substrate specificity engineering scenarios, users simply describe their target functional requirements in natural language, and the platform automatically orchestrates AI-directed evolution, AI enzyme discovery, and other core capabilities to complete the full-chain intelligent R&D workflow from sequence analysis and mutation design to performance prediction.
The platform has integrated AI protein design models including ProteinMPNN and SolubleMPNN, compressing traditional protein R&D timelines from 2–5 years down to 2–6 months. In multi-objective synergistic engineering of enzyme substrate specificity, activity enhancement, and stability optimization, the platform provides a complete "design-to-validation" closed loop, significantly lowering the R&D barriers and experimental costs of enzyme engineering.
VII. Summary
The methodological repertoire for enzyme substrate specificity engineering is now quite comprehensive. The choice of approach depends on the specific scenario: with high-resolution structures and clear targets, rational design and computational methods are optimal; in the absence of structural information, directed evolution is the preferred route; with a small amount of beneficial mutation data already available, protein language model-assisted approaches are particularly efficient; and when pursuing maximum efficiency, AI autonomous engineering platforms offer the fastest R&D cycles.
From "point-by-point" site-directed mutagenesis, to the "random exploration" of directed evolution, to AI-driven "precision design," enzyme substrate specificity engineering is undergoing a profound paradigm transformation. Traditional methods excel at "engineering the known," while AI methods are enabling "design of the unknown"—not only telling us "how to engineer" but also predicting "where to engineer" and "what to engineer it into." With the continued iteration of protein language models, graph neural networks, and other AI technologies, and the maturation of industrial-grade AI platforms such as MatwingsVenus™ (Xiaowu™) , enzyme substrate specificity engineering is evolving from a "craft" into a "science"—providing an increasingly powerful technological foundation for industrial biocatalysis, synthetic biology, and pharmaceutical synthesis.