Can AI Virtual Screening Compress Industrial Enzyme R&D Cycles from Years to Months?
Published on July 15, 2026

AI virtual screening is becoming a core driving force in industrial enzyme R&D. Traditional enzyme engineering follows the DBTL cycle but is constrained by bottlenecks including the vastness of sequence space, insufficient screening throughput, and difficulties in multi-objective co-optimization, resulting in R&D cycles typically spanning 2 to 5 years. AI virtual screening, by leveraging deep learning models to establish correlations between protein sequence, structure, and function, can rapidly identify high-potential candidate variants within computational space, steering enzyme R&D from "needle-in-a-haystack" random screening toward "precision-guided" intelligent design.
I. AI Virtual Screening: Conceptual Definition and Value Positioning in Enzyme Engineering
AI-driven virtual screening is a technological paradigm that employs artificial intelligence algorithms to rapidly identify candidate molecules with target functional potential from massive pools of protein sequences and structural variants within computational space. In the context of enzyme engineering, its core task is to construct end-to-end sequence–function prediction models for multi-dimensional performance metrics—including hydrolytic activity, thermostability, and pH tolerance—thereby effectively reducing the candidate library size required for traditional experimental screening.
Enzymes are the core components of biological catalysis, and their performance evaluation follows standardized systems. According to the IUBMB enzyme classification and nomenclature rules, one International Unit (IU) of enzyme activity is defined as the amount of enzyme that catalyzes the conversion of 1 micromole of substrate per minute under standard conditions of temperature, pH, and substrate concentration. Specific activity, expressed in U/mg protein, is a core metric for assessing enzyme preparation purity and catalytic efficiency. Catalytic efficiency is commonly characterized by kcat/KM, which comprehensively reflects the enzyme's catalytic turnover capability and substrate recognition ability. Hydrolases—including lipases, proteases, amylases, and PET depolymerases—all require substrate-specific assay systems tailored to their respective substrates.
The core value of AI virtual screening lies in advancing functional evaluation—traditionally dependent on wet-lab trial and error—to the computational pre-screening stage. Deep learning models uncover intrinsic correlations between protein sequence, structure, and function, enabling the prioritization of a small number of variants with the greatest potential for improved hydrolytic activity from virtual mutant libraries comprising hundreds of thousands to millions of candidates for subsequent experimental validation, driving a significant paradigm shift in R&D approaches.
II. Bottlenecks in the Traditional Enzyme R&D DBTL Cycle: The Screening Throughput Challenge

DBTL Cycle with Efficiency Bottleneck
Traditional enzyme engineering employs the DBTL (Design–Build–Test–Learn) cycle: designing mutation schemes, constructing mutant libraries, assaying enzyme activity, summarizing data patterns, and iteratively optimizing. This paradigm has yielded numerous successful industrial enzyme engineering outcomes over more than three decades, yet its inherent structural bottlenecks continue to intensify.
1. The "curse of dimensionality" in sequence space. A typical enzyme comprising 300 amino acid residues has 19 possible mutations at each position, yielding a theoretical combinatorial space of 19³⁰⁰—far beyond the reach of any experimental screening system. Even with state-of-the-art high-throughput screening, a single round can only test 10⁴–10⁷ variants, representing an infinitesimally small fraction of the theoretical sequence space.
2. High development costs and lengthy timelines for screening systems. Establishing screening systems for specific substrates and reaction conditions is a laborious process—coupling reaction design, biosensor construction, and signal detection optimization often demand months of effort.
3. Conflicts in multi-performance co-optimization. Industrial enzymes must simultaneously satisfy multiple requirements: high hydrolytic activity, thermostability, pH tolerance, and high expression yield. Traditional directed evolution can only apply selective pressure on a single metric at a time, and multi-objective optimization requires parallel experimentation across multiple rounds, further extending R&D timelines.
4. Data fragmentation creates data silos. In traditional R&D, sequence–activity data are scattered without systematic integration and modeling, making iterative learning heavily dependent on individual researcher experience and hindering data-driven optimization.
The compounding of these bottlenecks results in traditional industrial enzyme R&D cycles of 2–5 years, with ultimate success rates heavily reliant on the accumulated experience of R&D personnel.
III. Technical Principles of AI Virtual Screening: A Paradigm Shift from Empirical Trial-and-Error to Data-Driven Approaches

Active Site Precision Design
AI virtual screening is driving enzyme engineering from "experiment-driven" toward "data-driven" approaches, with a core technical architecture comprising four layers:
1. Deep learning modeling of sequence–function mapping.
Leveraging massive protein pre-training corpora containing billions of protein sequences (including substantial metagenomic diversity) and millions of functional experimental annotations, deep learning acquires statistical associations between sequence features and functional parameters through pre-training. Upon input of any enzyme variant sequence, the model outputs predictions for multiple performance metrics—including hydrolytic activity, thermostability, and expression yield—within milliseconds. The core advantage lies in its strong generalization capability: requiring only minimal or even zero experimental data from the target enzyme family (few-shot/zero-shot), it can provide valuable predictive rankings for unknown variants across multiple performance dimensions, effectively narrowing the experimental screening scope.
2. Structure generation and precise active-site evaluation.
Hydrolytic activity is highly dependent on the three-dimensional conformational arrangement of the active site, and generative AI has brought breakthroughs to protein structure design. RFdiffusion, a flow-matching diffusion model, can generate protein structures carrying specified functional motifs from random noise. PLACER evaluates atomic-level conformational ensembles of active sites on RFdiffusion-generated backbones, predicting whether the spatial arrangements of catalytic residues and substrates at each step of the catalytic reaction satisfy chemical and geometric constraints, thereby screening candidates with high active-site preorganization and multi-step catalytic capability from a vast pool of designs. In PDB benchmark tests, it achieves sub-angstrom accuracy for small-molecule and side-chain conformation predictions. In final experimental validation, the best-designed crystal structure exhibited an active-site all-atom RMSD of only 0.54 Å and a Cα backbone RMSD <1 Å relative to the design model.
A research team published a study in Science in February 2025 demonstrating the combined application of RFdiffusion and PLACER for de novo design of serine hydrolases: relying solely on conserved active-site motif constraints, they obtained artificial enzymes capable of efficient ester hydrolysis. The highest experimentally determined catalytic efficiency (kcat/KM) reached 2.2 × 10⁵ M⁻¹s⁻¹, with protein crystal structures matching design models (Cα RMSD <1 Å). The high-activity enzymes obtained spanned multiple distinct protein folding topologies, significantly expanding the achievable backbone space for serine hydrolases.
3. Intelligent multi-objective co-optimization search.
Industrial enzyme R&D requires finding Pareto-optimal solutions across hydrolytic activity, thermostability, pH tolerance, and expression yield. AI multi-objective optimization algorithms can batch-screen mutant variants simultaneously satisfying multiple performance thresholds within massive sequence spaces. A 2025 study published in PNAS demonstrated that a single round of AI design yielded optimized sequences with 10- to 20-fold improvements in catalytic efficiency and approximately 10°C enhancements in thermostability.
4. The dry–wet closed-loop data iteration flywheel.
The dry–wet closed-loop iteration system represents the mature implementation form of AI virtual screening. Real activity data generated from each round of wet-lab experimentation flow back to the AI model, correcting prediction biases and updating model parameters, continuously improving the accuracy of subsequent virtual screening rounds and progressively forming a self-improving iteration mechanism.
IV. MatwingsVenus™ (晓鹜™): An Industrial Deployment Platform for AI Virtual Screening
As AI virtual screening transitions from academic theory to industrial application, Matwings has developed MatwingsVenus™ (晓鹜™), China's first conversational, agent-centric one-stop protein R&D foundation model platform.
1. Platform positioning and core capabilities.
The platform is positioned as a dry–wet closed-loop protein R&D infrastructure encompassing "AI design – automated wet-lab experimentation – expert collaboration," building a full-chain capability system around intelligent agents:
Data layer: Built-in datasets of tens of billions of real labeled protein data, including extreme-environment enzyme sequence resources; this database has received a second-class award in a national data elements competition.
Tool layer: Integration of over 200 professional protein design tools, covering the full computational workflow of structure prediction, sequence engineering, molecular docking, and molecular dynamics simulation.
Skills layer: Equipped with over 30 enzyme engineering-specific functions, supporting wild-type enzyme mining, directed evolution, zero-shot de novo design of hydrolases, and multi-objective co-optimization.
Expert layer: A pool of over 50 platform-certified industry experts providing on-demand professional technical support.
2. Conversational AI virtual screening: lowering the barrier to computational tool adoption.
Users need not master complex computational biology operations; they simply input R&D requirements via natural language (e.g., "design a PET depolymerase with high hydrolytic activity at 65°C"), and the platform automatically decomposes tasks, orchestrates sequence design, performance prediction, and batch screening modules, generating candidate sequence libraries and outputting optimal variants ranked by comprehensive performance.
3. Dry–wet closed loop: bridging virtual screening and automated experimentation.
After the agent completes virtual screening, the platform automatically pushes candidate sequences to plasmid ordering and experimental scheduling modules, driving automated equipment to complete strain construction, protein purification, and batch hydrolytic activity assays. Experimental data automatically flows back to the AI module to initiate the next iteration, forming a complete conversational dry–wet closed loop.
According to Matwings official data, the platform compresses the traditional 2- to 5-year enzyme R&D cycle to 2–6 months, improves engineering success rates from the conventional 0.1%–1% to 30%, and has already delivered over 30 industrial enzyme-related projects.
4. Differentiation from academic models.
RFdiffusion and PLACER focus on protein structure generation and site evaluation, serving as laboratory research tools. MatwingsVenus™ (晓鹜™), by contrast, is oriented toward commercial industrialization, providing—in addition to integrated cutting-edge AI design algorithms—complete infrastructure encompassing data management, virtual screening, automated experimentation, and expert collaboration.
V. Industrial Applications and the Standardized R&D Closed Loop

In Silico-to-Wet Lab Closed Loop
1. Plastic biodegradation: AI virtual screening of PET depolymerases.
PET plastic biodegradation is a core track in carbon-neutral biomanufacturing. AI virtual screening enables high-throughput mining of PET depolymerase candidates from databases with simultaneous accurate prediction of protein structure and degradation performance. A research team developed a TurboPETase variant through rational remodeling of PET-binding groove flexibility—an approach distinct from deep learning/AI methods—achieving complete depolymerization of PET waste under high substrate loading within 8 hours.
2. Fermentation and food biomanufacturing.
Thermostability and hydrolytic activity in enzymes commonly exhibit a trade-off relationship. AI multi-objective optimization can screen for mutant combinations that balance both heat resistance and high activity. A case study on Matwings platform engineering of the sweet protein Monellin demonstrated that, through the "agent design → automated experimentation → data feedback iteration" cycle, multiple mutants achieved more than tenfold sweetness enhancement compared to the wild type while maintaining thermostability around 75°C.
3. Standardized R&D pipeline.
MatwingsVenus™ (晓鹜™) has established a standardized enzyme R&D workflow: natural language requirement input → AI multi-dimensional virtual screening → automated experimental validation → data-driven iterative optimization. This closed loop upgrades enzyme R&D from manual project-based work to standardized assembly-line operations, enabling the scalable, replicable engineering deployment of precisely enhanced hydrolytic activity.
VI. Conclusion and Outlook
AI virtual screening is driving transformation in industrial enzyme R&D paradigms—it is no longer merely an auxiliary tool for directed evolution and rational design, but a core R&D engine spanning the entire workflow of sequence mining, precise active-site engineering, and multi-performance co-optimization.
Academic models including RFdiffusion and PLACER continue to expand the scientific frontiers of de novo enzyme design, while industrial platforms such as MatwingsVenus™ (晓鹜™) translate these scientific capabilities into scalable, reusable industrial R&D infrastructure. From PET waste degradation and industrial fermentation enzyme upgrading to carbon-neutral biomanufacturing, AI virtual screening continues to drive the deployment of high-efficiency, precise, and sustainable novel industrial enzymes.
Looking ahead, as protein annotation databases continue to expand, AI prediction models undergo iterative upgrades, and dry–wet automated closed-loop systems mature, industrial enzyme R&D cycles are expected to shorten further, continuously empowering the synthetic biology and green biomanufacturing industries.