How to Discover Proteins

Billion-scale Search Space, One-Step Protein Discovery

Transform structural biology intuition into the starting point of AI search. Expand unprecedented search boundaries for your enzyme engineering.

蛋白发现-en

1. Case Study (How to Discover Your Protein?)

1.1 Follow the Prompted Dialogue

1.1.1 Input Requirement

I want to use UniProt ID P20261 (Candida rugosa lipase this is the protein’s ID) as a template to search for candidates with stronger lipolytic activity in the default database.

Reference Link

Please do not navigate to the homepage from this link

屏幕截图 2026-04-24 084055

The assistant establishes a workflow. You can supplement information to initiate execution.

image

1.1.2 Additional Information

Set 70% similarity as the threshold, apply multi-model cross-validation, and return 30 candidate amino acid sequences.

PixPin_2026-04-24_08-43-00

Once complete information is provided, the assistant begins execution.

image

Note: Sometimes the assistant will request feedback after completing a step to improve subsequent tasks (Tip: add "directly output final results" to skip interruptions).

proceed with the multi-model cross-validation step

After retrieval, the assistant outputs the final filtered results.

image

1.2 Direct Instruction Input

Provide all requirements at once, and the assistant will directly predict and output results.

I want to use UniProt ID P20261 (Candida rugosa lipase this is the protein’s ID) as a template to search for candidates with stronger lipolytic activity in the default database, with a similarity threshold of 70%. Use multi-model cross-validation. Provide 30 candidate amino acid sequences. Directly output final results.

Reference Link

PixPin_2026-04-24_08-58-19

The assistant directly outputs prediction results.

image

image

2. What Can MatwingsVenus™ Do?

  1. Protein structure prediction and visualization (generate 3D structures):

    1. ESMFold (Fast Mode): Fast with moderate accuracy (≤500 amino acids), charged per run.

    2. AlphaFold (High-Accuracy Mode): Slower but highly accurate (≤2000 amino acids), charged by runtime.

  2. Function and property prediction:

    1. Solubility: Ability of a protein to dissolve in a given environment

    2. Subcellular Localization: Location within a cell (e.g., nucleus, cytoplasm)

    3. Membrane Protein: Whether the protein is membrane-associated

    4. Metal Ion Binding: Ability to bind metal ions

    5. Stability: Ability to maintain structure and function

    6. Sorting Signal: Signal directing protein transport

    7. Optimum Temperature: Temperature of maximum activity

    8. Kcat: Catalytic efficiency

    9. Optimal pH: pH of maximum activity

    10. Immunogenicity Prediction - Virus: Immune activation potential for viruses

    11. Immunogenicity Prediction - Bacteria: Immune activation potential for bacteria

    12. Immunogenicity Prediction - Tumor: Immune activation potential for tumors

  3. Evolutionary Visualization: Generate phylogenetic trees to show relationships between sequences.


3. Input Tips (How to Improve Accuracy?)

3.1 Provide Detailed Template Information

  • Upload protein 3D structure (PDB) for best results

  • Input UniProt ID or sequence to auto-predict structure

  • Use built-in large-scale metagenomic data

蛋白发现-附件-en

3.2 Specify Your Goal

  • Desired function (e.g., hydrolysis, catalysis)

  • Source requirements (e.g., organism, tissue)

  • Desired properties (e.g., thermostability, solubility, activity)

蛋白发现-目标-en

3.3 Add Filtering Criteria

  • Set similarity thresholds (e.g., 70% / 80%) and number of outputs

  • Use multi-model validation for reliability

  • Or follow step-by-step prompts


4. Mode Selection

The platform provides three working modes:

  • Fast Mode: Lightweight intelligent retrieval with efficient output

  • Thinking Mode: Handles complex tasks with deeper reasoning

  • Thinking Mode Pro: Advanced reasoning and cross-domain problem solving


5. Model Overview

  1. VenusMineEnzyme Discovery. Discover novel enzymes beyond sequence similarity using structural insights. [Paper] [Code]

  2. VenusX SeriesProtein Site Prediction. Identify active sites, binding regions, and functional residues. [Paper] [Code]

  3. VenusG SeriesProtein Property Prediction. Analyze protein functions and properties. [Paper] [Code]


Platform design and commercial copyright belong to Tianwu