UniProt Protein Search Tutorial: What Exactly Should You Type in the Search Box?
Published on August 24, 2026

Note: The operations and data described in this article are based on the UniProt 2026 release. For the most current interface and functionality, please refer to the official website.
In protein engineering and bioinformatics research, querying protein sequence and functional information is one of the most fundamental and essential tasks. Whether studying the function of an unfamiliar protein, searching for homologous sequences, or laying the groundwork for subsequent experimental design, UniProt is the indispensable first stop.
UniProt (Universal Protein Resource) is the world‘s leading, high-quality, and freely accessible resource for protein sequence and functional information. It is jointly maintained by the UniProt Consortium, which comprises the European Bioinformatics Institute (EMBL-EBI), the Swiss Institute of Bioinformatics (SIB), and the Protein Information Resource (PIR) at Georgetown University.
I. The Core Components of UniProt: Understanding What You Are Searching For

Core components of the UniProt database
Before performing a query, it is essential to understand UniProt’s data structure. UniProt consists of multiple interconnected databases, the most central of which is UniProtKB (UniProt Knowledgebase).
UniProtKB comprises two complementary sections:
UniProtKB/Swiss-Prot (Reviewed): Entries in this section are manually curated by expert curators. Curators systematically review the scientific literature and integrate computational analyses to add high-quality annotations to each protein record, including functional descriptions, domain information, subcellular localization, disease associations, post-translational modifications, and more. Swiss-Prot data are highly reliable and richly annotated, making it the preferred data source for functional studies. As of June 2026, Swiss-Prot contains 575,503 manually reviewed protein sequences.
UniProtKB/TrEMBL (Unreviewed): Entries in this section are automatically annotated by computational systems and have not yet undergone manual review. TrEMBL data are primarily derived from translations of coding sequences in nucleotide sequence databases, offering extremely broad coverage; however, annotation quality varies, and users must exercise their own judgment when using these entries.
Important Update: Since 2026, UniProt has implemented a major shift in TrEMBL‘s inclusion strategy—moving from "comprehensive inclusion" to “selected inclusion,” retaining only entries with experimental evidence or biological importance, as well as reference proteome sequences. A large number of entries from unclassified species and metagenomic sources have been transferred to the UniParc archive. Consequently, the total number of TrEMBL entries has decreased significantly from its previous peak of approximately 250 million.
In short, use Swiss-Prot for detailed functional studies, TrEMBL for large-scale screening, and both in combination for the most comprehensive coverage.
Beyond UniProtKB, UniProt provides several other important datasets:
Proteomes: Complete proteome datasets organized by species, covering thousands of organisms from bacteria to humans.
UniRef: Non-redundant databases clustered by sequence similarity (available at three levels: UniRef100, UniRef90, and UniRef50), used to improve the efficiency of large-scale searches and comparisons.
UniParc: A sequence archive that stores every protein sequence that has ever appeared, with each sequence stored only once. It is the primary source for finding historical sequences and entries that have been removed—a role that has become particularly important following the 2026 TrEMBL streamlining.
II. Basic Queries: Starting from the Search Box
2.1 The Simplest Way to Search
On the UniProt homepage, you can enter the following directly into the search box at the top for a quick search:
Protein names, e.g., insulin, p53, interleukin-6 (IL-6)
Gene names, e.g., BRCA1, TP53
UniProt accession numbers, e.g., P01308 (the UniProt ID for insulin)
Species names, e.g., human, mouse
Keyword combinations, e.g., “kinase human”
Simply enter your keywords and click the search button or press Enter.
Tip: You can use field prefix syntax directly in the search box (e.g.,"gene: tp53 organism:“homo sapiens”) to quickly narrow your search without needing to navigate to the advanced search page.
2.2 Quick Filters on the Search Results Page
The search results page features a filter panel on the left side that allows you to quickly narrow down your results. You can filter by species to view proteins from a specific organism only; filter by data source to display only high-quality, manually reviewed entries; and filter by sequence length, molecular weight, and other physicochemical properties.
2.3 Understanding the Search Results List
Each record in the search results list contains key information:
Entry: The unique UniProt accession number for the protein (e.g., P01308)
Entry Name: A brief identifier for the protein (e.g., INS_HUMAN)
Protein names: The full name(s) of the protein
Gene names: The name(s) of the gene(s) encoding the protein
Organism: The species from which the protein originates
Length: The length of the amino acid sequence
Note: Entry (accession number) and Entry Name are two different identifiers—confusing them may lead to failed searches. The accession number is a stable, unchanging unique identifier and is recommended as the primary search term.

Workflow for UniProt protein query
III. The Protein Details Page: Reading a Protein Entry
After clicking an entry name to enter the details page, the navigation menu on the left allows you to quickly jump to different information sections. These sections contain nearly all known information about the protein:
Function: Describes the protein‘s biological functions, including the metabolic pathways it participates in, the reactions it catalyzes, and the molecules it binds.
Names & Taxonomy: Provides protein and gene names, synonyms, and taxonomic classification information.
Subcellular Location: Shows the protein’s specific location within the cell.
Disease & Variants: Records disease associations and pathogenic mutations related to the protein.
Expression: Shows the expression levels of the gene in different tissues or cell types.
PTM/Processing: Provides information on post-translational modifications (such as phosphorylation, glycosylation) and protein processing.
Interaction: Aggregates protein-protein interaction data from databases such as IntAct, including binary interactions, complex member relationships, and interaction evidence
Structure: Provides three-dimensional structural information for the protein, integrating experimentally determined structures from the PDB as well as predicted structural models from sources such as AlphaFold and SWISS-MODEL, with support for interactive viewing.
Family & Domains: Displays the protein family and domain information.
Sequence: Provides the amino acid sequence of the protein, including all isoforms.
IV. Advanced Search: Precisely Locating Your Target Protein
When a basic search returns too many results, the advanced search can help you pinpoint your target precisely.
Click the Advanced link next to the search box to access the advanced search page. The advanced search supports building complex logical queries using AND, OR, and NOT to combine multiple conditions; using specific fields such as gene name, protein name, and species; and filtering by data status, such as displaying only Swiss-Prot entries.
For example, to search for “manually reviewed proteins encoded by the human BRCA1 gene,” you can achieve precise targeting by combining three conditions: gene name, species name, and curation status.
The core logic of advanced search is the “field: value” pairing pattern. Once you master a few commonly used fields, you can construct extremely precise queries.
V. UniProt‘s Four Built-In Analysis Tools

Built-in analysis tools within UniProt
Beyond basic query functions, UniProt provides four powerful protein sequence analysis tools.
5.1 BLAST: Sequence Similarity Searching
BLAST (Basic Local Alignment Search Tool) is the most commonly used sequence alignment tool. On UniProt, you can access it by clicking BLAST in the toolbar at the top of the homepage.
To use it, paste your protein or nucleotide sequence into the input box, select the target database, and click Run. BLAST results are saved for seven days and can be accessed at any time.
When selecting a database, UniProtKB/Swiss-Prot is best for finding functionally characterized, reliably annotated homologous proteins; UniProtKB/TrEMBL is suitable for searching a broader sequence space to find distantly related homologs. The key difference between the two lies in annotation quality—Swiss-Prot is manually reviewed and highly reliable, while TrEMBL offers broader coverage but annotations are not manually verified.
Key points for interpreting BLAST results include: E-value—the lower the value, the more significant the match; sequence identity—reflects the degree of similarity in the aligned regions; and attention should be paid to matches with low E-values, high coverage, and high identity.
5.2 Align: Multiple Sequence Alignment
The Align tool supports multiple sequence alignment of protein or nucleotide sequences using mainstream alignment engines such as Clustal Omega. Use cases include aligning a target protein with multiple homologous sequences to identify conserved residues and functionally critical regions.
5.3 Peptide Search: Short Peptide Matching
The Peptide Search tool allows you to submit a short peptide sequence and search UniProtKB for all proteins containing that peptide. The primary use case is identifying the source protein of a peptide detected by mass spectrometry.
5.4 ID Mapping: Cross-Database Identifier Conversion
Different databases may use completely different identifiers for the same protein—UniProt accession, Gene ID, PDB ID, Ensembl ID, and so on.
The ID Mapping tool supports batch submission of large numbers of identifiers (up to approximately 100,000 IDs per submission, depending on the mapping type) for unified mapping to UniProt entries or external databases. Use cases include: converting a list of gene names to UniProt accession numbers; mapping UniProt IDs to PDB IDs to obtain structural information; and batch retrieval of sequences or annotations for multiple proteins.
Practical Tip: Specifying the species name during ID Mapping can significantly reduce irrelevant results and improve mapping accuracy.
VI. Recent Features: Important UniProt Updates
In recent years, the UniProt website has undergone multiple major revisions and introduced several new features:
Tools Dashboard. Following the website redesign, tool access has been unified. The Tools Dashboard allows users to view and manage their BLAST, Align, Peptide Search, and ID Mapping tasks in a single interface, with support for saving favorites and resubmitting jobs.
Continuous Improvement in TrEMBL Automatic Annotation Quality. Multiple computational methods, including machine learning models, have been introduced to assist with function prediction and domain annotation, raising the overall quality of automatic annotation.
Expansion of Reference Proteomes. The reference proteome dataset continues to expand to better capture biodiversity.
Enhanced Structure Visualization. In the Structure section, the display of AlphaFold-predicted structures has been enhanced, supporting more convenient interactive viewing and analysis.
Major Revision of TrEMBL Inclusion Strategy. Since 2026, UniProtKB/TrEMBL has shifted from "comprehensive inclusion" to "selected inclusion"—retaining only entries with experimental evidence or biological importance, as well as reference proteome sequences. A large number of entries from unclassified species and metagenomic sources have been transferred to the UniParc archive. This revision has significantly reduced TrEMBL redundancy, meaning that some sequences that could previously be found in TrEMBL now need to be searched for in UniParc.
VII. Conclusion
The core approach to querying proteins in UniProt is: first, locate the target entry through basic or advanced search; then, obtain comprehensive information through the details page; and finally, perform in-depth analysis using the built-in tools.
UniProt offers differentiated query pathways for different research scenarios. For quick queries of known proteins, simply enter the name or accession number directly into the search box to locate the protein. When searching for proteins from a specific species, you can quickly narrow the scope after a basic search using the species filter on the left side. For multi-condition precise targeting, advanced search with field combinations is the best choice. For function prediction of novel sequences, BLAST homology search combined with high-quality Swiss-Prot annotations can complete the inference. When batch data conversion is needed, the ID Mapping service can efficiently complete the task. And for multi-sequence conservation analysis, the Align multiple sequence alignment tool is highly practical.
From the most basic search box to advanced retrieval, from BLAST to ID Mapping, UniProt provides a complete toolchain for protein sequence and functional information queries—taking you from "finding something" to "finding precisely" to “finding everything.” Once you have mastered these methods, you will be able to quickly locate the specific protein you need among hundreds of millions of protein sequences, laying a solid data foundation for subsequent protein engineering research.