NAR-2022 DeepLoc 2.0:multi-label subcellular localization prediction using protein language models
Paper: DeepLoc 2.0:multi-label subcellular localization prediction using protein language models
Web server: https://services.healthtech.dtu.dk/services/DeepLoc-2.0/
DeepLoc-2.0: Predicting subcellular localization and protein sorting signals
Abstract
Predicting protein subcellular localization is important for proteomics research. DeepLoc-2 is introduced for multi-location prediction. For training and validation, multi-location protein datasets for eukaryotes and humans are curated under strict homology partitioning, together with sorting signal information compiled from the literature. Two approaches improve interpretability: attention outputs along the sequence and highly accurate prediction of nine types of protein sorting signals; attention outputs correlate well with the positions of sorting signals.
Introduction
Identifying where proteins localize in different cellular compartments plays a key role in functional annotation. It also helps identify drug targets and understand diseases linked to aberrant subcellular localization. Some proteins are known to localize to multiple compartments. Biological mechanisms that explain localization involve short sequences called sorting signals.
SwissProt localization dataset
Protein data were extracted from UniProt release 2021_03. Sequences and localization annotations were filtered with the following criteria: eukaryotes; not fragments (fragments may lack N- or C-terminal sorting signals); nuclear-encoded; length >40 amino acids; and experimentally annotated subcellular localization (ECO:0000269). Proteins can be assigned to one or more of ten locations: cytoplasm, nucleus, extracellular, cell membrane, mitochondria, plastid, endoplasmic reticulum, lysosome/vacuole, Golgi apparatus, and peroxisome.
Human protein atlas
The Human Protein Atlas (HPA) project provides subcellular localization of human proteins from confocal microscopy. Annotations carry four reliability labels—Enhanced, Supported, Approved, and Uncertain—based on criteria such as antibody validation in the literature and experimental evidence. Only Enhanced and Supported labels are used for the independent test set, as they are the most reliable. This dataset was constructed with no sequence sharing >30% global sequence identity with the SwissProt dataset above and was used for independent validation.
Sorting signals
Experimentally validated sorting signal annotations were compiled mainly from the literature, excluding proteins absent from the SwissProt localization dataset described above.
DeepLoc 2.0 Overview





<HR align=left color=#987cb9 SIZE=1>