NAR-2022 DeepLoc 2.0:multi-label subcellular localization prediction using protein language models

· Nucleic Acids Research· · protein, localization, PLM

Paper: DeepLoc 2.0:multi-label subcellular localization prediction using protein language models

Web server: https://services.healthtech.dtu.dk/services/DeepLoc-2.0/

DeepLoc-2.0: Predicting subcellular localization and protein sorting signals

Abstract

Predicting protein subcellular localization is important for proteomics research. DeepLoc-2 is introduced for multi-location prediction. For training and validation, multi-location protein datasets for eukaryotes and humans are curated under strict homology partitioning, together with sorting signal information compiled from the literature. Two approaches improve interpretability: attention outputs along the sequence and highly accurate prediction of nine types of protein sorting signals; attention outputs correlate well with the positions of sorting signals.

Introduction

Identifying where proteins localize in different cellular compartments plays a key role in functional annotation. It also helps identify drug targets and understand diseases linked to aberrant subcellular localization. Some proteins are known to localize to multiple compartments. Biological mechanisms that explain localization involve short sequences called sorting signals.

SwissProt localization dataset

Protein data were extracted from UniProt release 2021_03. Sequences and localization annotations were filtered with the following criteria: eukaryotes; not fragments (fragments may lack N- or C-terminal sorting signals); nuclear-encoded; length >40 amino acids; and experimentally annotated subcellular localization (ECO:0000269). Proteins can be assigned to one or more of ten locations: cytoplasm, nucleus, extracellular, cell membrane, mitochondria, plastid, endoplasmic reticulum, lysosome/vacuole, Golgi apparatus, and peroxisome.

Human protein atlas

The Human Protein Atlas (HPA) project provides subcellular localization of human proteins from confocal microscopy. Annotations carry four reliability labels—Enhanced, Supported, Approved, and Uncertain—based on criteria such as antibody validation in the literature and experimental evidence. Only Enhanced and Supported labels are used for the independent test set, as they are the most reliable. This dataset was constructed with no sequence sharing >30% global sequence identity with the SwissProt dataset above and was used for independent validation.

Sorting signals

Experimentally validated sorting signal annotations were compiled mainly from the literature, excluding proteins absent from the SwissProt localization dataset described above.

DeepLoc 2.0 Overview

avatar
avatar

avatar

avataravatar

<HR align=left color=#987cb9 SIZE=1>