warin.me · Research Lab EN · ไทย
HUMAN PROTEIN SUBCELLULAR LOCALIZATION · 2026

Human locations.
Local evidence.

A human-specific continuation of the PlantLoc systems study: experimentally supported microscopy annotations, local sequence predictors, a PSO-weighted multilabel vote, and visible evidence before weighting.

Why build a human lab?

Clinical and functional context

Location changes interpretation

A variant or expression change can have different consequences depending on whether the protein reaches the nucleus, plasma membrane, mitochondrion or secretory system.

Multi-location biology

One protein, several compartments

Human proteins shuttle, traffic through organelles and redistribute with cell state. A top-1 label can hide valid secondary localizations and context-dependent function.

Reproducible delivery

No remote RPA queue

WoLF PSORT, Light Attention and the animal ngLOC model execute locally in isolated containers. Their evidence is normalized and cached before PSO voting.

Scope. This is a development and systems experiment, not a diagnostic test. Predictions prioritize experimental hypotheses and do not replace microscopy, fractionation or other laboratory evidence.

Human-specific experimental dataset

The benchmark joins Human Protein Atlas v25.1 subcellular annotations of Enhanced or Supported reliability to reviewed Homo sapiens sequences from UniProt. A stable accession hash creates fit, validation and locked-test partitions.

4,222
matched human proteins
fit 2,933 · validation 625 · locked test 664

HPA's fine-grained microscopy structures are collapsed into nine shared compartments: nucleus, cytoplasm, plasma membrane, mitochondrion, extracellular, endoplasmic reticulum, Golgi, peroxisome and lysosome. Sequences are capped at 1,200 residues to define a reproducible CPU-feasible scope; this excludes 492 long-tail records from the initially matched set. The original fine labels remain attributable to HPA; this mapping is an explicit modelling choice.

Local predictors

PredictorLocal modeRoleImportant limitation
Light AttentionProtT5 + published checkpointProtein-language-model evidenceSingle-label probability head; public training overlap may exist
WoLF PSORTAnimal kNN modelSorting-signal and sequence evidenceLegacy model and label vocabulary
ngLOC 1.0Official animal n-gram modelIndependent generative sequence evidenceLegacy Swiss-Prot training set and normalized legacy score output

Why two PSO results?

The objective changes the scientific behaviour. Optimizing exact match produces a conservative voter and raises held-out exact accuracy from Light Attention's 37.05% to 37.80%, but misses several rare compartments. Optimizing macro F1 raises held-out macro F1 from 43.57% to 51.34%, while predicting too many label combinations and reducing exact accuracy to 3.16%. The live panel uses the validation-selected exact-match voter; both outcomes remain visible below.

Prediction panel

Paste one human protein FASTA sequence. The first uncached run can take several minutes because all models execute locally.

Vote output

Raw predictor answers, learned coefficients, weighted contributions and thresholds will appear here.

Locked-test results

Experiment still running; this table will be populated from the frozen artifact.

PSO weights and thresholds

Confusion matrices by localization

Each localization is evaluated as an independent binary decision because a protein may have several valid labels.

Credits and references

  1. Human Protein Atlas subcellular resource and downloadable data, version 25.1.
  2. Thul P.J. et al. (2017), A subcellular map of the human proteome, Science.
  3. Stärk H. et al. (2021), Light Attention, Bioinformatics Advances.
  4. Horton P. et al. (2007), WoLF PSORT, Nucleic Acids Research.
  5. King B.R. and Guda C. (2012 standalone release), ngLOC, BMC Bioinformatics.
  6. The PSO voting lineage follows PSO-LocBact (2019) and the PlantLoc weighted ensemble (2021). Base predictors remain the work of their credited creators.