Location changes interpretation
A variant or expression change can have different consequences depending on whether the protein reaches the nucleus, plasma membrane, mitochondrion or secretory system.
A human-specific continuation of the PlantLoc systems study: experimentally supported microscopy annotations, local sequence predictors, a PSO-weighted multilabel vote, and visible evidence before weighting.
A variant or expression change can have different consequences depending on whether the protein reaches the nucleus, plasma membrane, mitochondrion or secretory system.
Human proteins shuttle, traffic through organelles and redistribute with cell state. A top-1 label can hide valid secondary localizations and context-dependent function.
WoLF PSORT, Light Attention and the animal ngLOC model execute locally in isolated containers. Their evidence is normalized and cached before PSO voting.
Scope. This is a development and systems experiment, not a diagnostic test. Predictions prioritize experimental hypotheses and do not replace microscopy, fractionation or other laboratory evidence.
The benchmark joins Human Protein Atlas v25.1 subcellular annotations of Enhanced or Supported reliability to reviewed Homo sapiens sequences from UniProt. A stable accession hash creates fit, validation and locked-test partitions.
HPA's fine-grained microscopy structures are collapsed into nine shared compartments: nucleus, cytoplasm, plasma membrane, mitochondrion, extracellular, endoplasmic reticulum, Golgi, peroxisome and lysosome. Sequences are capped at 1,200 residues to define a reproducible CPU-feasible scope; this excludes 492 long-tail records from the initially matched set. The original fine labels remain attributable to HPA; this mapping is an explicit modelling choice.
| Predictor | Local mode | Role | Important limitation |
|---|---|---|---|
| Light Attention | ProtT5 + published checkpoint | Protein-language-model evidence | Single-label probability head; public training overlap may exist |
| WoLF PSORT | Animal kNN model | Sorting-signal and sequence evidence | Legacy model and label vocabulary |
| ngLOC 1.0 | Official animal n-gram model | Independent generative sequence evidence | Legacy Swiss-Prot training set and normalized legacy score output |
The objective changes the scientific behaviour. Optimizing exact match produces a conservative voter and raises held-out exact accuracy from Light Attention's 37.05% to 37.80%, but misses several rare compartments. Optimizing macro F1 raises held-out macro F1 from 43.57% to 51.34%, while predicting too many label combinations and reducing exact accuracy to 3.16%. The live panel uses the validation-selected exact-match voter; both outcomes remain visible below.
Paste one human protein FASTA sequence. The first uncached run can take several minutes because all models execute locally.
Raw predictor answers, learned coefficients, weighted contributions and thresholds will appear here.
| Experiment still running; this table will be populated from the frozen artifact. |
Each localization is evaluated as an independent binary decision because a protein may have several valid labels.