IEEE ICSEC 2017 · 2017 · RESEARCH DOSSIER

Improved Prediction of Eukaryotic Protein Subcellular Localization Using Particle Swarm Optimization of Multiple Classifiers

When predictors disagree, the decision should learn how much each predictor deserves to be trusted.

เมื่อ predictor ให้คำตอบไม่ตรงกัน การตัดสินใจควรเรียนรู้ว่าน้ำหนักความน่าเชื่อถือของแต่ละตัวควรเป็นเท่าไร

Sirapop Nuannimnoi · Supatcha Lertampaiporn · Chinae ThammarongthamDOI 10.1109/ICSEC.2017.8443775
Representative illustration for Improved Prediction of Eukaryotic Protein Subcellular Localization Using Particle Swarm Optimization of Multiple Classifiers

THE RESEARCH QUESTION

Begin with the biological decision, not the fashionable model.

เริ่มจากการตัดสินใจทางชีววิทยา ไม่ใช่โมเดลที่กำลังเป็นกระแส

Existing protein-localization predictors differ in training data, feature assumptions, location coverage and accuracy. Choosing a single tool discards complementary evidence.

predictor ตำแหน่งโปรตีนแต่ละตัวต่างกันทั้ง training data, สมมติฐานของ feature, ตำแหน่งที่รองรับ และ accuracy การเลือกเพียงเครื่องมือเดียวจึงทิ้งหลักฐานที่เสริมกัน

METHOD AS AN ARGUMENT

What the study actually built

สิ่งที่งานสร้างขึ้นจริง

A modified Particle Swarm Optimization algorithm learned weights over probability outputs from four prediction methods. Weighted scores were combined and the highest-scoring location was checked against known labels across nine eukaryotic compartments.

modified Particle Swarm Optimization เรียนรู้น้ำหนักของ probability output จาก prediction method สี่ตัว จากนั้นรวม weighted score และเลือกตำแหน่งคะแนนสูงสุดมาเทียบกับ label จริงในเก้า eukaryotic compartment

01

Define the object

กำหนดสิ่งที่ศึกษา

Specify the biological class, its context and what counts as a defensible negative.

ระบุ biological class, บริบท และความหมายของ negative ให้ชัด

02

Encode evidence

แปลงหลักฐาน

Represent sequence, structure, context or model outputs without losing the scientific question.

แทน sequence, structure, context หรือ model output โดยไม่ทำคำถามวิจัยหาย

03

Learn the boundary

เรียนรู้เส้นแบ่ง

Choose a learner whose assumptions fit the representation, imbalance and intended decision.

เลือก learner ที่เหมาะกับ representation, imbalance และการตัดสินใจปลายทาง

04

Test the claim

ทดสอบข้ออ้าง

Read validation, independent testing and errors as evidence with a defined scope.

อ่าน validation, independent test และ error เป็นหลักฐานที่มีขอบเขต

EVIDENCE, THEN INTERPRETATION

The result matters because of what it allows us to decide.

ผลลัพธ์มีความหมาย เพราะช่วยให้เราตัดสินใจบางอย่างได้ดีขึ้น

Reported evidence

หลักฐานที่รายงาน

The PSO integration achieved 75.63% overall accuracy and produced per-predictor weights that act as reliability signals. The value lies in both the combined answer and the explicit accounting of predictor trust.

การรวมด้วย PSO ได้ overall accuracy 75.63% และให้น้ำหนักราย predictor ที่ใช้เป็น reliability signal คุณค่าจึงอยู่ทั้งในคำตอบรวมและการแสดงความน่าเชื่อถือของเครื่องมืออย่างชัดเจน

Boundary of the claim

ขอบเขตของข้ออ้าง

The reported system addresses single-location proteins. Multi-localization, unseen compartments, calibration stability and comparisons with newer end-to-end protein models remain open questions.

ระบบที่รายงานรองรับ single-location protein ส่วน multi-localization, compartment ที่ไม่เคยเห็น เสถียรภาพของ calibration และการเทียบกับ end-to-end protein model รุ่นใหม่ยังเป็นคำถามเปิด

WHY THIS WORK STILL MATTERS

A model is temporary. A well-formed research question travels further.

โมเดลมีอายุจำกัด แต่คำถามวิจัยที่ตั้งดีเดินทางได้ไกลกว่า

This work establishes the consensus-learning line that later develops into bacterial and multilabel plant localization research.

งานนี้วางฐานให้เส้นทาง consensus learning ที่ต่อมาพัฒนาไปสู่งานตำแหน่งโปรตีนในแบคทีเรียและงาน multilabel ในพืช

FOR STUDENTS

Reproduce before extending

ทำซ้ำให้ได้ก่อนต่อยอด

Rebuild the data boundary, preprocessing and evaluation protocol before changing the model.

สร้าง data boundary, preprocessing และ evaluation protocol เดิมให้ได้ก่อนเปลี่ยนโมเดล

FOR RESEARCHERS

Stress-test the negatives

ทดสอบความยากของ negative

Ask whether errors come from biology, annotation, sampling, leakage or a feature that encodes the wrong shortcut.

ถามว่า error มาจาก biology, annotation, sampling, leakage หรือ feature ที่เรียนรู้ shortcut ผิด

FOR LABS

Connect scores to action

เชื่อมคะแนนกับการลงมือทำ

Define which candidates should be inspected, synthesized, assayed or held back at each threshold.

กำหนดว่าแต่ละ threshold จะส่ง candidate ใดไป inspect, synthesize, assay หรือพักไว้

TEN LAYERS BEHIND THE RESULT

Ten layers a headline metric cannot explain by itself

สิบชั้นความคิดที่ headline metric อธิบายด้วยตัวเองไม่ได้

01

The biological object

สิ่งที่เป็นวัตถุทางชีววิทยา

A sequence is not merely a string. Composition, order, folding, processing and cellular context carry different evidence. A useful model preserves enough of that object for its output to remain scientifically interpretable.

sequence ไม่ใช่เพียงตัวอักษรเรียงกัน composition, order, folding, processing และ cellular context ให้หลักฐานคนละแบบ โมเดลที่มีประโยชน์ต้องรักษาความหมายของสิ่งที่ศึกษาไว้มากพอให้ผลลัพธ์ยังตีความทางวิทยาศาสตร์ได้

02

Positive evidence

หลักฐานฝั่ง positive

Known positives reflect what experiments discovered and databases chose to curate. They are observations under historical sampling, not a complete map of nature. Redundancy, family size and annotation confidence therefore shape what the learner sees.

positive ที่รู้จักสะท้อนสิ่งที่การทดลองค้นพบและฐานข้อมูลเลือก curate มันคือ observation ภายใต้ sampling ในอดีต ไม่ใช่แผนที่ธรรมชาติทั้งหมด redundancy, family size และ annotation confidence จึงกำหนดสิ่งที่ learner มองเห็น

03

Negative construction

การสร้าง negative

A negative may mean experimentally inactive, unrelated, pseudo, shuffled, coding, unannotated or simply absent from a database. Those meanings are not interchangeable. Negative construction silently defines the scientific question the classifier answers.

negative อาจหมายถึงไม่ active จากการทดลอง ไม่เกี่ยวข้อง pseudo, shuffled, coding, ยังไม่ annotate หรือเพียงไม่อยู่ในฐานข้อมูล ความหมายเหล่านี้แทนกันไม่ได้ วิธีสร้าง negative เป็นตัวกำหนดคำถามวิทยาศาสตร์ของ classifier อย่างเงียบ ๆ

04

Representation and features

representation และ feature

Handcrafted features encode accumulated domain knowledge and make biological assumptions inspectable. Their risk is blindness to signals no one thought to calculate. Learned representations widen the view but require stronger controls against shortcuts and leakage.

handcrafted feature เก็บ domain knowledge และทำให้ตรวจสมมติฐานทางชีววิทยาได้ ความเสี่ยงคือมองไม่เห็นสัญญาณที่ไม่มีใครคิดคำนวณ ส่วน learned representation มองได้กว้างขึ้นแต่ต้องควบคุม shortcut และ leakage เข้มขึ้น

05

Model diversity

ความหลากหลายของโมเดล

Algorithms are not meaningfully diverse because their names differ. Useful diversity appears when learners respond differently to local neighbourhoods, margins, nonlinear interactions, imbalance or noise. Every additional model should earn its place through complementary error.

algorithm ไม่ได้ diverse เพียงเพราะชื่อไม่เหมือนกัน useful diversity เกิดเมื่อ learner ตอบสนองต่างกันต่อ neighbourhood, margin, nonlinear interaction, imbalance หรือ noise โมเดลที่เพิ่มเข้ามาต้องพิสูจน์คุณค่าด้วย error ที่เสริมกัน

06

Validation design

การออกแบบ validation

Cross-validation estimates performance in a particular resampling world. Independent, temporal, species-separated or cluster-aware tests ask whether the explanation survives elsewhere. The split is part of the scientific method, not a clerical step.

cross-validation ประเมิน performance ในโลกของ resampling แบบหนึ่ง independent, temporal, species-separated หรือ cluster-aware test ถามว่าคำอธิบายยังอยู่เมื่อบริบทเปลี่ยนหรือไม่ การ split เป็นส่วนหนึ่งของ scientific method ไม่ใช่งานธุรการ

07

Metric interpretation

การตีความ metric

Accuracy summarizes; sensitivity and specificity expose a trade; MCC remains informative under imbalance; PR-AUC reflects positive retrieval; calibration asks whether confidence is usable. None explains the biological cost of an error without a downstream decision.

accuracy สรุปภาพ sensitivity และ specificity เปิด trade-off, MCC มีประโยชน์เมื่อข้อมูลไม่สมดุล PR-AUC สะท้อนการค้น positive และ calibration ถามว่าความมั่นใจใช้ได้หรือไม่ แต่ไม่มี metric ใดอธิบาย biological cost โดยไม่รู้ decision ปลายทาง

08

Error analysis

การวิเคราะห์ error

False positives can become costly laboratory detours, yet some may be discoveries missing from current annotations. False negatives may hide unusual biology unlike the training canon. Error analysis is where the next hypothesis often begins.

false positive อาจกลายเป็นทางอ้อมราคาแพงในห้องทดลอง แต่บางตัวอาจเป็น discovery ที่ annotation ยังไม่รู้จัก ส่วน false negative อาจซ่อน biology ที่ไม่เหมือน training canon error analysis จึงมักเป็นจุดเริ่ม hypothesis ถัดไป

09

Reproducibility

การทำซ้ำได้

A paper records the argument, not every operational detail. Reproducibility needs accession lists, preprocessing rules, feature definitions, seeds, software versions, split files and raw model outputs. Without them, the headline number is easier to quote than to test.

paper เก็บ argument แต่ไม่เก็บ operational detail ทั้งหมด reproducibility ต้องมี accession list, preprocessing rule, feature definition, seed, software version, split file และ raw model output มิฉะนั้น headline number จะถูกอ้างได้ง่ายกว่าถูกทดสอบ

10

Translation to experiments

การส่งต่อสู่การทดลอง

The endpoint of prediction is not a label. It is a decision about what to inspect, synthesize, assay, annotate or postpone. Thresholds should therefore be chosen with capacity, cost, uncertainty and the value of discovery in view.

ปลายทางของ prediction ไม่ใช่ label แต่เป็นการตัดสินใจว่าจะ inspect, synthesize, assay, annotate หรือพักอะไรไว้ threshold จึงต้องมอง capacity, cost, uncertainty และคุณค่าของ discovery ร่วมกัน

FROM READING TO A DEFENSIBLE EXTENSION

Twenty moves for reproducing, auditing and extending this work

ยี่สิบขั้นสำหรับทำซ้ำ ตรวจสอบ และต่อยอดงานนี้อย่างปกป้องได้

A

Reconstruct the question

เขียนคำถามใหม่

State the biological object, intended decision and exact comparison before touching code.

ระบุ biological object, decision ที่ต้องการ และสิ่งที่เปรียบเทียบให้ชัดก่อนแตะ code

B

Trace provenance

ตามที่มาข้อมูล

Record database releases, access dates, inclusion rules and every transformation.

เก็บ database release, วันที่เข้าถึง inclusion rule และ transformation ทุกขั้น

C

Control redundancy

ควบคุมความซ้ำ

Cluster related sequences before splitting so close families cannot leak across evaluation boundaries.

cluster sequence ที่เกี่ยวข้องก่อน split เพื่อไม่ให้ family ใกล้กันรั่วข้าม evaluation boundary

D

Audit imbalance

ตรวจ class imbalance

Measure imbalance by class, species, family and source; do not let one aggregate ratio hide the problem.

วัด imbalance ตาม class, species, family และแหล่งข้อมูล อย่าให้อัตราส่วนรวมซ่อนปัญหา

E

Rebuild a baseline

สร้าง baseline

Start with a transparent representation and learner whose failure can be understood.

เริ่มจาก representation และ learner ที่โปร่งใสและเข้าใจ failure ได้

F

Separate feature from model

แยก feature จาก model

Compare representations under matched learners and learners under matched representations.

เทียบ representation ภายใต้ learner เดียวกัน และเทียบ learner ภายใต้ representation เดียวกัน

G

Run ablations

ทำ ablation

Remove one feature family or pipeline stage at a time to learn where improvement originates.

ถอด feature family หรือ pipeline stage ทีละส่วนเพื่อหาที่มาของ improvement

H

Tune inside folds

ปรับค่าภายใน fold

Keep feature selection, scaling and hyperparameter search inside training data to prevent optimistic leakage.

ทำ feature selection, scaling และ hyperparameter search ภายใน training data เพื่อกัน optimistic leakage

I

Build independent tests

สร้าง independent test

Reserve evidence that differs by time, source, species or family rather than another random slice.

กันหลักฐานที่ต่างตามเวลา แหล่ง species หรือ family แทน random slice อีกชุด

J

Read confusion biologically

อ่าน confusion matrix เชิงชีววิทยา

Inspect which families, lengths, structures and confidence ranges create each error type.

ตรวจว่า family, length, structure และช่วง confidence ใดสร้าง error แต่ละชนิด

K

Compare matched conditions

เทียบอย่างยุติธรรม

Use identical sequences, splits and metrics before declaring one method stronger than another.

ใช้ sequence, split และ metric เดียวกันก่อนประกาศว่าวิธีหนึ่งดีกว่า

L

Calibrate probabilities

calibrate probability

A score used for triage should correspond to observed risk, not only ranking position.

คะแนนเพื่อ triage ควรสัมพันธ์กับ observed risk ไม่ใช่เพียงอันดับ

M

Track computational cost

เก็บต้นทุนคำนวณ

Report feature extraction, embedding, training and inference separately, including hardware.

รายงานต้นทุน feature extraction, embedding, training และ inference แยกกันพร้อม hardware

N

Preserve software context

เก็บบริบทซอฟต์แวร์

Freeze environments and retain intermediate artifacts so the workflow can be reconstructed.

freeze environment และเก็บ intermediate artifact เพื่อสร้าง workflow ซ้ำได้

O

Design the handoff

ออกแบบการส่งต่อ

Specify the table, sequence metadata, rationale and uncertainty an experimental collaborator actually needs.

กำหนดตาราง sequence metadata, rationale และ uncertainty ที่ผู้ร่วมทดลองต้องใช้จริง

P

Define stopping rules

กำหนดจุดหยุด

Decide in advance when evidence is too weak, too shifted or too costly to continue.

กำหนดล่วงหน้าว่าเมื่อใดหลักฐานอ่อน shift มาก หรือต้นทุนสูงเกินไป

Q

Convert limits to tests

เปลี่ยนข้อจำกัดเป็นการทดลอง

Turn every limitation into a measurable follow-up rather than a ceremonial final paragraph.

เปลี่ยนข้อจำกัดทุกข้อเป็น follow-up ที่วัดได้ ไม่ใช่ย่อหน้าปิดตามพิธี

R

Create a student contribution

ออกแบบงานสำหรับนักศึกษา

Choose one bounded extension with a reproducible baseline and a clear success criterion.

เลือก extension ที่มีขอบเขต baseline ทำซ้ำได้ และ success criterion ชัด

S

Communicate uncertainty

สื่อสาร uncertainty

Show confidence, applicability boundaries and unresolved cases instead of one definitive badge.

แสดง confidence, applicability boundary และกรณียังไม่ resolved แทน badge ฟันธง

T

Ask what changed

ถามว่าโมเดลเปลี่ยนอะไร

Judge success by a better research or laboratory decision, not a decimal point alone.

ตัดสินความสำเร็จจาก decision วิจัยหรือห้องทดลองที่ดีขึ้น ไม่ใช่ทศนิยมอย่างเดียว

PAPER-SPECIFIC READING

Read this result with its own boundary attached.

อ่านผลของงานนี้พร้อมขอบเขตของมันเสมอ

The PSO integration achieved 75.63% overall accuracy and produced per-predictor weights that act as reliability signals. The value lies in both the combined answer and the explicit accounting of predictor trust. The reported system addresses single-location proteins. Multi-localization, unseen compartments, calibration stability and comparisons with newer end-to-end protein models remain open questions.

การรวมด้วย PSO ได้ overall accuracy 75.63% และให้น้ำหนักราย predictor ที่ใช้เป็น reliability signal คุณค่าจึงอยู่ทั้งในคำตอบรวมและการแสดงความน่าเชื่อถือของเครื่องมืออย่างชัดเจน ระบบที่รายงานรองรับ single-location protein ส่วน multi-localization, compartment ที่ไม่เคยเห็น เสถียรภาพของ calibration และการเทียบกับ end-to-end protein model รุ่นใหม่ยังเป็นคำถามเปิด

REPLICATION

Can the reported pipeline be rebuilt?

สร้าง pipeline เดิมกลับมาได้หรือไม่

Begin with the exact dataset boundary and evaluation protocol. A newer algorithm on a different split does not reproduce the original claim.

เริ่มจาก dataset boundary และ evaluation protocol เดิม algorithm ใหม่บน split คนละแบบไม่ถือว่าทำซ้ำข้ออ้างเดิม

ROBUSTNESS

Where does the result weaken?

ผลเริ่มอ่อนลงตรงไหน

Test source shift, family separation, harder negatives, missing features and confidence calibration before expanding the claim.

ทดสอบ source shift, family separation, negative ที่ยากขึ้น feature ที่หาย และ confidence calibration ก่อนขยายข้ออ้าง

TRANSLATION

What should happen after prediction?

หลัง prediction ควรเกิดอะไรขึ้น

Connect ranked candidates to an explicit inspection or experimental queue, then use outcomes to revise the representation and sampling strategy.

เชื่อม ranked candidate กับ inspection หรือ experimental queue ที่ชัด แล้วใช้ผลย้อนกลับมาปรับ representation และ sampling strategy

DATASET AND EVIDENCE AUDIT

Six questions to ask before trusting the next decimal place

หกคำถามที่ควรถามก่อนเชื่อทศนิยมตำแหน่งถัดไป

01 · IDENTITY

What exactly is one sample?

หนึ่ง sample คืออะไรแน่

Determine whether the unit is a mature sequence, precursor, full protein, peptide candidate, genomic window, fraction or aggregated prediction. Mixing levels can create impressive metrics that answer no coherent biological question.

ต้องรู้ว่า unit คือ mature sequence, precursor, full protein, peptide candidate, genomic window, fraction หรือ aggregated prediction การผสมคนละระดับอาจสร้าง metric สวยแต่ไม่ตอบคำถามชีววิทยาที่สอดคล้องกัน

02 · PROVENANCE

Who created the label?

ใครเป็นผู้สร้าง label

Separate experimentally confirmed records from computational annotation and database inheritance. A label copied through several resources is not several independent pieces of evidence.

แยก record ที่ยืนยันจากการทดลองออกจาก computational annotation และ label ที่สืบทอดผ่านฐานข้อมูล การคัดลอก label ผ่านหลาย resource ไม่ได้กลายเป็นหลักฐานอิสระหลายชิ้น

03 · SIMILARITY

Can relatives appear on both sides?

ญาติใกล้กันอยู่คนละฝั่งได้หรือไม่

Sequence identity, shared families and near-duplicate structures can make a random split test memory rather than generalization. The similarity threshold must follow the intended deployment claim.

sequence identity, shared family และโครงสร้างเกือบซ้ำทำให้ random split ทดสอบความจำแทน generalization ค่า similarity threshold ต้องสัมพันธ์กับข้ออ้างการใช้งานจริง

04 · NEGATIVES

Are negatives biologically plausible?

negative สมจริงทางชีววิทยาหรือไม่

Easy negatives reward superficial shortcuts. A deployment-oriented set should include candidates that pass early filters, share length or localization context, or resemble positives while lacking the target evidence.

negative ที่ง่ายให้รางวัล shortcut แบบผิวเผิน ชุดที่ใกล้การใช้งานควรรวม candidate ที่ผ่าน early filter มีความยาวหรือ localization context คล้ายกัน หรือดูเหมือน positive แต่ขาด target evidence

05 · SHIFT

What will change after publication?

อะไรจะเปลี่ยนหลังตีพิมพ์

Databases grow, taxonomic coverage expands, instruments change and annotations are corrected. Re-evaluation on a later release is a scientific experiment about temporal robustness, not routine maintenance.

ฐานข้อมูลโตขึ้น taxonomic coverage กว้างขึ้น เครื่องมือเปลี่ยนและ annotation ถูกแก้ การประเมินบน release ใหม่เป็นการทดลองเรื่อง temporal robustness ไม่ใช่เพียง maintenance

06 · ACTION

Which error is expensive?

error แบบใดมีราคาแพง

A false positive may consume synthesis and assay capacity; a false negative may hide an unusual family. The preferred operating point depends on budget, discovery value and whether a second-stage filter exists.

false positive อาจกิน capacity การสังเคราะห์และ assay ส่วน false negative อาจซ่อน family แปลกใหม่ operating point ที่เหมาะขึ้นกับงบ คุณค่าการค้นพบ และการมี second-stage filter

FOUR RESEARCH PROJECTS INSIDE THE NEXT QUESTION

Extensions that add evidence, not decoration

งานต่อยอดที่เพิ่มหลักฐาน ไม่ใช่เพียงเพิ่มเครื่องมือ

A · REPRODUCTION

Rebuild the original claim

สร้างข้ออ้างเดิมกลับมา

Recover accession lists, recreate features, preserve the original split logic and explain every deviation. The deliverable is a transparent baseline plus a discrepancy report, not merely code that runs.

กู้ accession list สร้าง feature ใหม่ รักษา split logic เดิมและอธิบายทุก deviation ผลงานคือ baseline โปร่งใสพร้อม discrepancy report ไม่ใช่แค่ code ที่รันได้

B · ROBUSTNESS

Replace convenience with harder evidence

เปลี่ยนความสะดวกเป็นหลักฐานที่ยากขึ้น

Introduce family-aware separation, harder negatives, later database releases and repeated external tests. Measure not only the drop, but which biological groups produce it.

เพิ่ม family-aware separation, negative ที่ยากขึ้น database release ใหม่ และ external test หลายชุด วัดไม่เพียงคะแนนที่ลด แต่ดู biological group ที่ทำให้ลดด้วย

C · EXPLANATION

Connect model evidence to mechanism

เชื่อมหลักฐานของโมเดลกับกลไก

Use ablation, grouped permutation, counterfactual sequence edits and structural review to distinguish a biologically plausible signal from an accidental dataset shortcut.

ใช้ ablation, grouped permutation, counterfactual sequence edit และ structural review เพื่อแยกสัญญาณสมเหตุผลทางชีววิทยาออกจาก dataset shortcut โดยบังเอิญ

D · TRANSLATION

Design an experimental queue

ออกแบบ experimental queue

Combine score, uncertainty, novelty, diversity and assay cost into a shortlist. Record why each candidate was selected so wet-lab feedback can improve the next computational cycle.

รวม score, uncertainty, novelty, diversity และ assay cost เป็น shortlist พร้อมเก็บเหตุผลการเลือกทุก candidate เพื่อให้ผล wet lab ปรับปรุง computational cycle ถัดไปได้

JOIN THE RESEARCH

A place for students who enjoy difficult boundaries

พื้นที่สำหรับนักศึกษาที่สนุกกับเส้นแบ่งซึ่งไม่ง่าย

This work establishes the consensus-learning line that later develops into bacterial and multilabel plant localization research. A student does not need to arrive knowing every biological database or model. The useful starting point is intellectual patience: trace one dataset carefully, question one negative definition, reproduce one result and make one improvement whose source can be explained.

งานนี้วางฐานให้เส้นทาง consensus learning ที่ต่อมาพัฒนาไปสู่งานตำแหน่งโปรตีนในแบคทีเรียและงาน multilabel ในพืช นักศึกษาไม่จำเป็นต้องเข้ามาพร้อมความรู้ทุกฐานข้อมูลหรือทุกโมเดล จุดเริ่มที่มีค่าคือความอดทนทางปัญญา ตาม dataset หนึ่งชุดให้ละเอียด ตั้งคำถามกับ negative definition หนึ่งแบบ ทำผลหนึ่งงานให้ซ้ำได้ และสร้าง improvement ที่อธิบายที่มาได้

PATTERN SEEKER

You notice what repeats—and what refuses to repeat.

คุณมองเห็นทั้งสิ่งที่ซ้ำและสิ่งที่ไม่ยอมซ้ำ

Feature analysis, sequence families and error clusters offer a disciplined route from curiosity to a testable biological hypothesis.

feature analysis, sequence family และ error cluster เปลี่ยนความสงสัยให้เป็น biological hypothesis ที่ทดสอบได้อย่างมีวินัย

SYSTEMS THINKER

You connect data, models and experiments.

คุณเชื่อม data, model และ experiment

The strongest contribution may be a reliable pipeline, a better split, an interpretable ranking or a handoff that makes laboratory work more selective.

contribution ที่แข็งแรงอาจเป็น pipeline ที่เชื่อถือได้ split ที่ดีขึ้น ranking ที่ตีความได้ หรือ handoff ที่ช่วยให้ห้องทดลองเลือกงานได้แม่นขึ้น

CAREFUL BUILDER

You prefer defensible progress to theatrical novelty.

คุณเลือกความก้าวหน้าที่ปกป้องได้ มากกว่าความใหม่ที่ดูตื่นเต้น

Reproducible code, documented assumptions and honest negative results are not secondary work. They are infrastructure for the next discovery.

code ที่ทำซ้ำได้ สมมติฐานที่บันทึกไว้ และ negative result ที่ซื่อตรงไม่ใช่งานรอง แต่เป็น infrastructure ของ discovery ถัดไป

PUBLICATION RECORD

Citation metadata

ข้อมูลบรรณานุกรม

APANuannimnoi, S., Lertampaiporn, S., & Thammarongtham, C. (2017). Improved Prediction of Eukaryotic Protein Subcellular Localization Using Particle Swarm Optimization of Multiple Classifiers. 2017 21st International Computer Science and Engineering Conference (ICSEC), 1–5. https://doi.org/10.1109/ICSEC.2017.8443775IEEESirapop Nuannimnoi, Supatcha Lertampaiporn, Chinae Thammarongtham, “Improved Prediction of Eukaryotic Protein Subcellular Localization Using Particle Swarm Optimization of Multiple Classifiers,” IEEE ICSEC 2017, 2017, doi: 10.1109/ICSEC.2017.8443775.BibTeX@article{nuannimnoi2017eukaryotic, title={Improved Prediction of Eukaryotic Protein Subcellular Localization Using Particle Swarm Optimization of Multiple Classifiers}, author={Sirapop Nuannimnoi and Supatcha Lertampaiporn and Chinae Thammarongtham}, year={2017}, doi={10.1109/ICSEC.2017.8443775} }