Define the object
กำหนดสิ่งที่ศึกษา
Specify the biological class, its context and what counts as a defensible negative.
ระบุ biological class, บริบท และความหมายของ negative ให้ชัด
Journal of Molecular Graphics and Modelling · 2026 · RESEARCH DOSSIER
Finding an AMP-like sequence is easy compared with deciding which candidate deserves synthesis and assay.
การหา sequence ที่ดูคล้าย AMP ง่ายกว่าการตัดสินใจมากว่า candidate ใดควรถูกนำไปสังเคราะห์และทดสอบจริง

THE RESEARCH QUESTION
Machine-learning AMP screens can produce biologically implausible false positives, particularly from hydrophobic or compositionally biased peptides. Peptidomics mixtures add a second difficulty: activity may reflect enrichment patterns rather than one dominant sequence.
machine-learning AMP screen อาจให้ false positive ที่ไม่สมเหตุผลทางชีววิทยา โดยเฉพาะ peptide ที่ hydrophobic หรือ composition เอนเอียง ส่วน peptidomics mixture เพิ่มความยากอีกชั้น เพราะ activity อาจมาจากรูปแบบการสะสมของหลาย peptide ไม่ใช่ sequence เด่นตัวเดียว
METHOD AS AN ARGUMENT
The framework separates conservative AMP screening from mechanism-guided prioritization. MAP grades membrane-interaction plausibility from interpretable physicochemical constraints; AIPx ranks candidates using weights calibrated from experimentally characterized anti-Vibrio peptides; structural visualization supports review.
framework แยก conservative AMP screening ออกจาก mechanism-guided prioritization โดย MAP ให้ระดับความเป็นไปได้ของ membrane interaction จากข้อจำกัดทาง physicochemical ที่ตีความได้ ส่วน AIPx จัดอันดับ candidate ด้วยน้ำหนักที่ calibrate จาก anti-Vibrio peptide ที่มีผลทดลอง พร้อม structural visualization เพื่อช่วยตรวจสอบ
Specify the biological class, its context and what counts as a defensible negative.
ระบุ biological class, บริบท และความหมายของ negative ให้ชัด
Represent sequence, structure, context or model outputs without losing the scientific question.
แทน sequence, structure, context หรือ model output โดยไม่ทำคำถามวิจัยหาย
Choose a learner whose assumptions fit the representation, imbalance and intended decision.
เลือก learner ที่เหมาะกับ representation, imbalance และการตัดสินใจปลายทาง
Read validation, independent testing and errors as evidence with a defined scope.
อ่าน validation, independent test และ error เป็นหลักฐานที่มีขอบเขต
EVIDENCE, THEN INTERPRETATION
Across peptidomics-derived fractions, experimental anti-Vibrio activity aligned more consistently with enrichment of high-ranking peptides than with total predicted AMP abundance. The result supports distribution-aware interpretation and sequential filtering.
ใน fraction จาก peptidomics ผล anti-Vibrio จากการทดลองสอดคล้องกับการ enrichment ของ peptide อันดับสูงมากกว่าจำนวน predicted AMP รวม สนับสนุนการอ่านผลแบบ distribution-aware และการกรองเป็นลำดับขั้น
AIPx uses MIC as a coarse ranking reference, not a direct MIC predictor. Calibration is target-oriented, mixture effects are not full causal attribution, and broader species transfer requires new experimental anchors.
AIPx ใช้ MIC เป็น coarse ranking reference ไม่ใช่ direct MIC predictor การ calibrate ผูกกับ target, mixture effect ยังไม่ใช่ causal attribution เต็มรูปแบบ และการย้ายไป species อื่นต้องมี experimental anchor ใหม่
WHY THIS WORK STILL MATTERS
This work closes the loop from sequence prediction toward experimentally actionable prioritization without hiding the mechanism behind a single opaque score.
งานนี้เชื่อมวงจรจาก sequence prediction ไปสู่การจัดลำดับที่ใช้วางแผนทดลองได้ โดยไม่ซ่อนกลไกทั้งหมดไว้หลังคะแนนทึบเพียงค่าเดียว
Rebuild the data boundary, preprocessing and evaluation protocol before changing the model.
สร้าง data boundary, preprocessing และ evaluation protocol เดิมให้ได้ก่อนเปลี่ยนโมเดล
Ask whether errors come from biology, annotation, sampling, leakage or a feature that encodes the wrong shortcut.
ถามว่า error มาจาก biology, annotation, sampling, leakage หรือ feature ที่เรียนรู้ shortcut ผิด
Define which candidates should be inspected, synthesized, assayed or held back at each threshold.
กำหนดว่าแต่ละ threshold จะส่ง candidate ใดไป inspect, synthesize, assay หรือพักไว้
TEN LAYERS BEHIND THE RESULT
A sequence is not merely a string. Composition, order, folding, processing and cellular context carry different evidence. A useful model preserves enough of that object for its output to remain scientifically interpretable.
sequence ไม่ใช่เพียงตัวอักษรเรียงกัน composition, order, folding, processing และ cellular context ให้หลักฐานคนละแบบ โมเดลที่มีประโยชน์ต้องรักษาความหมายของสิ่งที่ศึกษาไว้มากพอให้ผลลัพธ์ยังตีความทางวิทยาศาสตร์ได้
Known positives reflect what experiments discovered and databases chose to curate. They are observations under historical sampling, not a complete map of nature. Redundancy, family size and annotation confidence therefore shape what the learner sees.
positive ที่รู้จักสะท้อนสิ่งที่การทดลองค้นพบและฐานข้อมูลเลือก curate มันคือ observation ภายใต้ sampling ในอดีต ไม่ใช่แผนที่ธรรมชาติทั้งหมด redundancy, family size และ annotation confidence จึงกำหนดสิ่งที่ learner มองเห็น
A negative may mean experimentally inactive, unrelated, pseudo, shuffled, coding, unannotated or simply absent from a database. Those meanings are not interchangeable. Negative construction silently defines the scientific question the classifier answers.
negative อาจหมายถึงไม่ active จากการทดลอง ไม่เกี่ยวข้อง pseudo, shuffled, coding, ยังไม่ annotate หรือเพียงไม่อยู่ในฐานข้อมูล ความหมายเหล่านี้แทนกันไม่ได้ วิธีสร้าง negative เป็นตัวกำหนดคำถามวิทยาศาสตร์ของ classifier อย่างเงียบ ๆ
Handcrafted features encode accumulated domain knowledge and make biological assumptions inspectable. Their risk is blindness to signals no one thought to calculate. Learned representations widen the view but require stronger controls against shortcuts and leakage.
handcrafted feature เก็บ domain knowledge และทำให้ตรวจสมมติฐานทางชีววิทยาได้ ความเสี่ยงคือมองไม่เห็นสัญญาณที่ไม่มีใครคิดคำนวณ ส่วน learned representation มองได้กว้างขึ้นแต่ต้องควบคุม shortcut และ leakage เข้มขึ้น
Algorithms are not meaningfully diverse because their names differ. Useful diversity appears when learners respond differently to local neighbourhoods, margins, nonlinear interactions, imbalance or noise. Every additional model should earn its place through complementary error.
algorithm ไม่ได้ diverse เพียงเพราะชื่อไม่เหมือนกัน useful diversity เกิดเมื่อ learner ตอบสนองต่างกันต่อ neighbourhood, margin, nonlinear interaction, imbalance หรือ noise โมเดลที่เพิ่มเข้ามาต้องพิสูจน์คุณค่าด้วย error ที่เสริมกัน
Cross-validation estimates performance in a particular resampling world. Independent, temporal, species-separated or cluster-aware tests ask whether the explanation survives elsewhere. The split is part of the scientific method, not a clerical step.
cross-validation ประเมิน performance ในโลกของ resampling แบบหนึ่ง independent, temporal, species-separated หรือ cluster-aware test ถามว่าคำอธิบายยังอยู่เมื่อบริบทเปลี่ยนหรือไม่ การ split เป็นส่วนหนึ่งของ scientific method ไม่ใช่งานธุรการ
Accuracy summarizes; sensitivity and specificity expose a trade; MCC remains informative under imbalance; PR-AUC reflects positive retrieval; calibration asks whether confidence is usable. None explains the biological cost of an error without a downstream decision.
accuracy สรุปภาพ sensitivity และ specificity เปิด trade-off, MCC มีประโยชน์เมื่อข้อมูลไม่สมดุล PR-AUC สะท้อนการค้น positive และ calibration ถามว่าความมั่นใจใช้ได้หรือไม่ แต่ไม่มี metric ใดอธิบาย biological cost โดยไม่รู้ decision ปลายทาง
False positives can become costly laboratory detours, yet some may be discoveries missing from current annotations. False negatives may hide unusual biology unlike the training canon. Error analysis is where the next hypothesis often begins.
false positive อาจกลายเป็นทางอ้อมราคาแพงในห้องทดลอง แต่บางตัวอาจเป็น discovery ที่ annotation ยังไม่รู้จัก ส่วน false negative อาจซ่อน biology ที่ไม่เหมือน training canon error analysis จึงมักเป็นจุดเริ่ม hypothesis ถัดไป
A paper records the argument, not every operational detail. Reproducibility needs accession lists, preprocessing rules, feature definitions, seeds, software versions, split files and raw model outputs. Without them, the headline number is easier to quote than to test.
paper เก็บ argument แต่ไม่เก็บ operational detail ทั้งหมด reproducibility ต้องมี accession list, preprocessing rule, feature definition, seed, software version, split file และ raw model output มิฉะนั้น headline number จะถูกอ้างได้ง่ายกว่าถูกทดสอบ
The endpoint of prediction is not a label. It is a decision about what to inspect, synthesize, assay, annotate or postpone. Thresholds should therefore be chosen with capacity, cost, uncertainty and the value of discovery in view.
ปลายทางของ prediction ไม่ใช่ label แต่เป็นการตัดสินใจว่าจะ inspect, synthesize, assay, annotate หรือพักอะไรไว้ threshold จึงต้องมอง capacity, cost, uncertainty และคุณค่าของ discovery ร่วมกัน
FROM READING TO A DEFENSIBLE EXTENSION
State the biological object, intended decision and exact comparison before touching code.
ระบุ biological object, decision ที่ต้องการ และสิ่งที่เปรียบเทียบให้ชัดก่อนแตะ code
Record database releases, access dates, inclusion rules and every transformation.
เก็บ database release, วันที่เข้าถึง inclusion rule และ transformation ทุกขั้น
Cluster related sequences before splitting so close families cannot leak across evaluation boundaries.
cluster sequence ที่เกี่ยวข้องก่อน split เพื่อไม่ให้ family ใกล้กันรั่วข้าม evaluation boundary
Measure imbalance by class, species, family and source; do not let one aggregate ratio hide the problem.
วัด imbalance ตาม class, species, family และแหล่งข้อมูล อย่าให้อัตราส่วนรวมซ่อนปัญหา
Start with a transparent representation and learner whose failure can be understood.
เริ่มจาก representation และ learner ที่โปร่งใสและเข้าใจ failure ได้
Compare representations under matched learners and learners under matched representations.
เทียบ representation ภายใต้ learner เดียวกัน และเทียบ learner ภายใต้ representation เดียวกัน
Remove one feature family or pipeline stage at a time to learn where improvement originates.
ถอด feature family หรือ pipeline stage ทีละส่วนเพื่อหาที่มาของ improvement
Keep feature selection, scaling and hyperparameter search inside training data to prevent optimistic leakage.
ทำ feature selection, scaling และ hyperparameter search ภายใน training data เพื่อกัน optimistic leakage
Reserve evidence that differs by time, source, species or family rather than another random slice.
กันหลักฐานที่ต่างตามเวลา แหล่ง species หรือ family แทน random slice อีกชุด
Inspect which families, lengths, structures and confidence ranges create each error type.
ตรวจว่า family, length, structure และช่วง confidence ใดสร้าง error แต่ละชนิด
Use identical sequences, splits and metrics before declaring one method stronger than another.
ใช้ sequence, split และ metric เดียวกันก่อนประกาศว่าวิธีหนึ่งดีกว่า
A score used for triage should correspond to observed risk, not only ranking position.
คะแนนเพื่อ triage ควรสัมพันธ์กับ observed risk ไม่ใช่เพียงอันดับ
Report feature extraction, embedding, training and inference separately, including hardware.
รายงานต้นทุน feature extraction, embedding, training และ inference แยกกันพร้อม hardware
Freeze environments and retain intermediate artifacts so the workflow can be reconstructed.
freeze environment และเก็บ intermediate artifact เพื่อสร้าง workflow ซ้ำได้
Specify the table, sequence metadata, rationale and uncertainty an experimental collaborator actually needs.
กำหนดตาราง sequence metadata, rationale และ uncertainty ที่ผู้ร่วมทดลองต้องใช้จริง
Decide in advance when evidence is too weak, too shifted or too costly to continue.
กำหนดล่วงหน้าว่าเมื่อใดหลักฐานอ่อน shift มาก หรือต้นทุนสูงเกินไป
Turn every limitation into a measurable follow-up rather than a ceremonial final paragraph.
เปลี่ยนข้อจำกัดทุกข้อเป็น follow-up ที่วัดได้ ไม่ใช่ย่อหน้าปิดตามพิธี
Choose one bounded extension with a reproducible baseline and a clear success criterion.
เลือก extension ที่มีขอบเขต baseline ทำซ้ำได้ และ success criterion ชัด
Show confidence, applicability boundaries and unresolved cases instead of one definitive badge.
แสดง confidence, applicability boundary และกรณียังไม่ resolved แทน badge ฟันธง
Judge success by a better research or laboratory decision, not a decimal point alone.
ตัดสินความสำเร็จจาก decision วิจัยหรือห้องทดลองที่ดีขึ้น ไม่ใช่ทศนิยมอย่างเดียว
PAPER-SPECIFIC READING
Across peptidomics-derived fractions, experimental anti-Vibrio activity aligned more consistently with enrichment of high-ranking peptides than with total predicted AMP abundance. The result supports distribution-aware interpretation and sequential filtering. AIPx uses MIC as a coarse ranking reference, not a direct MIC predictor. Calibration is target-oriented, mixture effects are not full causal attribution, and broader species transfer requires new experimental anchors.
ใน fraction จาก peptidomics ผล anti-Vibrio จากการทดลองสอดคล้องกับการ enrichment ของ peptide อันดับสูงมากกว่าจำนวน predicted AMP รวม สนับสนุนการอ่านผลแบบ distribution-aware และการกรองเป็นลำดับขั้น AIPx ใช้ MIC เป็น coarse ranking reference ไม่ใช่ direct MIC predictor การ calibrate ผูกกับ target, mixture effect ยังไม่ใช่ causal attribution เต็มรูปแบบ และการย้ายไป species อื่นต้องมี experimental anchor ใหม่
Begin with the exact dataset boundary and evaluation protocol. A newer algorithm on a different split does not reproduce the original claim.
เริ่มจาก dataset boundary และ evaluation protocol เดิม algorithm ใหม่บน split คนละแบบไม่ถือว่าทำซ้ำข้ออ้างเดิม
Test source shift, family separation, harder negatives, missing features and confidence calibration before expanding the claim.
ทดสอบ source shift, family separation, negative ที่ยากขึ้น feature ที่หาย และ confidence calibration ก่อนขยายข้ออ้าง
Connect ranked candidates to an explicit inspection or experimental queue, then use outcomes to revise the representation and sampling strategy.
เชื่อม ranked candidate กับ inspection หรือ experimental queue ที่ชัด แล้วใช้ผลย้อนกลับมาปรับ representation และ sampling strategy
DATASET AND EVIDENCE AUDIT
Determine whether the unit is a mature sequence, precursor, full protein, peptide candidate, genomic window, fraction or aggregated prediction. Mixing levels can create impressive metrics that answer no coherent biological question.
ต้องรู้ว่า unit คือ mature sequence, precursor, full protein, peptide candidate, genomic window, fraction หรือ aggregated prediction การผสมคนละระดับอาจสร้าง metric สวยแต่ไม่ตอบคำถามชีววิทยาที่สอดคล้องกัน
Separate experimentally confirmed records from computational annotation and database inheritance. A label copied through several resources is not several independent pieces of evidence.
แยก record ที่ยืนยันจากการทดลองออกจาก computational annotation และ label ที่สืบทอดผ่านฐานข้อมูล การคัดลอก label ผ่านหลาย resource ไม่ได้กลายเป็นหลักฐานอิสระหลายชิ้น
Sequence identity, shared families and near-duplicate structures can make a random split test memory rather than generalization. The similarity threshold must follow the intended deployment claim.
sequence identity, shared family และโครงสร้างเกือบซ้ำทำให้ random split ทดสอบความจำแทน generalization ค่า similarity threshold ต้องสัมพันธ์กับข้ออ้างการใช้งานจริง
Easy negatives reward superficial shortcuts. A deployment-oriented set should include candidates that pass early filters, share length or localization context, or resemble positives while lacking the target evidence.
negative ที่ง่ายให้รางวัล shortcut แบบผิวเผิน ชุดที่ใกล้การใช้งานควรรวม candidate ที่ผ่าน early filter มีความยาวหรือ localization context คล้ายกัน หรือดูเหมือน positive แต่ขาด target evidence
Databases grow, taxonomic coverage expands, instruments change and annotations are corrected. Re-evaluation on a later release is a scientific experiment about temporal robustness, not routine maintenance.
ฐานข้อมูลโตขึ้น taxonomic coverage กว้างขึ้น เครื่องมือเปลี่ยนและ annotation ถูกแก้ การประเมินบน release ใหม่เป็นการทดลองเรื่อง temporal robustness ไม่ใช่เพียง maintenance
A false positive may consume synthesis and assay capacity; a false negative may hide an unusual family. The preferred operating point depends on budget, discovery value and whether a second-stage filter exists.
false positive อาจกิน capacity การสังเคราะห์และ assay ส่วน false negative อาจซ่อน family แปลกใหม่ operating point ที่เหมาะขึ้นกับงบ คุณค่าการค้นพบ และการมี second-stage filter
FOUR RESEARCH PROJECTS INSIDE THE NEXT QUESTION
Recover accession lists, recreate features, preserve the original split logic and explain every deviation. The deliverable is a transparent baseline plus a discrepancy report, not merely code that runs.
กู้ accession list สร้าง feature ใหม่ รักษา split logic เดิมและอธิบายทุก deviation ผลงานคือ baseline โปร่งใสพร้อม discrepancy report ไม่ใช่แค่ code ที่รันได้
Introduce family-aware separation, harder negatives, later database releases and repeated external tests. Measure not only the drop, but which biological groups produce it.
เพิ่ม family-aware separation, negative ที่ยากขึ้น database release ใหม่ และ external test หลายชุด วัดไม่เพียงคะแนนที่ลด แต่ดู biological group ที่ทำให้ลดด้วย
Use ablation, grouped permutation, counterfactual sequence edits and structural review to distinguish a biologically plausible signal from an accidental dataset shortcut.
ใช้ ablation, grouped permutation, counterfactual sequence edit และ structural review เพื่อแยกสัญญาณสมเหตุผลทางชีววิทยาออกจาก dataset shortcut โดยบังเอิญ
Combine score, uncertainty, novelty, diversity and assay cost into a shortlist. Record why each candidate was selected so wet-lab feedback can improve the next computational cycle.
รวม score, uncertainty, novelty, diversity และ assay cost เป็น shortlist พร้อมเก็บเหตุผลการเลือกทุก candidate เพื่อให้ผล wet lab ปรับปรุง computational cycle ถัดไปได้
JOIN THE RESEARCH
This work closes the loop from sequence prediction toward experimentally actionable prioritization without hiding the mechanism behind a single opaque score. A student does not need to arrive knowing every biological database or model. The useful starting point is intellectual patience: trace one dataset carefully, question one negative definition, reproduce one result and make one improvement whose source can be explained.
งานนี้เชื่อมวงจรจาก sequence prediction ไปสู่การจัดลำดับที่ใช้วางแผนทดลองได้ โดยไม่ซ่อนกลไกทั้งหมดไว้หลังคะแนนทึบเพียงค่าเดียว นักศึกษาไม่จำเป็นต้องเข้ามาพร้อมความรู้ทุกฐานข้อมูลหรือทุกโมเดล จุดเริ่มที่มีค่าคือความอดทนทางปัญญา ตาม dataset หนึ่งชุดให้ละเอียด ตั้งคำถามกับ negative definition หนึ่งแบบ ทำผลหนึ่งงานให้ซ้ำได้ และสร้าง improvement ที่อธิบายที่มาได้
Feature analysis, sequence families and error clusters offer a disciplined route from curiosity to a testable biological hypothesis.
feature analysis, sequence family และ error cluster เปลี่ยนความสงสัยให้เป็น biological hypothesis ที่ทดสอบได้อย่างมีวินัย
The strongest contribution may be a reliable pipeline, a better split, an interpretable ranking or a handoff that makes laboratory work more selective.
contribution ที่แข็งแรงอาจเป็น pipeline ที่เชื่อถือได้ split ที่ดีขึ้น ranking ที่ตีความได้ หรือ handoff ที่ช่วยให้ห้องทดลองเลือกงานได้แม่นขึ้น
Reproducible code, documented assumptions and honest negative results are not secondary work. They are infrastructure for the next discovery.
code ที่ทำซ้ำได้ สมมติฐานที่บันทึกไว้ และ negative result ที่ซื่อตรงไม่ใช่งานรอง แต่เป็น infrastructure ของ discovery ถัดไป
PUBLICATION RECORD
Lertampaiporn, S., Wattanapornprom, W., & Hongsthong, A. (2026). A mechanism-guided framework for prioritizing membrane-interaction anti-Vibrio peptides from peptidomics data. Journal of Molecular Graphics and Modelling, 148, 109497. https://doi.org/10.1016/j.jmgm.2026.109497IEEESupatcha Lertampaiporn, Warin Wattanapornprom, Apiradee Hongsthong, “A mechanism-guided framework for prioritizing membrane-interaction anti-Vibrio peptides from peptidomics data,” Journal of Molecular Graphics and Modelling, 2026, doi: 10.1016/j.jmgm.2026.109497.BibTeX@article{lertampaiporn2026anti,
title={A mechanism-guided framework for prioritizing membrane-interaction anti-Vibrio peptides from peptidomics data},
author={Supatcha Lertampaiporn and Warin Wattanapornprom and Apiradee Hongsthong},
year={2026},
doi={10.1016/j.jmgm.2026.109497}
}