Define the object
กำหนดสิ่งที่ศึกษา
Specify the biological class, its context and what counts as a defensible negative.
ระบุ biological class, บริบท และความหมายของ negative ให้ชัด
BioData Mining · 2022 · RESEARCH DOSSIER
A general ncRNA detector asks what all ncRNAs share. mSRFR asks what microalgal ncRNAs reveal about their own lineage.
ncRNA detector ทั่วไปถามว่า ncRNA ทั้งหมดเหมือนกันตรงไหน แต่ mSRFR ถามว่า microalgal ncRNA เปิดเผยลายเซ็นของ lineage ตัวเองอย่างไร

THE RESEARCH QUESTION
Microalgal ncRNA data span green algae, diatoms, golden algae and cyanobacteria with strongly unequal class and species counts. A model can easily learn abundance instead of biology.
ข้อมูล microalgal ncRNA ครอบคลุม green algae, diatom, golden algae และ cyanobacteria โดยจำนวนแต่ละ class และ species ไม่สมดุลอย่างมาก โมเดลจึงอาจเรียนรู้ความถี่แทนชีววิทยา
METHOD AS AN ARGUMENT
SMOTE addressed imbalance; Relief selected 20 variables from 106 sequence, secondary-structure, base-pair and triplet features; six classifier families were compared by ten-fold cross-validation; Random Forest was then evaluated against established coding-potential and ncRNA tools.
ใช้ SMOTE แก้ imbalance, Relief เลือก 20 ตัวแปรจาก 106 sequence, secondary-structure, base-pair และ triplet feature เปรียบเทียบ classifier หกตระกูลด้วย ten-fold cross-validation แล้วนำ Random Forest ไปเทียบกับเครื่องมือ coding-potential และ ncRNA เดิม
Specify the biological class, its context and what counts as a defensible negative.
ระบุ biological class, บริบท และความหมายของ negative ให้ชัด
Represent sequence, structure, context or model outputs without losing the scientific question.
แทน sequence, structure, context หรือ model output โดยไม่ทำคำถามวิจัยหาย
Choose a learner whose assumptions fit the representation, imbalance and intended decision.
เลือก learner ที่เหมาะกับ representation, imbalance และการตัดสินใจปลายทาง
Read validation, independent testing and errors as evidence with a defined scope.
อ่าน validation, independent test และ error เป็นหลักฐานที่มีขอบเขต
EVIDENCE, THEN INTERPRETATION
Random Forest reached ROC area 0.992 in model selection. On a 7,292-sequence test set, mSRFR achieved about 97% accuracy with about 2% false positives. Relief analysis identified %GA dinucleotide as a notable microalgal signature.
Random Forest ได้ ROC area 0.992 ในขั้นเลือกโมเดล บน test set 7,292 sequence, mSRFR ได้ accuracy ราว 97% และ false positive ราว 2% การวิเคราะห์ด้วย Relief ชี้ว่า %GA dinucleotide เป็น signature สำคัญของ microalgae
SMOTE creates synthetic points in feature space, not new biological observations. Species-wise independence, raw-read support and validation on newly curated microalgal datasets are necessary before operational deployment.
SMOTE สร้างจุดสังเคราะห์ใน feature space ไม่ใช่ biological observation ใหม่ การแยกทดสอบตาม species การรองรับ raw read และ validation บนฐานข้อมูล microalgae รุ่นใหม่จำเป็นก่อนใช้งานจริง
WHY THIS WORK STILL MATTERS
The study moves from generic prediction toward lineage-aware bioinformatics and shows how feature importance can produce a biological hypothesis, not only a score.
งานขยับจาก generic prediction ไปสู่ lineage-aware bioinformatics และแสดงว่า feature importance สามารถนำไปสู่ biological hypothesis ไม่ใช่เพียงคะแนน
Rebuild the data boundary, preprocessing and evaluation protocol before changing the model.
สร้าง data boundary, preprocessing และ evaluation protocol เดิมให้ได้ก่อนเปลี่ยนโมเดล
Ask whether errors come from biology, annotation, sampling, leakage or a feature that encodes the wrong shortcut.
ถามว่า error มาจาก biology, annotation, sampling, leakage หรือ feature ที่เรียนรู้ shortcut ผิด
Define which candidates should be inspected, synthesized, assayed or held back at each threshold.
กำหนดว่าแต่ละ threshold จะส่ง candidate ใดไป inspect, synthesize, assay หรือพักไว้
TEN LAYERS BEHIND THE RESULT
A sequence is not merely a string. Composition, order, folding, processing and cellular context carry different evidence. A useful model preserves enough of that object for its output to remain scientifically interpretable.
sequence ไม่ใช่เพียงตัวอักษรเรียงกัน composition, order, folding, processing และ cellular context ให้หลักฐานคนละแบบ โมเดลที่มีประโยชน์ต้องรักษาความหมายของสิ่งที่ศึกษาไว้มากพอให้ผลลัพธ์ยังตีความทางวิทยาศาสตร์ได้
Known positives reflect what experiments discovered and databases chose to curate. They are observations under historical sampling, not a complete map of nature. Redundancy, family size and annotation confidence therefore shape what the learner sees.
positive ที่รู้จักสะท้อนสิ่งที่การทดลองค้นพบและฐานข้อมูลเลือก curate มันคือ observation ภายใต้ sampling ในอดีต ไม่ใช่แผนที่ธรรมชาติทั้งหมด redundancy, family size และ annotation confidence จึงกำหนดสิ่งที่ learner มองเห็น
A negative may mean experimentally inactive, unrelated, pseudo, shuffled, coding, unannotated or simply absent from a database. Those meanings are not interchangeable. Negative construction silently defines the scientific question the classifier answers.
negative อาจหมายถึงไม่ active จากการทดลอง ไม่เกี่ยวข้อง pseudo, shuffled, coding, ยังไม่ annotate หรือเพียงไม่อยู่ในฐานข้อมูล ความหมายเหล่านี้แทนกันไม่ได้ วิธีสร้าง negative เป็นตัวกำหนดคำถามวิทยาศาสตร์ของ classifier อย่างเงียบ ๆ
Handcrafted features encode accumulated domain knowledge and make biological assumptions inspectable. Their risk is blindness to signals no one thought to calculate. Learned representations widen the view but require stronger controls against shortcuts and leakage.
handcrafted feature เก็บ domain knowledge และทำให้ตรวจสมมติฐานทางชีววิทยาได้ ความเสี่ยงคือมองไม่เห็นสัญญาณที่ไม่มีใครคิดคำนวณ ส่วน learned representation มองได้กว้างขึ้นแต่ต้องควบคุม shortcut และ leakage เข้มขึ้น
Algorithms are not meaningfully diverse because their names differ. Useful diversity appears when learners respond differently to local neighbourhoods, margins, nonlinear interactions, imbalance or noise. Every additional model should earn its place through complementary error.
algorithm ไม่ได้ diverse เพียงเพราะชื่อไม่เหมือนกัน useful diversity เกิดเมื่อ learner ตอบสนองต่างกันต่อ neighbourhood, margin, nonlinear interaction, imbalance หรือ noise โมเดลที่เพิ่มเข้ามาต้องพิสูจน์คุณค่าด้วย error ที่เสริมกัน
Cross-validation estimates performance in a particular resampling world. Independent, temporal, species-separated or cluster-aware tests ask whether the explanation survives elsewhere. The split is part of the scientific method, not a clerical step.
cross-validation ประเมิน performance ในโลกของ resampling แบบหนึ่ง independent, temporal, species-separated หรือ cluster-aware test ถามว่าคำอธิบายยังอยู่เมื่อบริบทเปลี่ยนหรือไม่ การ split เป็นส่วนหนึ่งของ scientific method ไม่ใช่งานธุรการ
Accuracy summarizes; sensitivity and specificity expose a trade; MCC remains informative under imbalance; PR-AUC reflects positive retrieval; calibration asks whether confidence is usable. None explains the biological cost of an error without a downstream decision.
accuracy สรุปภาพ sensitivity และ specificity เปิด trade-off, MCC มีประโยชน์เมื่อข้อมูลไม่สมดุล PR-AUC สะท้อนการค้น positive และ calibration ถามว่าความมั่นใจใช้ได้หรือไม่ แต่ไม่มี metric ใดอธิบาย biological cost โดยไม่รู้ decision ปลายทาง
False positives can become costly laboratory detours, yet some may be discoveries missing from current annotations. False negatives may hide unusual biology unlike the training canon. Error analysis is where the next hypothesis often begins.
false positive อาจกลายเป็นทางอ้อมราคาแพงในห้องทดลอง แต่บางตัวอาจเป็น discovery ที่ annotation ยังไม่รู้จัก ส่วน false negative อาจซ่อน biology ที่ไม่เหมือน training canon error analysis จึงมักเป็นจุดเริ่ม hypothesis ถัดไป
A paper records the argument, not every operational detail. Reproducibility needs accession lists, preprocessing rules, feature definitions, seeds, software versions, split files and raw model outputs. Without them, the headline number is easier to quote than to test.
paper เก็บ argument แต่ไม่เก็บ operational detail ทั้งหมด reproducibility ต้องมี accession list, preprocessing rule, feature definition, seed, software version, split file และ raw model output มิฉะนั้น headline number จะถูกอ้างได้ง่ายกว่าถูกทดสอบ
The endpoint of prediction is not a label. It is a decision about what to inspect, synthesize, assay, annotate or postpone. Thresholds should therefore be chosen with capacity, cost, uncertainty and the value of discovery in view.
ปลายทางของ prediction ไม่ใช่ label แต่เป็นการตัดสินใจว่าจะ inspect, synthesize, assay, annotate หรือพักอะไรไว้ threshold จึงต้องมอง capacity, cost, uncertainty และคุณค่าของ discovery ร่วมกัน
FROM READING TO A DEFENSIBLE EXTENSION
State the biological object, intended decision and exact comparison before touching code.
ระบุ biological object, decision ที่ต้องการ และสิ่งที่เปรียบเทียบให้ชัดก่อนแตะ code
Record database releases, access dates, inclusion rules and every transformation.
เก็บ database release, วันที่เข้าถึง inclusion rule และ transformation ทุกขั้น
Cluster related sequences before splitting so close families cannot leak across evaluation boundaries.
cluster sequence ที่เกี่ยวข้องก่อน split เพื่อไม่ให้ family ใกล้กันรั่วข้าม evaluation boundary
Measure imbalance by class, species, family and source; do not let one aggregate ratio hide the problem.
วัด imbalance ตาม class, species, family และแหล่งข้อมูล อย่าให้อัตราส่วนรวมซ่อนปัญหา
Start with a transparent representation and learner whose failure can be understood.
เริ่มจาก representation และ learner ที่โปร่งใสและเข้าใจ failure ได้
Compare representations under matched learners and learners under matched representations.
เทียบ representation ภายใต้ learner เดียวกัน และเทียบ learner ภายใต้ representation เดียวกัน
Remove one feature family or pipeline stage at a time to learn where improvement originates.
ถอด feature family หรือ pipeline stage ทีละส่วนเพื่อหาที่มาของ improvement
Keep feature selection, scaling and hyperparameter search inside training data to prevent optimistic leakage.
ทำ feature selection, scaling และ hyperparameter search ภายใน training data เพื่อกัน optimistic leakage
Reserve evidence that differs by time, source, species or family rather than another random slice.
กันหลักฐานที่ต่างตามเวลา แหล่ง species หรือ family แทน random slice อีกชุด
Inspect which families, lengths, structures and confidence ranges create each error type.
ตรวจว่า family, length, structure และช่วง confidence ใดสร้าง error แต่ละชนิด
Use identical sequences, splits and metrics before declaring one method stronger than another.
ใช้ sequence, split และ metric เดียวกันก่อนประกาศว่าวิธีหนึ่งดีกว่า
A score used for triage should correspond to observed risk, not only ranking position.
คะแนนเพื่อ triage ควรสัมพันธ์กับ observed risk ไม่ใช่เพียงอันดับ
Report feature extraction, embedding, training and inference separately, including hardware.
รายงานต้นทุน feature extraction, embedding, training และ inference แยกกันพร้อม hardware
Freeze environments and retain intermediate artifacts so the workflow can be reconstructed.
freeze environment และเก็บ intermediate artifact เพื่อสร้าง workflow ซ้ำได้
Specify the table, sequence metadata, rationale and uncertainty an experimental collaborator actually needs.
กำหนดตาราง sequence metadata, rationale และ uncertainty ที่ผู้ร่วมทดลองต้องใช้จริง
Decide in advance when evidence is too weak, too shifted or too costly to continue.
กำหนดล่วงหน้าว่าเมื่อใดหลักฐานอ่อน shift มาก หรือต้นทุนสูงเกินไป
Turn every limitation into a measurable follow-up rather than a ceremonial final paragraph.
เปลี่ยนข้อจำกัดทุกข้อเป็น follow-up ที่วัดได้ ไม่ใช่ย่อหน้าปิดตามพิธี
Choose one bounded extension with a reproducible baseline and a clear success criterion.
เลือก extension ที่มีขอบเขต baseline ทำซ้ำได้ และ success criterion ชัด
Show confidence, applicability boundaries and unresolved cases instead of one definitive badge.
แสดง confidence, applicability boundary และกรณียังไม่ resolved แทน badge ฟันธง
Judge success by a better research or laboratory decision, not a decimal point alone.
ตัดสินความสำเร็จจาก decision วิจัยหรือห้องทดลองที่ดีขึ้น ไม่ใช่ทศนิยมอย่างเดียว
PAPER-SPECIFIC READING
Random Forest reached ROC area 0.992 in model selection. On a 7,292-sequence test set, mSRFR achieved about 97% accuracy with about 2% false positives. Relief analysis identified %GA dinucleotide as a notable microalgal signature. SMOTE creates synthetic points in feature space, not new biological observations. Species-wise independence, raw-read support and validation on newly curated microalgal datasets are necessary before operational deployment.
Random Forest ได้ ROC area 0.992 ในขั้นเลือกโมเดล บน test set 7,292 sequence, mSRFR ได้ accuracy ราว 97% และ false positive ราว 2% การวิเคราะห์ด้วย Relief ชี้ว่า %GA dinucleotide เป็น signature สำคัญของ microalgae SMOTE สร้างจุดสังเคราะห์ใน feature space ไม่ใช่ biological observation ใหม่ การแยกทดสอบตาม species การรองรับ raw read และ validation บนฐานข้อมูล microalgae รุ่นใหม่จำเป็นก่อนใช้งานจริง
Begin with the exact dataset boundary and evaluation protocol. A newer algorithm on a different split does not reproduce the original claim.
เริ่มจาก dataset boundary และ evaluation protocol เดิม algorithm ใหม่บน split คนละแบบไม่ถือว่าทำซ้ำข้ออ้างเดิม
Test source shift, family separation, harder negatives, missing features and confidence calibration before expanding the claim.
ทดสอบ source shift, family separation, negative ที่ยากขึ้น feature ที่หาย และ confidence calibration ก่อนขยายข้ออ้าง
Connect ranked candidates to an explicit inspection or experimental queue, then use outcomes to revise the representation and sampling strategy.
เชื่อม ranked candidate กับ inspection หรือ experimental queue ที่ชัด แล้วใช้ผลย้อนกลับมาปรับ representation และ sampling strategy
DATASET AND EVIDENCE AUDIT
Determine whether the unit is a mature sequence, precursor, full protein, peptide candidate, genomic window, fraction or aggregated prediction. Mixing levels can create impressive metrics that answer no coherent biological question.
ต้องรู้ว่า unit คือ mature sequence, precursor, full protein, peptide candidate, genomic window, fraction หรือ aggregated prediction การผสมคนละระดับอาจสร้าง metric สวยแต่ไม่ตอบคำถามชีววิทยาที่สอดคล้องกัน
Separate experimentally confirmed records from computational annotation and database inheritance. A label copied through several resources is not several independent pieces of evidence.
แยก record ที่ยืนยันจากการทดลองออกจาก computational annotation และ label ที่สืบทอดผ่านฐานข้อมูล การคัดลอก label ผ่านหลาย resource ไม่ได้กลายเป็นหลักฐานอิสระหลายชิ้น
Sequence identity, shared families and near-duplicate structures can make a random split test memory rather than generalization. The similarity threshold must follow the intended deployment claim.
sequence identity, shared family และโครงสร้างเกือบซ้ำทำให้ random split ทดสอบความจำแทน generalization ค่า similarity threshold ต้องสัมพันธ์กับข้ออ้างการใช้งานจริง
Easy negatives reward superficial shortcuts. A deployment-oriented set should include candidates that pass early filters, share length or localization context, or resemble positives while lacking the target evidence.
negative ที่ง่ายให้รางวัล shortcut แบบผิวเผิน ชุดที่ใกล้การใช้งานควรรวม candidate ที่ผ่าน early filter มีความยาวหรือ localization context คล้ายกัน หรือดูเหมือน positive แต่ขาด target evidence
Databases grow, taxonomic coverage expands, instruments change and annotations are corrected. Re-evaluation on a later release is a scientific experiment about temporal robustness, not routine maintenance.
ฐานข้อมูลโตขึ้น taxonomic coverage กว้างขึ้น เครื่องมือเปลี่ยนและ annotation ถูกแก้ การประเมินบน release ใหม่เป็นการทดลองเรื่อง temporal robustness ไม่ใช่เพียง maintenance
A false positive may consume synthesis and assay capacity; a false negative may hide an unusual family. The preferred operating point depends on budget, discovery value and whether a second-stage filter exists.
false positive อาจกิน capacity การสังเคราะห์และ assay ส่วน false negative อาจซ่อน family แปลกใหม่ operating point ที่เหมาะขึ้นกับงบ คุณค่าการค้นพบ และการมี second-stage filter
FOUR RESEARCH PROJECTS INSIDE THE NEXT QUESTION
Recover accession lists, recreate features, preserve the original split logic and explain every deviation. The deliverable is a transparent baseline plus a discrepancy report, not merely code that runs.
กู้ accession list สร้าง feature ใหม่ รักษา split logic เดิมและอธิบายทุก deviation ผลงานคือ baseline โปร่งใสพร้อม discrepancy report ไม่ใช่แค่ code ที่รันได้
Introduce family-aware separation, harder negatives, later database releases and repeated external tests. Measure not only the drop, but which biological groups produce it.
เพิ่ม family-aware separation, negative ที่ยากขึ้น database release ใหม่ และ external test หลายชุด วัดไม่เพียงคะแนนที่ลด แต่ดู biological group ที่ทำให้ลดด้วย
Use ablation, grouped permutation, counterfactual sequence edits and structural review to distinguish a biologically plausible signal from an accidental dataset shortcut.
ใช้ ablation, grouped permutation, counterfactual sequence edit และ structural review เพื่อแยกสัญญาณสมเหตุผลทางชีววิทยาออกจาก dataset shortcut โดยบังเอิญ
Combine score, uncertainty, novelty, diversity and assay cost into a shortlist. Record why each candidate was selected so wet-lab feedback can improve the next computational cycle.
รวม score, uncertainty, novelty, diversity และ assay cost เป็น shortlist พร้อมเก็บเหตุผลการเลือกทุก candidate เพื่อให้ผล wet lab ปรับปรุง computational cycle ถัดไปได้
JOIN THE RESEARCH
The study moves from generic prediction toward lineage-aware bioinformatics and shows how feature importance can produce a biological hypothesis, not only a score. A student does not need to arrive knowing every biological database or model. The useful starting point is intellectual patience: trace one dataset carefully, question one negative definition, reproduce one result and make one improvement whose source can be explained.
งานขยับจาก generic prediction ไปสู่ lineage-aware bioinformatics และแสดงว่า feature importance สามารถนำไปสู่ biological hypothesis ไม่ใช่เพียงคะแนน นักศึกษาไม่จำเป็นต้องเข้ามาพร้อมความรู้ทุกฐานข้อมูลหรือทุกโมเดล จุดเริ่มที่มีค่าคือความอดทนทางปัญญา ตาม dataset หนึ่งชุดให้ละเอียด ตั้งคำถามกับ negative definition หนึ่งแบบ ทำผลหนึ่งงานให้ซ้ำได้ และสร้าง improvement ที่อธิบายที่มาได้
Feature analysis, sequence families and error clusters offer a disciplined route from curiosity to a testable biological hypothesis.
feature analysis, sequence family และ error cluster เปลี่ยนความสงสัยให้เป็น biological hypothesis ที่ทดสอบได้อย่างมีวินัย
The strongest contribution may be a reliable pipeline, a better split, an interpretable ranking or a handoff that makes laboratory work more selective.
contribution ที่แข็งแรงอาจเป็น pipeline ที่เชื่อถือได้ split ที่ดีขึ้น ranking ที่ตีความได้ หรือ handoff ที่ช่วยให้ห้องทดลองเลือกงานได้แม่นขึ้น
Reproducible code, documented assumptions and honest negative results are not secondary work. They are infrastructure for the next discovery.
code ที่ทำซ้ำได้ สมมติฐานที่บันทึกไว้ และ negative result ที่ซื่อตรงไม่ใช่งานรอง แต่เป็น infrastructure ของ discovery ถัดไป
PUBLICATION RECORD
Anuntakarun, S., Lertampaiporn, S., Laomettachit, T., Wattanapornprom, W., & Ruengjitchatchawalya, M. (2022). mSRFR: a machine learning model using microalgal signature features for ncRNA classification. BioData Mining, 15, 8. https://doi.org/10.1186/s13040-022-00291-0IEEESongtham Anuntakarun, Supatcha Lertampaiporn, Teeraphan Laomettachit, Warin Wattanapornprom, Marasri Ruengjitchatchawalya, “mSRFR: a machine learning model using microalgal signature features for ncRNA classification,” BioData Mining, 2022, doi: 10.1186/s13040-022-00291-0.BibTeX@article{anuntakarun2022msrfr,
title={mSRFR: a machine learning model using microalgal signature features for ncRNA classification},
author={Songtham Anuntakarun and Supatcha Lertampaiporn and Teeraphan Laomettachit and Warin Wattanapornprom and Marasri Ruengjitchatchawalya},
year={2022},
doi={10.1186/s13040-022-00291-0}
}