SUPATCHA / RESEARCH
Dossiersแฟ้มงานวิจัยPublicationsผลงานตีพิมพ์

Nucleic Acids Research · 2014

Identification of Non-Coding RNAs with a New Composite Feature in the Hybrid Random Forest Ensemble Algorithm

Short RNA, long RNA, bacterial genome, human region: one detector must recognize a family resemblance without pretending every ncRNA looks the same.

ทั้ง short RNA, long RNA, bacterial genome และ human genomic region ต้องถูกอ่านด้วยระบบเดียวที่มองเห็นความคล้ายร่วมกัน โดยไม่แกล้งทำว่า ncRNA ทุกชนิดมีหน้าตาเหมือนกัน

Supatcha Lertampaiporn · Chinae Thammarongtham · Chakarida Nukoolkit · Boonserm Kaewkamnerdpong · Marasri RuengjitchatchawalyaDOI 10.1093/nar/gku325

The research problemโจทย์วิจัย

The difficult question under the model

คำถามยากที่อยู่ใต้โมเดล

Non-coding RNAs vary widely in length, structure and sequence composition. Methods tuned for one family may not generalize to long or heterogeneous ncRNAs. The study asks whether a composite feature can compress several biological signals into a form a random-forest ensemble can use.

ทั้ง short RNA, long RNA, bacterial genome และ human genomic region ต้องถูกอ่านด้วยระบบเดียวที่มองเห็นความคล้ายร่วมกัน โดยไม่แกล้งทำว่า ncRNA ทุกชนิดมีหน้าตาเหมือนกัน

This study treats computational prediction as scientific triage: the purpose is not to replace biological validation, but to make the next experiment more informed. Data design, biological representation, model diversity and evaluation therefore belong to one argument.

งานนี้มอง computational prediction เป็น scientific triage จุดประสงค์ไม่ใช่แทนที่การยืนยันทางชีววิทยา แต่ทำให้การทดลองถัดไปมีข้อมูลรองรับมากขึ้น การออกแบบข้อมูล การแทนความรู้ชีววิทยา ความหลากหลายของโมเดล และการประเมินผลจึงเป็น argument เดียวกัน

Hybrid random-forest ncRNA detection workflow
FIGURE 1Hybrid random-forest ncRNA detection workflow · Original figure recovered from ncrna-pred.com / ภาพต้นฉบับจาก ncrna-pred.com
92.11%classification accuracy
90.7%sensitivity
93.5%specificity
0.972reported cross-validation AUC

Why the work mattersเหตุผลที่งานนี้สำคัญ

A prediction changes the cost of the next decision

prediction ที่ดีเปลี่ยนต้นทุนของการตัดสินใจถัดไป

Biological value

The system organizes difficult sequence evidence into candidates researchers can inspect, compare and prioritize.

ระบบจัดระเบียบหลักฐานจาก sequence ที่ซับซ้อนให้เป็น candidate ที่นักวิจัยตรวจสอบ เปรียบเทียบ และจัดลำดับได้

Computational value

The study demonstrates how feature engineering, imbalance handling and model combination solve different parts of one problem.

งานแสดงให้เห็นว่า feature engineering, imbalance handling และ model combination แก้ปัญหาคนละส่วนแต่ต้องทำงานร่วมกัน

Practical value

Better prioritization can reduce wasted experimental effort while keeping laboratory validation central.

การจัดลำดับที่ดีช่วยลดงานทดลองที่สูญเปล่า โดยยังคงให้ laboratory validation เป็นศูนย์กลาง

Training value

Students can see a complete research chain: biological question, dataset, representation, model, benchmark and interpretation.

นักศึกษาเห็น research chain ครบตั้งแต่คำถามชีววิทยา ชุดข้อมูล representation โมเดล benchmark และการตีความ

Method journeyเส้นทางวิธีวิจัย

Five decisions that turn sequences into evidence

ห้าการตัดสินใจที่เปลี่ยน sequence ให้เป็นหลักฐาน

Construct balanced ncRNA and negative training sets

Define what counts as a trustworthy positive and a meaningful negative before the model sees either.

ขั้นตอนที่ 1 ทำให้สมมติฐานส่วนหนึ่งตรวจสอบได้ชัดขึ้น พร้อมเปิดเผยต้นทุน ความไม่แน่นอน และ failure mode ที่ต้องติดตาม

Extract sequence, structure, modularity, robustness and coding-potential features

Convert biological knowledge into measurable signals while acknowledging what each descriptor leaves out.

ขั้นตอนที่ 2 ทำให้สมมติฐานส่วนหนึ่งตรวจสอบได้ชัดขึ้น พร้อมเปิดเผยต้นทุน ความไม่แน่นอน และ failure mode ที่ต้องติดตาม

Combine five significant signals through a logistic-regression SCORE feature

Control complexity so the model learns signal rather than accidental detail.

ขั้นตอนที่ 3 ทำให้สมมติฐานส่วนหนึ่งตรวจสอบได้ชัดขึ้น พร้อมเปิดเผยต้นทุน ความไม่แน่นอน และ failure mode ที่ต้องติดตาม

Train a hybrid random-forest classifier

Combine complementary views because different algorithms expose different boundaries.

ขั้นตอนที่ 4 ทำให้สมมติฐานส่วนหนึ่งตรวจสอบได้ชัดขึ้น พร้อมเปิดเผยต้นทุน ความไม่แน่นอน และ failure mode ที่ต้องติดตาม

Scan prokaryotic and eukaryotic genomic regions and evaluate known ncRNA recovery

Test outside the easiest training setting and interpret errors as information.

ขั้นตอนที่ 5 ทำให้สมมติฐานส่วนหนึ่งตรวจสอบได้ชัดขึ้น พร้อมเปิดเผยต้นทุน ความไม่แน่นอน และ failure mode ที่ต้องติดตาม

Evidence summaryสรุปหลักฐาน

What the paper supports

ข้อสรุปที่ paper รองรับ

01The SCORE composite feature improved characterization of long ncRNA elements.ผลลัพธ์ลำดับที่ 1 เชื่อมการออกแบบของระบบเข้ากับหลักฐานที่วัดได้ และควรอ่านภายในขอบเขตของชุดข้อมูลกับ protocol ที่ระบุในงาน
02The random forest handled heterogeneous ncRNA families through multiple trees and decision rules.ผลลัพธ์ลำดับที่ 2 เชื่อมการออกแบบของระบบเข้ากับหลักฐานที่วัดได้ และควรอ่านภายในขอบเขตของชุดข้อมูลกับ protocol ที่ระบุในงาน
03Genome-wide screening recovered known ncRNAs with sensitivity above 90% in prokaryotic sequences and 77.7% in eukaryotic sequences.ผลลัพธ์ลำดับที่ 3 เชื่อมการออกแบบของระบบเข้ากับหลักฐานที่วัดได้ และควรอ่านภายในขอบเขตของชุดข้อมูลกับ protocol ที่ระบุในงาน

Extended research readingการอ่านงานวิจัยเชิงลึก

Ten layers behind the headline result

สิบชั้นความคิดที่อยู่หลังผลลัพธ์หลัก

1. The biological object

A sequence is not merely a row of characters. Its composition, order, structure, processing and cellular context carry different kinds of evidence. The model becomes scientifically useful only when its representation respects enough of that biological object to make the prediction meaningful.

sequence ไม่ใช่เพียงแถวของตัวอักษร composition, order, structure, processing และ cellular context ให้หลักฐานคนละแบบ โมเดลจะมีคุณค่าทางวิทยาศาสตร์เมื่อ representation เคารพ biological object มากพอให้ prediction มีความหมาย

2. Positive examples

Known positives are shaped by what experiments were able to discover and what databases chose to curate. They are evidence, not a complete map of nature. Sampling and redundancy control therefore determine whether a model learns a general pattern or memorizes a historical collection.

positive example ถูกกำหนดทั้งจากสิ่งที่การทดลองค้นพบได้และสิ่งที่ database เลือก curate จึงเป็นหลักฐานแต่ไม่ใช่แผนที่ธรรมชาติทั้งหมด sampling และ redundancy control ตัดสินว่าโมเดลเรียนรู้ pattern ทั่วไปหรือจำคอลเลกชันเดิม

3. Negative examples

In bioinformatics, “negative” can mean experimentally absent, unrelated, shuffled, pseudo, or simply not yet annotated. Those meanings are not interchangeable. The negative set silently defines the scientific question the classifier is actually answering.

ใน bioinformatics คำว่า negative อาจหมายถึงไม่พบจากการทดลอง ไม่เกี่ยวข้อง shuffled, pseudo หรือเพียงยังไม่ถูก annotate ความหมายเหล่านี้แทนกันไม่ได้ และ negative set เป็นตัวกำหนดคำถามวิทยาศาสตร์จริงของ classifier

4. Feature design

Handcrafted features encode accumulated domain knowledge. Their advantage is interpretability; their risk is blindness to patterns no one thought to calculate. A strong study makes the choice explicit and examines which feature families actually carry the decision.

handcrafted feature บรรจุ domain knowledge ที่สะสมมา ข้อดีคือ interpretability ความเสี่ยงคือมองไม่เห็น pattern ที่ไม่มีใครคิดคำนวณ งานที่ดีต้องบอกเหตุผลของการเลือกและตรวจสอบว่า feature family ใดมีผลต่อ decision

5. Model diversity

Two algorithms are not meaningfully diverse merely because their names differ. Useful diversity appears when models respond differently to local neighborhoods, margins, nonlinear interactions, imbalance or noise. Ensemble design should earn diversity rather than count models.

algorithm สองตัวไม่ได้ diverse เพียงเพราะชื่อไม่เหมือนกัน useful diversity เกิดเมื่อโมเดลตอบสนองต่างกันต่อ neighborhood, margin, nonlinear interaction, imbalance หรือ noise การออกแบบ ensemble ต้องสร้าง diversity ที่มีเหตุผล

6. Validation design

Cross-validation estimates performance under a particular resampling world. Independent tests ask a harder question: does the explanation survive data collected elsewhere? Species separation, temporal separation and homology control can be more informative than another decimal point.

cross-validation ประเมิน performance ภายใต้โลกของ resampling แบบหนึ่ง independent test ถามยากกว่าว่าคำอธิบายยังอยู่เมื่อข้อมูลมาจากที่อื่นหรือไม่ species separation, temporal separation และ homology control อาจมีข้อมูลมากกว่าทศนิยมอีกหนึ่งตำแหน่ง

7. Metric interpretation

Accuracy summarizes; sensitivity and specificity expose the trade; MCC remains informative under imbalance; AUC describes ranking over thresholds. None explains the biological cost of an error. Metric choice should follow the decision the system is meant to support.

accuracy สรุปภาพรวม sensitivity และ specificity เปิด trade-off, MCC มีประโยชน์เมื่อข้อมูลไม่สมดุล และ AUC อธิบาย ranking หลาย threshold แต่ไม่มีตัวใดอธิบาย biological cost ของ error ได้เอง metric ต้องตาม decision ที่ระบบจะสนับสนุน

8. Error analysis

False positives may be expensive laboratory detours, but some may also be unannotated discoveries. False negatives can hide unusual biology that does not resemble the training canon. Error analysis is therefore not housekeeping; it is where the next hypothesis often begins.

false positive อาจเป็นต้นทุนทดลองที่สูญเปล่า แต่บางตัวอาจเป็น discovery ที่ยังไม่ถูก annotate ส่วน false negative อาจซ่อน biology ที่ไม่เหมือน training canon error analysis จึงไม่ใช่งานเก็บกวาด แต่เป็นจุดเริ่มของ hypothesis ถัดไป

9. Reproducibility

A paper records the argument; a migration archive preserves the operational memory around it: figures, readmes, datasets, legacy pages and downloadable artifacts. Reproducibility improves when future students can see both the polished result and the working ecosystem.

paper เก็บ argument ส่วน migration archive เก็บ operational memory รอบงาน ทั้ง figure, readme, dataset, legacy page และ artifact reproducibility ดีขึ้นเมื่อรุ่นถัดไปเห็นทั้งผลลัพธ์ที่เรียบเรียงแล้วและ working ecosystem

10. Translation to experiments

The endpoint of a predictor is not a label. It is a decision about what to inspect, synthesize, assay, annotate or discuss next. The strongest computational biology keeps that downstream action visible from the beginning.

ปลายทางของ predictor ไม่ใช่ label แต่เป็นการตัดสินใจว่าจะ inspect, synthesize, assay, annotate หรืออภิปรายอะไรต่อ computational biology ที่แข็งแรงต้องมอง downstream action นี้ตั้งแต่เริ่ม

Reproducibility practicumpracticum สำหรับการทำซ้ำและต่อยอด

Twenty moves from reading to a defensible extension

ยี่สิบขั้นจากการอ่านไปสู่งานต่อยอดที่ปกป้องได้

This layer is designed for students who want to turn the paper into a working research programme. It makes the invisible decisions explicit: how to rebuild the evidence, challenge it fairly and prepare a result for biological collaboration.

ส่วนนี้ออกแบบสำหรับนักศึกษาที่ต้องการเปลี่ยน paper ให้เป็น working research programme โดยเปิด decision ที่มักมองไม่เห็น ทั้งการสร้าง evidence ใหม่ การท้าทายผลอย่างเป็นธรรม และการเตรียมผลลัพธ์เพื่อร่วมงานกับนักชีววิทยา

A. Reconstruct the question

Write the prediction task as a decision that a biologist must make. Name the unit being classified, the biological scope, the positive definition, the negative definition and the downstream action. If those five items are vague, model comparison will create precise numbers for an imprecise question.

เขียน prediction task ให้เป็น decision ที่นักชีววิทยาต้องใช้จริง ระบุ unit, biological scope, positive definition, negative definition และ downstream action ถ้าห้าส่วนนี้ยังคลุมเครือ การเปรียบเทียบโมเดลจะสร้างตัวเลขที่แม่นยำให้กับคำถามที่ไม่แม่นยำ

B. Trace dataset provenance

Record where every sequence came from, which database release was used, when it was downloaded, how labels were assigned and whether evidence was experimental, inferred or predicted. Provenance is not administration; it determines which claims the dataset can support and which historical biases it carries.

บันทึกที่มาของทุก sequence, database release, วันที่ download, วิธีให้ label และชนิดของ evidence ว่า experimental, inferred หรือ predicted provenance ไม่ใช่งานธุรการ แต่กำหนดว่า dataset รองรับ claim แบบใดและพา historical bias อะไรมาด้วย

C. Control redundancy

Cluster sequences before splitting, inspect identity thresholds and keep close homologues from leaking across training and testing. A random split can reward memorization of biological families while appearing to measure generalization. Report both the threshold and the direction in which performance changes when control becomes stricter.

cluster sequence ก่อน split ตรวจ identity threshold และป้องกัน close homologue รั่วจาก training ไป testing random split อาจให้รางวัลกับการจำ biological family ทั้งที่ดูเหมือนวัด generalization ควรรายงานทั้ง threshold และทิศทางที่ performance เปลี่ยนเมื่อควบคุมเข้มขึ้น

D. Audit class imbalance

Count examples globally and within biologically meaningful subgroups. Compare resampling, class weighting, threshold adjustment and ensemble strategies without letting the test set influence the choice. Imbalance is not only a ratio; it changes which errors the model sees often enough to learn from.

นับตัวอย่างทั้งภาพรวมและ subgroup ที่มีความหมายทางชีววิทยา เปรียบเทียบ resampling, class weighting, threshold adjustment และ ensemble โดยไม่ให้ test set มีอิทธิพลต่อการเลือก imbalance ไม่ใช่เพียง ratio แต่เปลี่ยนว่า error แบบใดเกิดบ่อยพอให้โมเดลเรียนรู้

E. Build an honest baseline

Start with a transparent model and a minimal feature set. The baseline establishes how much signal exists before complex architecture is introduced. A larger model earns its place only when it improves a relevant outcome, survives independent testing and adds value that cannot be obtained by threshold tuning alone.

เริ่มด้วยโมเดลโปร่งใสและ feature set ขั้นต่ำ baseline บอกว่ามี signal เท่าไร ก่อนเพิ่ม architecture ที่ซับซ้อน โมเดลใหญ่ต้องพิสูจน์ตัวเองด้วย outcome ที่สำคัญ independent test และ value ที่ไม่สามารถได้จาก threshold tuning เพียงอย่างเดียว

F. Separate representation from classifier

Test whether improvement comes from better biological representation or from a different decision algorithm. Cross the feature sets and model families systematically. Without this separation, a study may attribute success to the classifier when the real contribution is a composite feature, an embedding or a cleaner dataset.

แยกให้ได้ว่า improvement มาจาก biological representation ที่ดีขึ้นหรือ decision algorithm ที่ต่างออกไป ทดลองข้าม feature set กับ model family อย่างเป็นระบบ มิฉะนั้นงานอาจให้เครดิต classifier ทั้งที่ contribution จริงคือ composite feature, embedding หรือ dataset ที่สะอาดขึ้น

G. Perform ablation studies

Remove one component at a time: a feature family, balancing method, base learner, voting rule or post-processing stage. Measure not only average performance but also subgroup behavior. Ablation converts architecture from a diagram into an explanation of which component does what and under which conditions.

ถอด component ทีละส่วน ทั้ง feature family, balancing method, base learner, voting rule หรือ post-processing วัดทั้ง average performance และ subgroup behavior ablation เปลี่ยน architecture จากภาพประกอบให้เป็นคำอธิบายว่า component ใดทำอะไรและทำงานในเงื่อนไขแบบไหน

H. Calibrate probabilities

A candidate score should mean something operational. Use reliability plots, calibration error and decision curves to check whether a reported probability corresponds to observed frequency. Calibrated outputs help collaborators choose thresholds according to laboratory capacity and tolerance for missed candidates.

candidate score ควรมี operational meaning ใช้ reliability plot, calibration error และ decision curve ตรวจว่า probability สอดคล้องกับ observed frequency หรือไม่ calibrated output ช่วย collaborator เลือก threshold ตาม laboratory capacity และความยอมรับต่อ missed candidate

I. Design independent tests

Use data that differ in time, source, species, laboratory or family from the training set. Explain exactly what independence means. One external dataset is more informative when its relationship to training data is understood; a dataset called independent without overlap analysis may still be quietly familiar.

ใช้ข้อมูลที่ต่างจาก training ในด้านเวลา แหล่ง species, laboratory หรือ family และอธิบายให้ชัดว่า independence หมายถึงอะไร external dataset มีข้อมูลมากขึ้นเมื่อเข้าใจความสัมพันธ์กับ training ส่วน dataset ที่เรียก independent โดยไม่วิเคราะห์ overlap อาจยังคุ้นเคยกับโมเดลอยู่

J. Read the confusion matrix biologically

Inspect false positives and false negatives by length, family, species, composition, structural class and annotation confidence. Ask whether errors cluster around genuinely ambiguous biology. Some errors reveal missing features; others reveal label uncertainty or a boundary that nature itself does not draw sharply.

อ่าน false positive และ false negative ตาม length, family, species, composition, structural class และ annotation confidence ถามว่า error กระจุกใน biology ที่กำกวมจริงหรือไม่ บาง error ชี้ missing feature บางส่วนชี้ label uncertainty หรือ boundary ที่ธรรมชาติไม่ได้แบ่งคมชัด

K. Compare at matched conditions

Do not place metrics from different papers in one league table without checking datasets, splits, redundancy control, label definitions and software versions. Re-evaluate competing tools on a shared test when possible. Fair comparison is experimental design, not formatting.

อย่าวาง metric จาก paper ต่างกันใน league table เดียวโดยไม่ตรวจ dataset, split, redundancy control, label definition และ software version หากทำได้ควร re-evaluate tool ต่าง ๆ บน shared test fair comparison เป็น experimental design ไม่ใช่ formatting

L. Track computational cost

Measure feature-computation time, memory, training time, inference latency and dependencies. A model intended for genome-scale screening has different constraints from a small candidate-ranking tool. Scientific usefulness includes whether another group can run the method with realistic resources.

วัด feature-computation time, memory, training time, inference latency และ dependency โมเดลสำหรับ genome-scale screening มี constraint ต่างจาก candidate-ranking tool ขนาดเล็ก scientific usefulness รวมถึงความสามารถที่กลุ่มอื่นจะรันวิธีนี้ด้วยทรัพยากรจริง

M. Preserve software context

Archive code, environment versions, parameter files, feature definitions and example inputs. A web interface is helpful, but the reproducible object is the complete path from raw input to output. Migration should retain legacy pages and downloads while documenting how a maintained replacement will differ.

เก็บ code, environment version, parameter file, feature definition และ example input web interface มีประโยชน์ แต่ reproducible object คือเส้นทางครบจาก raw input ถึง output migration ควรเก็บ legacy page และ download พร้อมบอกว่า maintained replacement จะแตกต่างอย่างไร

N. Plan a wet-lab handoff

Define how many candidates can be tested, what experimental assay is suitable, what controls are required and how computational uncertainty will influence selection. The model should produce a ranked, documented handoff rather than an unexplained spreadsheet of positives.

กำหนดจำนวน candidate ที่ทดสอบได้ assay ที่เหมาะสม control ที่ต้องมี และวิธีใช้ computational uncertainty ในการเลือก โมเดลควรส่งมอบ ranked, documented handoff ไม่ใช่ spreadsheet ของ positive ที่ไม่มีคำอธิบาย

O. Treat ethics as method

For biological prediction, ethics includes responsible claims, transparent uncertainty, appropriate use of public data, licensing, biosafety and avoidance of clinical overstatement. These are methodological constraints because they shape which outputs should be produced and how they may be communicated.

ethics ใน biological prediction รวม responsible claim, transparent uncertainty, การใช้ public data และ license ที่เหมาะสม biosafety และการไม่กล่าวเกิน clinical evidence สิ่งเหล่านี้เป็น methodological constraint เพราะกำหนดว่า output ใดควรสร้างและควรสื่อสารอย่างไร

P. Convert limitations into experiments

A limitation becomes useful when translated into a variable, a comparison and an observable outcome. “More data are needed” is weak; “test performance on post-2024 sequences from unseen families with homology below a stated threshold” is a research plan.

limitation มีประโยชน์เมื่อเปลี่ยนเป็น variable, comparison และ observable outcome คำว่า more data are needed ยังอ่อน แต่การทดสอบ post-2024 sequence จาก unseen family ที่ homology ต่ำกว่า threshold ที่ระบุคือ research plan

Q. Build a student contribution map

Separate tasks into data curation, biological interpretation, feature engineering, model development, software design, visualization, reproducibility and experimental coordination. Students can enter from different strengths while still understanding how their contribution changes the whole evidence chain.

แยกงานเป็น data curation, biological interpretation, feature engineering, model development, software design, visualization, reproducibility และ experimental coordination นักศึกษาเข้าจาก strength ต่างกันได้ แต่ต้องเข้าใจว่า contribution ของตนเปลี่ยน evidence chain ทั้งระบบอย่างไร

R. Define success before training

Pre-register the primary metric, secondary metrics, test set, subgroup analyses and stopping rule. Decide what improvement would matter scientifically, not just statistically. This protects the project from selecting whichever result looks best after many experiments.

กำหนด primary metric, secondary metric, test set, subgroup analysis และ stopping rule ก่อน train ตัดสินว่า improvement แบบใดมีความหมายทางวิทยาศาสตร์ ไม่ใช่เพียง statistical วิธีนี้ป้องกันการเลือกเฉพาะผลที่ดูดีที่สุดหลังทดลองหลายครั้ง

S. Communicate uncertainty visibly

Show ranges, confidence intervals, calibration and known failure regions near the prediction, not hidden at the end of documentation. A user who sees uncertainty can make a better decision than a user given a confident label detached from evidence.

แสดง range, confidence interval, calibration และ known failure region ใกล้ prediction ไม่ใช่ซ่อนไว้ท้าย documentation ผู้ใช้ที่เห็น uncertainty ตัดสินใจได้ดีกว่าผู้ใช้ที่ได้รับ confident label ซึ่งหลุดจาก evidence

T. Ask what the model changed

At the end, identify the decision that became faster, cheaper, more reliable or newly possible. If no downstream decision changes, the model may still be an interesting benchmark, but its practical claim should remain modest. Impact begins where a prediction changes responsible action.

ท้ายที่สุดต้องบอกว่า decision ใดเร็วขึ้น ถูกลง น่าเชื่อถือขึ้น หรือเพิ่งเป็นไปได้ ถ้าไม่มี downstream decision เปลี่ยน โมเดลอาจยังเป็น benchmark ที่น่าสนใจ แต่ practical claim ควรพอดี impact เริ่มตรงที่ prediction เปลี่ยน responsible action

Research downloadsไฟล์ดาวน์โหลดงานวิจัย

Datasets, documentation and research tools

ชุดข้อมูล เอกสาร และเครื่องมือวิจัย

Download the available resources for study, reproducibility and further research. Please review the documentation accompanying each file before use.

ดาวน์โหลด resource ที่เกี่ยวข้องเพื่อการศึกษา การทำซ้ำผล และการต่อยอดงานวิจัย กรุณาอ่านเอกสารประกอบของแต่ละไฟล์ก่อนนำไปใช้

Research interpretationการตีความงานวิจัย

The result is useful. The structure of the reasoning is reusable.

ผลลัพธ์มีประโยชน์ แต่โครงสร้างการคิดนำกลับมาใช้ได้ไกลกว่า

A composite feature is a small theory embedded inside a larger model. Logistic regression decides how five biological signals should travel together; random forest then learns the nonlinear boundary. The architecture is hybrid because interpretation and flexibility are split deliberately.

ตัวเลขบอกว่าระบบทำงานได้ในเงื่อนไขหนึ่ง แต่ intellectual value อยู่ที่การอธิบายว่าทำไม representation และ architecture แบบนี้จึงสร้างผลลัพธ์นั้นได้ เมื่ออ่านแบบนี้ paper จะไม่จบที่คะแนน แต่กลายเป็นจุดเริ่มต้นของ experiment ถัดไป

Evidence boundaryขอบเขตของหลักฐาน

Where confidence should stop

จุดที่ความมั่นใจควรหยุด

Genome-wide windows create a different distribution from curated training sequences. Performance varies between prokaryotic and eukaryotic settings, and uncharacterized ncRNAs cannot be counted as ordinary false positives with certainty.

ผลลัพธ์ไม่ใช่คำสัญญาว่าจะทำงานเหมือนเดิมกับทุก species, database version หรือ experimental setting การต่อยอดที่ดีต้องเปลี่ยนเงื่อนไขสำคัญทีละส่วน แล้วดูว่าคำอธิบายเดิมยังยืนอยู่หรือไม่

Student research tracksเส้นทางต่อยอดสำหรับนักศึกษา

Four projects hiding inside the next question

สี่โครงงานที่ซ่อนอยู่ในคำถามถัดไป

Begin with reproduction. Earn the right to add complexity by first understanding the baseline, its errors and its evidence boundary.

เริ่มจาก reproduction ก่อน ทำความเข้าใจ baseline, error และ evidence boundary ให้ชัด แล้วค่อยเพิ่ม complexity อย่างมีเหตุผล ไม่ใช่เพราะโมเดลใหม่ดูน่าตื่นเต้นกว่า

01 · Evaluate on current Rfam and lncRNA annotations
02 · Add strand-aware and context-aware genomic features
03 · Compare SCORE with learned RNA embeddings
04 · Design candidate-ranking interfaces for experimental collaborators

Who belongs in this research space?ใครเหมาะกับพื้นที่วิจัยนี้

Curiosity first. Discipline immediately after.

เริ่มด้วยความอยากรู้ แล้วตามด้วยวินัยทันที

Pattern seeker

You enjoy finding structure in sequences and asking whether the pattern is biological, statistical or accidental.

คุณชอบหา pattern ใน sequence และถามต่อว่ามันเป็น biological, statistical หรือ accidental

Systems thinker

You naturally connect data quality, feature design, algorithms, software and experiments instead of optimizing one box.

คุณเชื่อม data quality, feature design, algorithm, software และ experiment แทนที่จะ optimize เพียงกล่องเดียว

Quietly persistent builder

You do not need constant certainty. You can debug, read, rerun and refine until the evidence becomes clear.

คุณไม่ต้องมั่นใจตลอดเวลา แต่สามารถ debug อ่าน rerun และ refine จนหลักฐานชัดขึ้น

Bridge researcher

You want enough biology to ask a meaningful question and enough computation to test it honestly.

คุณอยากเข้าใจชีววิทยาพอที่จะตั้งคำถามมีความหมาย และเข้าใจ computation พอที่จะทดสอบอย่างตรงไปตรงมา

Publication recordข้อมูลการตีพิมพ์

Research metadata

ข้อมูลบรรณานุกรม

APALertampaiporn, S., Thammarongtham, C., Nukoolkit, C., Kaewkamnerdpong, B., Ruengjitchatchawalya, M. (2014). Identification of Non-Coding RNAs with a New Composite Feature in the Hybrid Random Forest Ensemble Algorithm. Nucleic Acids Research, 42(11), e93. https://doi.org/10.1093/nar/gku325IEEES. Lertampaiporn, C. Thammarongtham, C. Nukoolkit, B. Kaewkamnerdpong, M. Ruengjitchatchawalya, “Identification of Non-Coding RNAs with a New Composite Feature in the Hybrid Random Forest Ensemble Algorithm,” Nucleic Acids Research, vol. 42, no. 11, e93, 2014, doi: 10.1093/nar/gku325.BibTeX@article{lertampaiporn2014hybrid, title={Identification of Non-Coding RNAs with a New Composite Feature in the Hybrid Random Forest Ensemble Algorithm}, author={Supatcha Lertampaiporn and Chinae Thammarongtham and Chakarida Nukoolkit and Boonserm Kaewkamnerdpong and Marasri Ruengjitchatchawalya}, journal={Nucleic Acids Research}, volume={42}, number={11}, pages={e93}, year={2014}, doi={10.1093/nar/gku325} }Open official articleเปิดบทความทางการ

Research FAQ

Questions a serious reader should keep open

คำถามที่ผู้อ่านงานวิจัยควรเปิดไว้

What is the contribution?

The complete relationship between biological problem, representation, model architecture and evaluation—not the presence of a familiar algorithm alone.

ความสัมพันธ์ระหว่างปัญหาชีววิทยา representation, model architecture และ evaluation ไม่ใช่เพียงการมีอัลกอริทึมที่คุ้นเคย

Can headline metrics be compared?

Only when datasets, redundancy control, negative definitions, splits and evaluation protocols are comparable.

เปรียบเทียบได้เมื่อ dataset, redundancy control, negative definition, split และ evaluation protocol ใกล้เคียงกัน

Does prediction prove function?

No. It prioritizes evidence. Biological function still requires appropriate experimental confirmation.

ไม่ prediction ทำหน้าที่จัดลำดับหลักฐาน ส่วน biological function ยังต้องยืนยันด้วยการทดลองที่เหมาะสม

Where should a student begin?

Reproduce the smallest end-to-end baseline, verify one published metric and perform error analysis before changing the model.

เริ่มจาก baseline end-to-end ที่เล็กที่สุด ตรวจสอบ metric หนึ่งตัว และทำ error analysis ก่อนเปลี่ยนโมเดล

Your question could be nextคำถามถัดไปอาจเป็นของคุณ

Join a lab where biology and computation challenge each other.

เข้าร่วมพื้นที่ที่ชีววิทยาและ computation ตั้งคำถามให้กัน

Bring a sequence, a biological puzzle, an imperfect dataset or the patience to make a difficult question testable.

นำ sequence, biological puzzle, dataset ที่ยังไม่สมบูรณ์ หรือความอดทนที่จะเปลี่ยนคำถามยากให้ทดสอบได้เข้ามา

Talk to the labรู้จักห้องปฏิบัติการ