ประกอบค่าที่มีความหมายอยู่แล้ว
risk_pressure = volatility × illiquidity × exposureDependency ทุกตัวต้องมีเวลา หน่วย Population และ Version ที่เข้ากันได้ การนำ Feature มาคูณกันไม่ได้ทำให้ Composite มีความหมายโดยอัตโนมัติ
FEATURE DAG · SHARED STATE · MATERIALIZATION
Composite Feature ประกอบจาก Observation ที่ใช้ซ้ำได้ ส่วน Feature ที่มี Shared Structure อาศัย Normalization, Window, Graph, Spectrum, Tokenization หรือ Embedding เดียวกัน งานสำคัญคือค้นหาโครงสร้างร่วม คำนวณเพียงครั้งเดียว และรักษาความหมายให้คงเดิม
risk_pressure = volatility × illiquidity × exposureDependency ทุกตัวต้องมีเวลา หน่วย Population และ Version ที่เข้ากันได้ การนำ Feature มาคูณกันไม่ได้ทำให้ Composite มีความหมายโดยอัตโนมัติ
FFT(window) → RMS · energy · fault bandsการ Reuse เกิดที่ Intermediate Representation แม้ Final Feature จะมีสูตรและวัตถุประสงค์ต่างกัน
01 · SEE THE FEATURE DAG
คำนวณร่วมได้เมื่อ Source, Entity, Event-time policy, Window, Parameters และ Version ตรงกัน หากต่างกันเพียงชื่อแต่ความหมายต่าง ห้าม Share อย่างเงียบๆ
02 · COMPOSITE FEATURE ANATOMY
8 orders0.220.758 × .22 × .75 = 1.32ค่าที่ประกอบกันสัมพันธ์กับแนวคิดเดียวกันจริงหรือเพียง Correlate ในชุดข้อมูลหนึ่ง?
ผลลัพธ์มีหน่วยอะไร ต้อง Normalize ก่อนบวก/คูณหรือไม่?
ทุก Dependency มาจาก Cut-off เดียวกันและพร้อมใช้ก่อน Decision Time หรือไม่?
Feature ตัวเดียว Drift หรือ Missing จะทำให้ Composite เปลี่ยนเพียงใด?
03 · COST OF RECOMPUTATION
เป็น Conceptual Cost Model สำหรับสอน ไม่ใช่ Benchmark ของระบบใดระบบหนึ่ง
F × (S + A)S + F × A04 · TEN SHARED-STRUCTURE FAMILIES
ต้นทุนที่ควรมองหาไม่ได้มีเพียงสูตร Final Feature แต่รวมการอ่านข้อมูล การ Join, Sort, Normalize, Window, FFT, Tokenize, Build Graph, Spatial Index และ Model Inference ซึ่งมักแพงกว่า Final Aggregation หลายเท่า เลือกแต่ละกลุ่มเพื่อดูเส้นทางคำนวณและ Reuse Contract
ขั้นตอนนี้ถูกเรียกโดย Feature หรือ Model กี่ตัวต่อวัน?
ต้นทุนอยู่ที่ CPU/GPU, I/O, Shuffle, Memory หรือ External API?
Intermediate เปลี่ยนบ่อยแค่ไหน และตรวจ Invalid ได้หรือไม่?
ทุก Consumer ต้องการนิยาม เวลา Population และ Version เดียวกันจริงหรือไม่?
Adjusted price → returns → rolling windows50 ล้าน OHLCV rows · 5,000 assets · 10 yearsmomentum_20drealized_vol_20dRSI_14rolling_beta_60dคำนวณ Return และ Window State ครั้งเดียว ไม่อ่านราคาใหม่ทุก Feature
ถ้า 12 Features อ่านและปรับราคาแยกกัน จะ scan 600 ล้าน row-equivalents; ถ้าสร้าง adjusted return series ครั้งเดียว scan ฐาน 50 ล้านแถว แล้วทำ final aggregation ที่เบากว่า
materialize adjusted_returns_1d และ rolling_state(sum, sum_sq, count, min, max)
instrument_id + trading_date + adjustment_policy + return_versionSplit/dividend revision, late price correction, calendar หรือ benchmark version เปลี่ยน
Calibrated waveform → window → FFT100 sensors × 20 kHz × 8 hours ≈ 57.6 พันล้าน samplesRMSkurtosiscrest_factorband_energyfault_frequency_powerFFT เป็นงานแพง ควรสร้าง Spectrum หนึ่งครั้งแล้วสรุปหลาย Feature
5 spectral Features ที่รัน FFT แยกกัน = FFT 5 ครั้ง/window; shared spectrum ลดเหลือ 1 FFT + 5 cheap reductions
windowed_fft หรือ selected spectral bands; raw waveform เก็บตาม retention policy
machine_id + sensor_id + window_start + sample_rate + calibration_version + fft_configCalibration, sample rate, window length/overlap หรือ window function เปลี่ยน
Canonical sequence → tokens/counts1 ล้าน sequences × ความยาวเฉลี่ย 1,000 basesGC_contentk_mer_frequencycodon_usagesequence_entropyการ Scan ลำดับและนับ Token เป็นโครงสร้างร่วมของหลาย Feature
GC, entropy, 3-mer และ 6-mer ที่ scan แยกกันอ่าน sequence ซ้ำ; one-pass counter อัปเดต composition และหลาย k พร้อมกันได้
canonical_sequence_hash, length, base_counts, canonical_kmer_counts
sequence_hash + alphabet + strand_policy + k + tokenizer_versionSequence/database revision, ambiguity policy, strand convention หรือ genome build เปลี่ยน
Normalized document → tokens → corpus statistics10 ล้าน documents · หลายพันล้าน tokensTF_IDFBM25entity_densityreadabilityembedding_inputTokenization และ Corpus Statistics ต้องมี Version เดียวกันทั้ง Pipeline
TF-IDF กับ BM25 ไม่ควร tokenize corpus คนละรอบ; entity density/readability reuse sentence/token boundaries; embedding reuse normalized model input
document tokens, offsets, sentence boundaries, corpus DF และ embedding เมื่อมี reuse สูง
document_hash + language + tokenizer_version + corpus_snapshot + model_versionDocument edit, tokenizer/model/corpus reference หรือ privacy redaction เปลี่ยน
Clean transaction lines → customer timeline500 ล้าน transaction lines · หลาย channel · returns/cancellationsRFMbasket_sizecategory_affinitypromo_exposurestockout_adjusted_demandNormalize คืนสินค้าและ Identity ครั้งเดียวก่อน Aggregate หลายแบบ
RFM, affinity, basket และ promotion exposure ที่ join transaction ใหม่ทุกตัวทำให้ logic คืนสินค้าซ้ำและอาจไม่ตรงกัน; shared clean line fact ลดทั้ง compute และ semantic drift
canonical_transaction_line และ customer_daily_aggregate
customer_id_version + order_id + event_date + currency_policy + logic_versionLate return, identity merge/split, exchange-rate revision หรือ product taxonomy เปลี่ยน
Packets/logs → sessionized flowsล้าน packets/sec หรือหลายพันล้าน log events/dayflow_durationbytes_per_packetdestination_entropylogin_velocityrare_process_scoreSessionization เป็น Shared State; ห้ามให้แต่ละ Feature นิยาม Timeout ต่างกันโดยไม่ตั้งใจ
duration, bytes/packet, entropy และ velocity ควรอ่าน flow state เดียวกัน; sessionize แยกอาจให้ Timeout และ Direction ต่างกันจน Feature ขัดกัน
canonical_flow, identity_event และ rolling_counter state ตาม retention
sensor + five_tuple + direction + window + timeout_policy + parser_versionLate packet, NAT mapping, asset identity, timeout/parser หรือ detection-policy เปลี่ยน
Edges → canonical graph snapshot100 ล้าน nodes · 1 พันล้าน edges ต่อ snapshotdegreecentralitycycle_featurescommunityneighbor_countAdjacency และ Snapshot แพงกว่าสรุป Feature จึงควร reuse และ version
Degree, neighbor count และ cycle detector reuse adjacency; centrality/community อาจ reuse snapshot และ partitions แม้ algorithm ต่าง
graph_snapshot, adjacency partitions, node mapping และ selected expensive embeddings
graph_id + snapshot_time + edge_policy + directed_flag + algorithm_versionEdge correction, snapshot cutoff, directed/weighted policy หรือ node-resolution เปลี่ยน
Coordinates → projected/indexed geometry100 ล้าน GPS points · 5 ล้าน POIs/road segmentsdistancespeedPOI_densityregion_membershipdwell_timeProjection และ Spatial Index ใช้ร่วมกันได้หลาย Spatial Join
Speed, dwell, region membership และ POI density ไม่ควร project/map-match จุดเดิมใหม่ทุก Feature; reuse matched trajectory และ index
projected_point, matched_segment, trajectory_session และ versioned spatial index
entity + event_time + CRS + map_version + matching_configBase map/POI revision, CRS, GPS correction, trajectory gap หรือ matching threshold เปลี่ยน
Decoded image → normalized tensor/backbone10 ล้าน images หรือ video หลายล้าน framescolor_histogramedge_densitytextureshapeCNN_embeddingอย่า Decode/Resize หรือรัน Backbone ซ้ำเพื่อ Head แต่ละตัว
หลาย Classification/Detection Heads ไม่ควรรัน backbone ซ้ำ; color/edge/texture ก็ reuse decoded normalized image แม้ branch ต่าง
normalized tensor, augmentation manifest และ backbone activation ตามชั้นที่กำหนด
image_hash + crop/resize + color_space + weights + layer + preprocessing_versionImage/crop, augmentation seed, model weights/layer หรือ normalization เปลี่ยน
Normalized clinical events → patient windowsพันล้าน clinical events จากหลายระบบและหลายหน่วยlab_deltamedication_exposurecomorbidity_countvital_trendvisit_frequencyUnit normalization และ Point-in-time Timeline เป็นฐานร่วมที่ต้องกำกับอย่างเข้มงวด
lab delta, vital trend, medication exposure และ visit frequency ควรใช้ timeline ที่ normalize ครั้งเดียว มิฉะนั้นหน่วยและ availability cutoff อาจไม่ตรงกัน
normalized clinical event และ patient window index; จำกัดตาม governance
patient_key_version + event_id + event_time + available_at + terminology/unit_versionLate result, corrected lab, identity merge, terminology mapping, consent หรือ access policy เปลี่ยน
05 · IMPLEMENTATION PATTERNS
Canonicalize Expression และหา Common Subexpression ก่อนสร้าง Execution Plan
return → rolling_sum
return → rolling_sum_sqรักษา sum, count, sum², min/max queue ต่อ Entity แล้วอัปเดตแบบ Incremental
update: O(1) / eventเก็บ FFT, Embedding หรือ Graph Snapshot เมื่อแพง ใช้บ่อย และมี Freshness Contract ชัดเจน
key = entity + as_of + versionใช้ Definition เดียวกันแต่ Execution ต่างกัน และทำ Parity Test ระหว่าง Batch กับ Streaming
|offline-online| ≤ tolerance06 · WHEN SHARING BECOMES WRONG
30 calendar days ≠ 30 observations ≠ 30 trading days
Event Time เดียวกัน แต่ข้อมูลหนึ่งออกผลภายหลัง
Z-score เทียบลูกค้าทั้งหมด ≠ เทียบลูกค้าใน Segment
Raw price ≠ split-adjusted ≠ total-return adjusted
Tokenizer, Genome Build, CRS หรือ Model Weight ต่างกัน
Intermediate ที่สร้างจากข้อมูลอนาคตถูกนำไปใช้ใน Training Row อดีต
07 · DESIGN CHECKLIST
FINAL PRINCIPLE
เป้าหมายไม่ใช่ลดจำนวนสูตร แต่ลดการทำงานซ้ำโดยไม่ทำลายความหมาย ความถูกต้องตามเวลา และความสามารถในการสร้างซ้ำ