ออกแบบ
แปลงเจตนาของโมเดลเป็นนิยาม feature ที่ชัดเจน
ML DATA ENGINEERING · FEATURE PLATFORM
Feature ที่มีประโยชน์ไม่ใช่เพียงสูตรใน notebook แต่เป็น data product ที่เข้าใจเวลา ผ่านการทดสอบ สร้างซ้ำได้ และต้องให้ความหมายสอดคล้องกันทั้งตอน train และตอนทำนาย Feature Store คือระบบที่ทำให้คำสัญญานั้นดูแลได้
แปลงเจตนาของโมเดลเป็นนิยาม feature ที่ชัดเจน
สร้างค่าประวัติศาสตร์ซ้ำโดยไม่รั่วข้อมูลจากอนาคต
ส่งมอบความหมายเดียวกันให้ batch training และ online inference
ติดตาม owner, lineage, freshness, version และคุณภาพ
01 · THE ECONOMICS
อัตราส่วนหนึ่งค่าอาจใช้ SQL เพียงบรรทัดเดียว แต่การทำให้ถูกต้องสำหรับข้อมูล train ย้อนหลังห้าปี ตอบได้ในไม่กี่ millisecond ถูก monitor ทุกวัน และมีทีมอื่นเข้าใจตรงกัน คืออีกปัญหาหนึ่งทางวิศวกรรม
สัมภาษณ์ผู้รู้ วิเคราะห์ข้อมูล ทดลอง และสร้าง feature จำนวนมากที่ไม่เคยไปถึง production
join ขนาดใหญ่ rolling window, backfill และ training dataset ที่สร้างซ้ำ ใช้ทั้ง compute และเวลาวิศวกร
Point-in-time join ต้องใช้เฉพาะข้อมูลที่มีอยู่ ณ เวลาที่การทำนายนั้นควรเกิดขึ้น
หลังเปิดใช้ยังมี freshness, pipeline ล้มเหลว schema เปลี่ยน skew, drift, latency, incident และ ownership
หลายทีมสร้าง “customer activity” ใหม่ด้วย window, filter และนโยบาย missing value ที่ต่างกัน
Leakage หรือ training-serving skew อาจทำให้คะแนน offline ดูดี แต่การทำนายจริงแย่ลง
02 · FROM COLUMN TO DATA PRODUCT
“orders_30d” ยังไม่สมบูรณ์ หากไม่รู้ entity, event time, กฎการนับ, นโยบาย null, freshness, owner และ version
ช่วงเวลาสิ้นสุดที่ t แต่ไม่รวม t เพื่อไม่ให้เหตุการณ์ ณ เวลาทำนายหลุดเข้า training row ก่อนเวลาที่ระบบควรรู้เหตุการณ์นั้น
03 · WHAT A FEATURE STORE DOES
Feature Store ไม่ใช่เพียงฐานข้อมูลที่ตั้งชื่อใหม่ให้ทันสมัย แต่เป็นระบบข้อมูลสำหรับ ML ที่เชื่อม registry, historical retrieval และในกรณี latency ต่ำ ก็รวม online serving
04 · OFFLINE AND ONLINE ARE DIFFERENT JOBS
Training ต้องการ dataset ประวัติศาสตร์ขนาดใหญ่ที่ถูกต้อง ส่วน real-time inference อาจต้องการค่าล่าสุดในระดับ millisecond Store ประสานความหมาย แต่ไม่ได้ลบ trade-off ทางกายภาพ
เหมาะกับ scan, history, point-in-time retrieval, backfill และการสร้าง training set มักอยู่บน warehouse หรือ lakehouse
เหมาะกับ keyed lookup และ latency ต่ำที่คาดการณ์ได้ โดยทั่วไปเก็บค่าล่าสุด ไม่ได้เก็บประวัติศาสตร์ทั้งหมด
อธิบายตัวตนและเจ้าของ Feature ช่วยให้ค้นพบและใช้ซ้ำ แต่ไม่สามารถทำให้นิยามที่ไม่ดีมีความหมายขึ้นมาเอง
05 · LAB: DEFINE HISTORICAL FEATURES
ตัวอย่างแบบ PostgreSQL-compatible นี้ใช้สอนหลักการ ระบบ production ยังต้องออกแบบ storage, incremental computation และ orchestration ที่ผ่านการทดสอบ
CREATE TABLE ml.training_events (
customer_id INTEGER NOT NULL,
prediction_time TIMESTAMP NOT NULL,
label INTEGER
);
-- One row represents one customer at one prediction time.
CREATE VIEW ml.v_customer_features_training AS
SELECT
t.customer_id,
t.prediction_time,
t.label,
COUNT(o.order_id) FILTER (
WHERE o.status = 'completed'
) AS completed_orders_30d,
COALESCE(SUM(o.amount) FILTER (
WHERE o.status = 'completed'
), 0) AS completed_value_30d,
MAX(o.ordered_at) AS latest_known_order_time
FROM ml.training_events AS t
LEFT JOIN raw.orders AS o
ON o.customer_id = t.customer_id
AND o.ordered_at >= t.prediction_time - INTERVAL '30 days'
AND o.ordered_at < t.prediction_time
GROUP BY t.customer_id, t.prediction_time, t.label;06 · LAB: MATERIALIZE THE LATEST VALUES
Primary key แทน feature vector ล่าสุดหนึ่งชุดต่อ entity ส่วน online store จริงยังต้องมี atomic write, freshness guarantee, availability, latency SLO และ recovery
CREATE TABLE ml.customer_features_online (
customer_id INTEGER PRIMARY KEY,
completed_orders_30d INTEGER NOT NULL,
completed_value_30d DECIMAL(14,2) NOT NULL,
feature_timestamp TIMESTAMP NOT NULL,
feature_version VARCHAR(20) NOT NULL
);
INSERT INTO ml.customer_features_online
SELECT
c.customer_id,
COUNT(o.order_id) FILTER (WHERE o.status='completed'),
COALESCE(SUM(o.amount) FILTER (WHERE o.status='completed'), 0),
CURRENT_TIMESTAMP,
'customer_activity_v1'
FROM raw.customers c
LEFT JOIN raw.orders o
ON o.customer_id=c.customer_id
AND o.ordered_at >= CURRENT_TIMESTAMP - INTERVAL '30 days'
GROUP BY c.customer_id
ON CONFLICT (customer_id) DO UPDATE SET
completed_orders_30d=EXCLUDED.completed_orders_30d,
completed_value_30d=EXCLUDED.completed_value_30d,
feature_timestamp=EXCLUDED.feature_timestamp,
feature_version=EXCLUDED.feature_version;07 · WHEN A FEATURE STORE EARNS ITS COST
Feature Store มีต้นทุนของ platform เอง จึงคุ้มเมื่อ reuse, time correctness, serving consistency และ governance ช่วยประหยัดมากกว่าค่าใช้จ่ายในการดูแลระบบ
08 · OPERATING CHECKLIST
Feature platform ควรลดงานที่ซ่อนอยู่ ไม่ใช่ซ่อนกระบวนการจนไม่มีใครตรวจสอบสมมติฐานได้
ทดสอบว่า historical row ไม่อ่านเหตุการณ์ในอนาคต
เปรียบเทียบค่า offline และ online ของ entity ตัวอย่าง
วัดอายุ feature และ retrieval latency เทียบ SLO ที่ระบุชัด
ติดตาม null rate, range, distribution shift และ entity coverage
สืบย้อน code, source, window, owner และโมเดลที่ใช้แต่ละ version
วัด compute และ storage แล้วเลิกใช้ feature ที่ไม่มีผู้ใช้อย่างปลอดภัย
CHALLENGES
หาจุดที่เกิด leakage ใน historical join ซึ่งใช้ CURRENT_TIMESTAMP แทน prediction_time
นิยาม “customer spend in 7 days” ให้ครบ entity, ขอบเขต window, event time, default และ owner
ออกแบบ query ตรวจ offline-to-online parity และกำหนด mismatch threshold ที่ยอมรับได้
ออกแบบ backfill เมื่อกฎ cancelled order เปลี่ยน พร้อมระบุโมเดลและ dataset ที่ได้รับผล
ประมาณต้นทุนรายเดือนของ 200 features สำหรับ 20 ล้าน entities โดยเขียนสมมติฐานทุกข้อ
โต้แย้งการใช้ Feature Store ในโครงการ batch ขนาดเล็ก แล้วระบุเงื่อนไขที่จะทำให้เปลี่ยนการตัดสินใจ
THE CENTRAL IDEA