- Home
- Courses
- ML in Practice — Lightweight MLOps
- Monitoring, Retraining, and Lifecycle
Micro Free Course — 100% Free Learning
All lessons in this module are free to learn. Sign in with Google to save your progress.
Monitoring, Retraining, and Lifecycle
Data drift vs concept drift, performance monitoring when labels are delayed, retraining strategies (scheduled / triggered / continual), incident response and rollback playbooks.
Module Content
Drift Detection — Data Drift vs Concept Drift
Two distinct kinds of drift. Different symptoms, different fixes. FIND_VIDEO: search 'data drift concept drift machine learning' — recommended channel: Evidently / Made with ML. Aim for 10 min or under.
Recap — Detecting and Acting on Drift
Most production ML systems degrade slowly. The slow degradation has two causes; detection is different for each.
Performance Monitoring When Labels Are Delayed
Most real ML labels arrive in days, weeks, or months. Don't wait — use proxies. FIND_VIDEO: search 'ml monitoring delayed labels proxy metrics' — recommended channel: Made with ML / Chip Huyen. Aim for 10 min or under.
Recap — Proxies and Indirect Metrics
When you can't measure accuracy directly, you measure things that correlate with accuracy. Done well, this catches problems before labels arrive.
Retraining Strategies — Scheduled, Triggered, Continual
Three retraining patterns. Pick by your data velocity and drift rate. FIND_VIDEO: search 'retraining strategy scheduled triggered continual learning' — recommended channel: Made with ML / Chip Huyen. Aim for 10 min or under.
Recap — When and How to Retrain
Retraining cadence is a system design decision, not a script-rerun. Each strategy has cost and risk profiles.
Incidents and Rollbacks — When Models Go Wrong
The plan you wish you had before the model started misbehaving in production. FIND_VIDEO: search 'ml incident response model rollback' — recommended channel: Made with ML / Chip Huyen. Aim for 10 min or under.
Recap — Incident Response Playbook for ML
ML incidents are different from software incidents — silent failures, cascading errors. The playbook that handles both.