MLOps Roadmap
From data contracts to production deployment, monitoring and rollback — build a repeatable ML delivery process.
MLOps is not a tool — it is the process that turns a notebook experiment into a production system you can deploy, monitor and recover from. This roadmap covers the full lifecycle: data versioning and contracts, reproducible training with experiment tracking, model packaging, serving patterns, pipeline orchestration, CI/CD with rollback, monitoring for both service health and model quality, and governance with retraining. Each stage has a build task, acceptance checks and a measurable outcome. By the end, you can take a model from experiment to production and keep it running.
What is the right MLOps roadmap?
The right MLOps roadmap does not start with tools — it starts with a question: can you take a model from a data scientist's notebook to a production service, keep it running and recover when it breaks? That means versioning your data so training is reproducible, packaging the model with its preprocessing so inference matches training, choosing a serving pattern that fits your latency budget, automating releases with gates that block bad models, monitoring both service health and prediction quality, and treating retraining as a candidate — not an automatic promotion. If you cannot reproduce your promoted model from recorded inputs, you do not have MLOps — you have a deployment script.
Sources and methodology · This roadmap is reviewed when production practices, tools or platform patterns materially change.
Stages
9
Last reviewed
16 September 2026
Stage 1: ML lifecycle and operational foundations
Before you touch any MLOps tool, understand how a model travels from a data scientist's notebook to a production service that real users depend on. This stage maps the full lifecycle — training artifacts, inference artifacts, the people who own each stage, and the failure points where things go wrong.
Engineers who skip lifecycle understanding end up with tool-first thinking — they install MLflow before they know what artifacts they need to track, or set up Kubernetes before they know their serving latency requirements. Map the lifecycle first, then choose tools that fit it.
- What you learn
- Training versus inference.
- Artifacts and owners.
- Git, Linux and identity.
- Failure points in the lifecycle.
- What you should build
- Take one ML model — yours or an open-source one — and map its full lifecycle: what artifacts are produced during training, what gets deployed for inference, who owns each stage, and where the top 3 failure points are. Write this as a one-page document.
- Ready when
- You can draw the lifecycle of one ML model from training to serving to monitoring, name the artifacts at each stage, identify who owns each transition, and point to where the model would fail if data changed, code changed or infrastructure changed.
- Common mistake
- Jumping straight into tools — installing Airflow, MLflow or Kubernetes — before you understand what you are orchestrating, tracking or deploying. The tool is not the lifecycle.
- Acceptance checks
- Map one model's training, release and operational dependencies.
- Related resources
- Google Cloud MLOps architecture — CI/CD and continuous training
- Docker getting started — Packaging basics
Stage 2: Data contracts, versioning and lineage
Your model is only as trustworthy as the data it was trained on. If the data changes and you cannot reproduce the exact dataset that produced your production model, you cannot debug quality regressions, audit for compliance, or roll back to a known-good state.
In production ML, data drifts. Schemas change. Columns get renamed. New categories appear. Without versioned data and contracts, a training run that produced a great model last month may produce a broken one today — and you will not know why. Data contracts catch these changes before they reach training.
- What you learn
- Snapshots and schemas.
- Data quality checks.
- Access and point-in-time correctness.
- Lineage tracking.
- What you should build
- Set up data versioning with DVC or equivalent for your training dataset. Write a schema contract that validates column names, types and ranges before training starts. Then deliberately break the schema (rename a column, change a type) and verify that your pipeline rejects the incompatible data instead of silently training on wrong input.
- Ready when
- Your training pipeline rejects incompatible data with a clear error message naming the violated contract. You can reproduce the exact dataset that produced your last promoted model from its version hash.
- Common mistake
- Versioning code with Git but leaving data unversioned. The most common production ML failure is a model that degraded because the training data changed silently — a column was dropped, a join changed or a source system was updated.
- Acceptance checks
- Reject incompatible data and reproduce the selected dataset.
- Related resources
- DVC getting started — Data versioning
- scikit-learn common pitfalls — Data quality and leakage
Stage 3: Reproducible training and experiment tracking
A reproducible training run means: given the same code, data, configuration and environment, you can produce the same model artifacts. This requires tracking not just the final model file, but every input and decision that produced it — including hyperparameters, library versions and random seeds.
When a production model degrades, the first question is 'what changed?' Without experiment tracking, you cannot answer that question. You do not know which data version, which code commit or which hyperparameters produced the current model. Reproducibility is not about academic purity — it is about being able to debug and roll back in production.
- What you learn
- Code, data and configuration versions.
- Metrics and artifacts.
- Experiment tracking.
- Seeds and determinism.
- What you should build
- Set up MLflow or equivalent to track: code commit hash, data version hash, hyperparameters, environment (library versions), metrics and model artifacts for every training run. Then reproduce your promoted model from scratch using only the recorded inputs — no manual steps, no 'I think I used learning rate 0.01.'
- Ready when
- You can reproduce your promoted model from recorded inputs alone — the same code commit, data version, config and environment — and get the same metrics within expected variance. Someone else on your team could do the same from your experiment records.
- Common mistake
- Tracking only the final model file and metrics, not the full input chain. When the model degrades in production, you have no way to trace back what changed.
- Acceptance checks
- Reproduce the promoted model from recorded inputs.
- Related resources
- MLflow ML documentation — Experiment tracking
- DVC getting started — Pipeline reproducibility
Stage 4: Package models and preprocessing
A model artifact without its preprocessing pipeline is a time bomb. If you fit a scaler on training data and then apply different preprocessing at inference time, your model will silently produce wrong predictions. Packaging means the model and its preprocessing travel together as one deployable unit.
Train/serve skew — when training preprocessing differs from inference preprocessing — is one of the most common and hardest-to-debug production ML failures. The model looks fine in evaluation but serves wrong predictions in production because the preprocessing was not packaged with it. A container that bundles the model, preprocessing and dependencies eliminates this class of failure.
- What you learn
- Dependencies and images.
- CPU/GPU compatibility.
- Readiness and health checks.
- Preprocessing pipeline packaging.
- What you should build
- Package your model and its full preprocessing pipeline (scaler, encoder, feature engineering) into a single Docker container. Include a health check endpoint that verifies the model loads and can make a prediction. Then load the container in a clean environment — no training dependencies installed — and verify it produces the same predictions as your local model.
- Ready when
- Your container starts cleanly in a fresh environment, the health check passes only when the model is loaded and ready to serve, and predictions match your local model within expected precision. No training-only dependencies are needed at inference time.
- Common mistake
- Saving just the model weights without the preprocessing pipeline. The model works in the notebook where the scaler was fitted, but produces garbage in production where the scaler is missing or was fitted on different data.
- Acceptance checks
- Load the saved preprocessing/model pipeline in a clean environment.
- Related resources
- Docker getting started — Container packaging
- scikit-learn common pitfalls — Training/serving consistency
Stage 5: Choose a serving pattern
How you serve a model determines its latency, throughput, cost and operational complexity. Batch inference processes records in bulk on a schedule. Synchronous inference responds to individual requests in real time. Asynchronous inference handles requests that take seconds or minutes. Event-driven inference reacts to data changes. The right choice depends on your latency budget, traffic pattern and cost constraints.
Choosing the wrong serving pattern means either overpaying for infrastructure you do not need (synchronous serving for a batch job) or failing your latency SLO (batch serving for a real-time use case). In production, the serving pattern also determines your failure modes — a synchronous endpoint that takes 30 seconds will queue requests and cascade failures under load.
- What you learn
- Batch inference.
- Online inference.
- Asynchronous inference.
- Event-driven inference.
- What you should build
- Implement two serving patterns for the same model — synchronous (FastAPI endpoint) and batch (scheduled job that processes a CSV). Benchmark both: measure p50/p95 latency, throughput, cost per 1000 predictions and resource usage. Write a one-paragraph justification for which pattern fits a real-time fraud detection use case vs a nightly batch scoring use case.
- Ready when
- You can benchmark two serving patterns against their request contracts and explain which fits which use case based on latency, throughput and cost — not just 'synchronous is for real-time.'
- Common mistake
- Defaulting to synchronous serving for everything because it is the simplest to build, then discovering at 3am that your 5-second inference time queues 1000 requests and takes down the service.
- Acceptance checks
- Benchmark the selected pattern against its request contract.
- Related resources
- FastAPI tutorial — API contracts
- Google Cloud MLOps architecture — Serving patterns
Stage 6: Orchestrate pipelines and infrastructure
Pipeline orchestration coordinates the steps of your ML workflow — data validation, training, evaluation, packaging and deployment — with dependencies, retries and scheduling. Cluster orchestration (Kubernetes) manages the infrastructure those steps run on. They are different layers, and confusing them leads to over-engineered or under-engineered systems.
SCAI instructors see two opposite mistakes: teams that build no orchestration and run pipelines manually (forgetting steps, running on stale data), and teams that over-engineer with Kubernetes before they need it (spending weeks on cluster setup for a model that trains in 10 minutes on a laptop). The right orchestration layer depends on your pipeline complexity, team size and infrastructure budget.
- What you learn
- Pipeline dependencies and retries.
- Backfills and scheduling.
- Workflow orchestration.
- Cluster orchestration as separate layer.
- What you should build
- Set up a pipeline orchestrator (Airflow, Prefect or equivalent) for your training workflow: data validation → training → evaluation → packaging. Configure it with dependencies (evaluation runs only if training succeeds), retries (one retry on transient failure) and a backfill capability (rerun for a past date). Then deliberately fail one step and verify the pipeline stops, the completed steps are not rerun, and the failure is visible.
- Ready when
- You can rerun a failed pipeline step safely — completed steps are not repeated, the failure is logged and visible, and the pipeline resumes from the failure point. You can justify whether your workload needs Kubernetes or runs fine on a single machine.
- Common mistake
- Conflating workflow orchestration with cluster orchestration. Airflow coordinates jobs and dependencies; Kubernetes manages containers and infrastructure. You can use Airflow without Kubernetes, and Kubernetes without Airflow. They solve different problems.
- Acceptance checks
- Rerun a failed pipeline step safely.
- Related resources
- Google Cloud MLOps architecture — Pipeline orchestration
- Kubernetes overview — Optional cluster orchestration
Stage 7: Test, release and roll back models
Releasing a model to production is not a git push — it is a decision that affects real users. A release pipeline needs gates: code tests pass, data validation passes, model metrics meet thresholds, and the model is registered with a version. Then it needs staged rollout (canary or blue-green) and rollback capability for when the new model is worse than the old one.
A bad model in production is worse than no model — it makes wrong predictions at scale, erodes user trust and may violate compliance requirements. Release gates catch degraded models before they reach users. Rollback capability means a bad release is a 2-minute fix, not a 2-hour crisis.
- What you learn
- Code, data and model gates.
- Registry promotion.
- Staged rollout.
- Rollback capability.
- What you should build
- Set up a CI/CD pipeline for your model: code tests → data validation → training → evaluation against a threshold → registry promotion. Then deploy with a canary strategy: route 10% of traffic to the new model, monitor metrics for 10 minutes, and promote to 100% only if metrics hold. Finally, deliberately introduce a regression (lower the evaluation threshold so a bad model passes) and demonstrate rollback to the previous model version.
- Ready when
- A degraded candidate is blocked by your gates before reaching production. If a bad model somehow passes gates, you can roll back to the previous model version within 5 minutes. Your release history shows which model version is serving at any point in time.
- Common mistake
- Promoting models without gates or rollback — treating model deployment like code deployment. A code bug is usually obvious; a model quality regression may be subtle and only visible in specific user segments.
- Acceptance checks
- Block a degraded candidate and restore a previous artifact.
- Related resources
- MLflow ML documentation — Model registry and promotion
- Google Cloud MLOps architecture — CI/CD for ML
Stage 8: Monitor services, data and model performance
A deployed ML model has two independent health signals: service health (is the endpoint up, fast and not erroring?) and model quality (is it still making accurate predictions?). Service monitoring catches outages. Model monitoring catches silent degradation — the service is healthy but predictions are wrong because the input data distribution shifted.
Model quality can degrade while the service stays perfectly healthy. No alerts fire. No errors appear. But predictions are wrong — maybe because a feature distribution shifted, a data source changed or the population the model was trained on no longer matches the population it is serving. Without model monitoring, you discover this when a business stakeholder asks 'why did our model start making bad decisions last month?'
- What you learn
- Service latency and errors.
- Input change detection.
- Delayed labels.
- Useful alerts.
- What you should build
- Set up monitoring for both layers: service metrics (latency p50/p95, error rate, request volume) with alerts tied to SLOs, and model metrics (input distribution drift, prediction distribution shift, delayed ground-truth comparison when labels arrive). Then inject two failures: a service outage (kill the process) and a model quality regression (shift the input distribution) — and verify that each triggers a different alert with a different response.
- Ready when
- You can diagnose a service failure (endpoint down, latency spike) separately from a model quality regression (predictions drifting, input distribution changed). Your monitoring distinguishes 'the model is unreachable' from 'the model is reachable but wrong.'
- Common mistake
- Monitoring only service metrics — CPU, memory, latency, error rate — and assuming the model is fine because the endpoint is healthy. Model quality degradation is silent and requires its own monitoring layer.
- Acceptance checks
- Distinguish a service outage from a model-quality regression.
- Related resources
- OpenTelemetry signals — Service observability
- Google SRE monitoring — Actionable monitoring
Stage 9: Retraining, governance and incident response
Drift detection tells you something changed. It does not tell you the new model will be better. Retraining should be triggered automatically when drift is detected, but promotion to production must pass the same gates as any other release: evaluation against thresholds, comparison with the current model, and operational checks. A newly trained model is a candidate, not a promotion.
Automatic retraining with automatic promotion is the most expensive mistake in MLOps. It means every data drift triggers a new model that may be worse than the current one — and if you have no evaluation gate, the worse model goes to production silently. The result is oscillating model quality that nobody can explain. Retrain automatically; promote with gates.
- What you learn
- Candidate evaluation.
- Approvals and auditability.
- Incident runbooks.
- Governance policies.
- What you should build
- Set up a retraining trigger that fires when input drift is detected. The trigger runs a new training job and registers the model as a candidate. Then add a promotion gate: the candidate must beat the current model on held-out metrics AND pass operational checks (latency, memory) before it replaces the production model. If the candidate fails the gate, the current model stays in production and the failed candidate is logged for analysis.
- Ready when
- Retraining is triggered automatically by drift, but promotion requires the candidate to pass evaluation and operational gates. A weaker candidate is blocked and the current production model continues serving. You can explain why a retraining run did or did not result in a production deployment.
- Common mistake
- Automatically promoting newly trained models without evaluation. Drift does not mean the new model is better — it means the old model may be worse. You still need to prove the new one is an improvement before replacing what works.
- Acceptance checks
- Trigger retraining without automatically promoting a weaker model.
- Related resources
- Google Cloud MLOps architecture — Governance and retraining
- Google SRE monitoring — Incident response
Stage 1: ML lifecycle and operational foundations
Before you touch any MLOps tool, understand how a model travels from a data scientist's notebook to a production service that real users depend on. This stage maps the full lifecycle — training artifacts, inference artifacts, the people who own each stage, and the failure points where things go wrong.
Engineers who skip lifecycle understanding end up with tool-first thinking — they install MLflow before they know what artifacts they need to track, or set up Kubernetes before they know their serving latency requirements. Map the lifecycle first, then choose tools that fit it.
- What you learn
- Training versus inference.
- Artifacts and owners.
- Git, Linux and identity.
- Failure points in the lifecycle.
- What you should build
- Take one ML model — yours or an open-source one — and map its full lifecycle: what artifacts are produced during training, what gets deployed for inference, who owns each stage, and where the top 3 failure points are. Write this as a one-page document.
- Ready when
- You can draw the lifecycle of one ML model from training to serving to monitoring, name the artifacts at each stage, identify who owns each transition, and point to where the model would fail if data changed, code changed or infrastructure changed.
- Common mistake
- Jumping straight into tools — installing Airflow, MLflow or Kubernetes — before you understand what you are orchestrating, tracking or deploying. The tool is not the lifecycle.
- Acceptance checks
- Map one model's training, release and operational dependencies.
- Related resources
- Google Cloud MLOps architecture — CI/CD and continuous training
- Docker getting started — Packaging basics
Stage 2: Data contracts, versioning and lineage
Your model is only as trustworthy as the data it was trained on. If the data changes and you cannot reproduce the exact dataset that produced your production model, you cannot debug quality regressions, audit for compliance, or roll back to a known-good state.
In production ML, data drifts. Schemas change. Columns get renamed. New categories appear. Without versioned data and contracts, a training run that produced a great model last month may produce a broken one today — and you will not know why. Data contracts catch these changes before they reach training.
- What you learn
- Snapshots and schemas.
- Data quality checks.
- Access and point-in-time correctness.
- Lineage tracking.
- What you should build
- Set up data versioning with DVC or equivalent for your training dataset. Write a schema contract that validates column names, types and ranges before training starts. Then deliberately break the schema (rename a column, change a type) and verify that your pipeline rejects the incompatible data instead of silently training on wrong input.
- Ready when
- Your training pipeline rejects incompatible data with a clear error message naming the violated contract. You can reproduce the exact dataset that produced your last promoted model from its version hash.
- Common mistake
- Versioning code with Git but leaving data unversioned. The most common production ML failure is a model that degraded because the training data changed silently — a column was dropped, a join changed or a source system was updated.
- Acceptance checks
- Reject incompatible data and reproduce the selected dataset.
- Related resources
- DVC getting started — Data versioning
- scikit-learn common pitfalls — Data quality and leakage
Stage 3: Reproducible training and experiment tracking
A reproducible training run means: given the same code, data, configuration and environment, you can produce the same model artifacts. This requires tracking not just the final model file, but every input and decision that produced it — including hyperparameters, library versions and random seeds.
When a production model degrades, the first question is 'what changed?' Without experiment tracking, you cannot answer that question. You do not know which data version, which code commit or which hyperparameters produced the current model. Reproducibility is not about academic purity — it is about being able to debug and roll back in production.
- What you learn
- Code, data and configuration versions.
- Metrics and artifacts.
- Experiment tracking.
- Seeds and determinism.
- What you should build
- Set up MLflow or equivalent to track: code commit hash, data version hash, hyperparameters, environment (library versions), metrics and model artifacts for every training run. Then reproduce your promoted model from scratch using only the recorded inputs — no manual steps, no 'I think I used learning rate 0.01.'
- Ready when
- You can reproduce your promoted model from recorded inputs alone — the same code commit, data version, config and environment — and get the same metrics within expected variance. Someone else on your team could do the same from your experiment records.
- Common mistake
- Tracking only the final model file and metrics, not the full input chain. When the model degrades in production, you have no way to trace back what changed.
- Acceptance checks
- Reproduce the promoted model from recorded inputs.
- Related resources
- MLflow ML documentation — Experiment tracking
- DVC getting started — Pipeline reproducibility
Stage 4: Package models and preprocessing
A model artifact without its preprocessing pipeline is a time bomb. If you fit a scaler on training data and then apply different preprocessing at inference time, your model will silently produce wrong predictions. Packaging means the model and its preprocessing travel together as one deployable unit.
Train/serve skew — when training preprocessing differs from inference preprocessing — is one of the most common and hardest-to-debug production ML failures. The model looks fine in evaluation but serves wrong predictions in production because the preprocessing was not packaged with it. A container that bundles the model, preprocessing and dependencies eliminates this class of failure.
- What you learn
- Dependencies and images.
- CPU/GPU compatibility.
- Readiness and health checks.
- Preprocessing pipeline packaging.
- What you should build
- Package your model and its full preprocessing pipeline (scaler, encoder, feature engineering) into a single Docker container. Include a health check endpoint that verifies the model loads and can make a prediction. Then load the container in a clean environment — no training dependencies installed — and verify it produces the same predictions as your local model.
- Ready when
- Your container starts cleanly in a fresh environment, the health check passes only when the model is loaded and ready to serve, and predictions match your local model within expected precision. No training-only dependencies are needed at inference time.
- Common mistake
- Saving just the model weights without the preprocessing pipeline. The model works in the notebook where the scaler was fitted, but produces garbage in production where the scaler is missing or was fitted on different data.
- Acceptance checks
- Load the saved preprocessing/model pipeline in a clean environment.
- Related resources
- Docker getting started — Container packaging
- scikit-learn common pitfalls — Training/serving consistency
Stage 5: Choose a serving pattern
How you serve a model determines its latency, throughput, cost and operational complexity. Batch inference processes records in bulk on a schedule. Synchronous inference responds to individual requests in real time. Asynchronous inference handles requests that take seconds or minutes. Event-driven inference reacts to data changes. The right choice depends on your latency budget, traffic pattern and cost constraints.
Choosing the wrong serving pattern means either overpaying for infrastructure you do not need (synchronous serving for a batch job) or failing your latency SLO (batch serving for a real-time use case). In production, the serving pattern also determines your failure modes — a synchronous endpoint that takes 30 seconds will queue requests and cascade failures under load.
- What you learn
- Batch inference.
- Online inference.
- Asynchronous inference.
- Event-driven inference.
- What you should build
- Implement two serving patterns for the same model — synchronous (FastAPI endpoint) and batch (scheduled job that processes a CSV). Benchmark both: measure p50/p95 latency, throughput, cost per 1000 predictions and resource usage. Write a one-paragraph justification for which pattern fits a real-time fraud detection use case vs a nightly batch scoring use case.
- Ready when
- You can benchmark two serving patterns against their request contracts and explain which fits which use case based on latency, throughput and cost — not just 'synchronous is for real-time.'
- Common mistake
- Defaulting to synchronous serving for everything because it is the simplest to build, then discovering at 3am that your 5-second inference time queues 1000 requests and takes down the service.
- Acceptance checks
- Benchmark the selected pattern against its request contract.
- Related resources
- FastAPI tutorial — API contracts
- Google Cloud MLOps architecture — Serving patterns
Stage 6: Orchestrate pipelines and infrastructure
Pipeline orchestration coordinates the steps of your ML workflow — data validation, training, evaluation, packaging and deployment — with dependencies, retries and scheduling. Cluster orchestration (Kubernetes) manages the infrastructure those steps run on. They are different layers, and confusing them leads to over-engineered or under-engineered systems.
SCAI instructors see two opposite mistakes: teams that build no orchestration and run pipelines manually (forgetting steps, running on stale data), and teams that over-engineer with Kubernetes before they need it (spending weeks on cluster setup for a model that trains in 10 minutes on a laptop). The right orchestration layer depends on your pipeline complexity, team size and infrastructure budget.
- What you learn
- Pipeline dependencies and retries.
- Backfills and scheduling.
- Workflow orchestration.
- Cluster orchestration as separate layer.
- What you should build
- Set up a pipeline orchestrator (Airflow, Prefect or equivalent) for your training workflow: data validation → training → evaluation → packaging. Configure it with dependencies (evaluation runs only if training succeeds), retries (one retry on transient failure) and a backfill capability (rerun for a past date). Then deliberately fail one step and verify the pipeline stops, the completed steps are not rerun, and the failure is visible.
- Ready when
- You can rerun a failed pipeline step safely — completed steps are not repeated, the failure is logged and visible, and the pipeline resumes from the failure point. You can justify whether your workload needs Kubernetes or runs fine on a single machine.
- Common mistake
- Conflating workflow orchestration with cluster orchestration. Airflow coordinates jobs and dependencies; Kubernetes manages containers and infrastructure. You can use Airflow without Kubernetes, and Kubernetes without Airflow. They solve different problems.
- Acceptance checks
- Rerun a failed pipeline step safely.
- Related resources
- Google Cloud MLOps architecture — Pipeline orchestration
- Kubernetes overview — Optional cluster orchestration
Stage 7: Test, release and roll back models
Releasing a model to production is not a git push — it is a decision that affects real users. A release pipeline needs gates: code tests pass, data validation passes, model metrics meet thresholds, and the model is registered with a version. Then it needs staged rollout (canary or blue-green) and rollback capability for when the new model is worse than the old one.
A bad model in production is worse than no model — it makes wrong predictions at scale, erodes user trust and may violate compliance requirements. Release gates catch degraded models before they reach users. Rollback capability means a bad release is a 2-minute fix, not a 2-hour crisis.
- What you learn
- Code, data and model gates.
- Registry promotion.
- Staged rollout.
- Rollback capability.
- What you should build
- Set up a CI/CD pipeline for your model: code tests → data validation → training → evaluation against a threshold → registry promotion. Then deploy with a canary strategy: route 10% of traffic to the new model, monitor metrics for 10 minutes, and promote to 100% only if metrics hold. Finally, deliberately introduce a regression (lower the evaluation threshold so a bad model passes) and demonstrate rollback to the previous model version.
- Ready when
- A degraded candidate is blocked by your gates before reaching production. If a bad model somehow passes gates, you can roll back to the previous model version within 5 minutes. Your release history shows which model version is serving at any point in time.
- Common mistake
- Promoting models without gates or rollback — treating model deployment like code deployment. A code bug is usually obvious; a model quality regression may be subtle and only visible in specific user segments.
- Acceptance checks
- Block a degraded candidate and restore a previous artifact.
- Related resources
- MLflow ML documentation — Model registry and promotion
- Google Cloud MLOps architecture — CI/CD for ML
Stage 8: Monitor services, data and model performance
A deployed ML model has two independent health signals: service health (is the endpoint up, fast and not erroring?) and model quality (is it still making accurate predictions?). Service monitoring catches outages. Model monitoring catches silent degradation — the service is healthy but predictions are wrong because the input data distribution shifted.
Model quality can degrade while the service stays perfectly healthy. No alerts fire. No errors appear. But predictions are wrong — maybe because a feature distribution shifted, a data source changed or the population the model was trained on no longer matches the population it is serving. Without model monitoring, you discover this when a business stakeholder asks 'why did our model start making bad decisions last month?'
- What you learn
- Service latency and errors.
- Input change detection.
- Delayed labels.
- Useful alerts.
- What you should build
- Set up monitoring for both layers: service metrics (latency p50/p95, error rate, request volume) with alerts tied to SLOs, and model metrics (input distribution drift, prediction distribution shift, delayed ground-truth comparison when labels arrive). Then inject two failures: a service outage (kill the process) and a model quality regression (shift the input distribution) — and verify that each triggers a different alert with a different response.
- Ready when
- You can diagnose a service failure (endpoint down, latency spike) separately from a model quality regression (predictions drifting, input distribution changed). Your monitoring distinguishes 'the model is unreachable' from 'the model is reachable but wrong.'
- Common mistake
- Monitoring only service metrics — CPU, memory, latency, error rate — and assuming the model is fine because the endpoint is healthy. Model quality degradation is silent and requires its own monitoring layer.
- Acceptance checks
- Distinguish a service outage from a model-quality regression.
- Related resources
- OpenTelemetry signals — Service observability
- Google SRE monitoring — Actionable monitoring
Stage 9: Retraining, governance and incident response
Drift detection tells you something changed. It does not tell you the new model will be better. Retraining should be triggered automatically when drift is detected, but promotion to production must pass the same gates as any other release: evaluation against thresholds, comparison with the current model, and operational checks. A newly trained model is a candidate, not a promotion.
Automatic retraining with automatic promotion is the most expensive mistake in MLOps. It means every data drift triggers a new model that may be worse than the current one — and if you have no evaluation gate, the worse model goes to production silently. The result is oscillating model quality that nobody can explain. Retrain automatically; promote with gates.
- What you learn
- Candidate evaluation.
- Approvals and auditability.
- Incident runbooks.
- Governance policies.
- What you should build
- Set up a retraining trigger that fires when input drift is detected. The trigger runs a new training job and registers the model as a candidate. Then add a promotion gate: the candidate must beat the current model on held-out metrics AND pass operational checks (latency, memory) before it replaces the production model. If the candidate fails the gate, the current model stays in production and the failed candidate is logged for analysis.
- Ready when
- Retraining is triggered automatically by drift, but promotion requires the candidate to pass evaluation and operational gates. A weaker candidate is blocked and the current production model continues serving. You can explain why a retraining run did or did not result in a production deployment.
- Common mistake
- Automatically promoting newly trained models without evaluation. Drift does not mean the new model is better — it means the old model may be worse. You still need to prove the new one is an improvement before replacing what works.
- Acceptance checks
- Trigger retraining without automatically promoting a weaker model.
- Related resources
- Google Cloud MLOps architecture — Governance and retraining
- Google SRE monitoring — Incident response
From roadmap to production
Build production MLOps systems with instructor feedback
You have the framework. The View the MLOps syllabus adds what self-study cannot: live instruction, instructor-reviewed labs, production deployment drills and a capstone that proves you can ship and operate — not just understand.
Fees, schedules and enrolment details are on the course page. No placement, salary or outcome is guaranteed.
Capstone
Build a reproducible ML release pipeline end to end
A column is renamed in the source — your schema validation rejects it before training runs on wrong data.
A retrained model scores worse than the current one — your evaluation gate blocks it from production.
The serving endpoint goes down — your monitoring detects it, alerts fire and rollback restores the previous version.
What you deliver
A reproducible training pipeline, a containerized serving endpoint, a CI/CD pipeline with evaluation gates, a monitoring dashboard, and a runbook covering the three failure scenarios. This is what an MLOps interviewer wants to see — not "I installed MLflow," but "I built a pipeline that catches bad models before they reach users and recovers when things break."
Take MLOps further
Go from understanding the framework to building production systems with feedback
| Stage | This roadmap (free) | MLOps course adds |
|---|---|---|
| Data + Training | Self-guided DVC + MLflow setup | Guided pipeline labs with instructor-reviewed experiment tracking |
| Packaging + Serving | Build Docker + FastAPI yourself | AWS SageMaker deployment with instructor feedback on your container |
| CI/CD + Monitoring | Set up gates and dashboards independently | Deployment drills with failure injection and rollback practice |
| Capstone | No feedback on your work | Reviewed capstone with instructor feedback on your pipeline and runbook |
You have the nine-stage framework — data contracts, reproducible training, packaging, serving, orchestration, CI/CD, monitoring and retraining. SCAI's 5-month live MLOps course helps you build it for real: data pipelines, distributed training, model release, drift detection and an AWS SageMaker deployment with instructor feedback at every step.
Ready to build production ML systems with instructor feedback?
Explore the MLOps courseWhat to read next
What to read next
ML Engineer Roadmap
Feature engineering, validation strategies and serving contracts — the model development side that MLOps operationalizes
LLMOps Roadmap
Prompt versioning, LLM evaluation, token cost control and serving — operating LLM applications in production
AIOps Roadmap
Telemetry, anomaly detection and AI-assisted incident investigation — detecting the incidents MLOps prevents
Related learning
- Continue to the ML Engineer roadmapFor the model-development foundation MLOps operationalizes.
- Continue to the AIOps roadmapTo broaden into production AI operations across LLMs and agents.
- Continue to the LLMOps roadmapTo specialize in LLM application operations.
- Compare MLOps and LLMOps and AIOpsWhere each operations track starts and ends.
- Compare MLOps and DevOpsWhat changes when data and models enter the pipeline.
- Compare the MLOps and AIOps coursesCourse-level decision between ML operations and broad AI operations.
FAQ
MLOps Roadmap — Frequently Asked Questions
Direct answers for engineers and hiring leads on building, transitioning into and operating ML systems in production.
I am a data scientist — what do I need to learn to move into MLOps?
You are roughly 40% of the way there — you can build the model. The 60% you are missing is everything that happens after model.fit(). Notebooks do not survive contact with production: they run on your laptop, depend on your local Python, and nobody can reproduce the run that produced your "best" model. Learn four things in this order:
1. Docker — ship the model + preprocessing together so it runs identically on your laptop, CI and production.
2. MLflow — log every run's params, metrics and artefact so any result is reproducible and traceable.
3. CI/CD (GitHub Actions) — automate training, tests and deployment so a merge does not silently change behaviour.
4. Monitoring — track drift and prediction quality after release, because a model that scored 0.93 at training can be 0.71 in production two months later and you will not know unless you measure.
You do not need to become a Kubernetes admin. This roadmap is built around exactly that 60%.
I am a DevOps or SRE engineer — can I move into MLOps, and what do I need to add?
Yes — you are roughly 60% of the way there. CI/CD, containers, IaC, monitoring, rollback, alerting: already in your toolkit. The 40% you are missing is ML literacy, and it is the part that will quietly break your pipelines if you skip it.
The three things that trip up DevOps engineers:
1. A model is not deterministic code. The same build can produce different behaviour because the data changed. Your standard "green build = safe to deploy" mental model does not hold.
2. You need to read a confusion matrix, understand precision vs. recall, and know what data leakage looks like — not to build models, but to reason about why a deployed model degraded.
3. Evaluation gates are not unit tests. A candidate model must beat the current production model on held-out data, not just pass a correctness check.
Add: data validation (schema + distribution checks), MLflow for experiment tracking, and model evaluation gates. The MLOps layer on top of your existing skills is thin — but it is the thin layer that makes or breaks production ML.
I am a software engineer with no ML background — can I learn MLOps directly?
Yes, but not by skipping ML. You do not need to design neural networks — you need enough ML literacy to treat a model as a deployable artefact that behaves differently from code.
Spend 2–3 weeks on five concepts: (1) training vs. inference (a model learns weights, then uses them — it is not "run the function"); (2) metrics — accuracy, precision, recall, AUC, and why accuracy alone lies on imbalanced data; (3) overfitting — a model that memorises training data and fails on new data; (4) data leakage — when test data sneaks into training and inflates your score; (5) why a model that scored 0.95 in the notebook can score 0.70 in production (distribution shift).
Then this roadmap is directly walkable. The MLOps skills themselves — pipelines, serving, monitoring — are closer to DevOps than to ML. Skip the literacy step and you will build pipelines that look correct but miss every real failure mode.
What does an MLOps engineer actually do day to day?
Closer to platform engineering for ML than to model research. A typical week breaks into four buckets:
Pipeline work (~40%) — maintain data validation, training, packaging and deployment. A scientist hands you a notebook; you turn it into a reproducible, tested pipeline that runs on schedule and produces a registered model.
Incidents (~25%) — a drift alert fires at 2am: investigate, decide rollback vs. retrain, write the runbook so the next on-call does not debug from scratch.
Collaboration (~20%) — review data scientists' code for leakage, help them log experiments properly, ship their model behind a versioned API.
Reliability improvements (~15%) — add evaluation gates, improve monitoring coverage, tighten the rollback path.
If you want to build novel models, that is an ML Engineer role. MLOps is the engineering that makes those models reliably reach and stay in production.
How is MLOps different from DevOps — and from ML Engineering?
Three roles, one artefact, different halves of its life:
DevOps ships deterministic code. Same build → same behaviour. CI/CD, IaC, rollback. The artefact does not change after deploy.
ML Engineering builds the model — features, validation strategy, serving contracts, model architecture. The artefact is the model itself.
MLOps ships a model that depends on data, and the data drifts. So MLOps adds three things DevOps does not need: data validation (schema + distribution checks before training), evaluation gates (candidate must beat the current model on held-out data, not just pass tests), and model monitoring (track prediction quality, not just service health).
In small teams one person does all three. In larger teams they split: ML Engineers build models, MLOps engineers make them reach and stay in production, DevOps owns the underlying platform.
What tools should I learn first in MLOps?
Five, in this order — each maps to one lifecycle stage:
1. DVC — version your datasets and preprocessing the way Git versions code. Without this, "retrain on the latest data" is not reproducible.
2. MLflow — log params, metrics and the model artefact for every run, and keep a model registry. This is how you prove which run was best.
3. Docker — package the model + preprocessing together so it runs identically in dev, CI and prod.
4. FastAPI — serve the containerised model behind a REST endpoint with a /health check.
5. GitHub Actions — wire training, tests, evaluation gates and deployment into a pipeline that runs on every merge.
Add Prometheus + Grafana for monitoring once something is in production. Do not start with Kubernetes — it is the most over-prescribed tool in MLOps and will cost you weeks before you need it. Airflow and feature stores come later, only when a real problem demands them.
Is Kubernetes required for MLOps?
No. Containers — yes, because the model and its preprocessing must ship together. Kubernetes — no, unless you have a specific reason.
Most production ML systems run on a managed service (AWS SageMaker, GCP Vertex AI, Azure ML) or even a single beefy host. Reach for Kubernetes only when you hit one of: distributed training that one machine cannot handle, multi-tenant serving with autoscaling, or shared GPU infrastructure across teams.
Rule of thumb: if your model trains in under an hour on a single machine and serves under 100 QPS, Kubernetes is overhead you do not need. Adding it before you need it will slow your delivery more than it helps — you will spend more time on cluster YAML than on the model.
Do I need a feature store?
Only when you have a specific problem it solves. A feature store adds real infrastructure cost — do not adopt one on FOMO.
You need one when:
• Multiple teams share feature definitions and you are tired of three teams recomputing the same feature three different ways.
• You serve in real time and need online/offline consistency (the feature value at serving time must match what was used in training).
• You need point-in-time correctness for training (avoiding future leakage into historical features).
For a single model or a small team, versioned datasets (DVC) plus a preprocessing pipeline packaged inside the model container are enough. Adopt a feature store when the pain of not having one is concrete, not because a conference talk made it sound modern.
What is data drift and how do I detect it?
Data drift = the input distribution changed between training and production. The world moved; your model did not. A churn model trained on 2024 customer behaviour starts mispredicting in 2026 because usage patterns shifted — the code is fine, the data is not.
Detect it three ways:
1. Feature-level statistical tests — KS test or Population Stability Index (PSI) on each feature; flag when a feature's distribution shifts beyond a threshold.
2. Prediction distribution shift — if the model's output distribution changes (e.g. suddenly 30% more "churn" predictions), inputs likely shifted.
3. Embedding distance — for text/image inputs, track distance between production embeddings and a reference set.
Tools: Evidently and WhyLabs compute these automatically and generate reports. Drift is a signal to investigate and retrain a candidate — not to auto-promote a new model.
Should drift automatically trigger a new production model?
No. This is the single most expensive mistake in MLOps.
Drift means the data changed. It does not mean a model retrained on the new data will be better — it might overfit to a temporary shift, or the new data might be noisy.
Split it:
• Auto-trigger candidate retraining — yes. Run the pipeline, produce a candidate, log it.
• Auto-promote to production — no. The candidate must pass evaluation gates first: held-out metrics against a fixed test set, head-to-head comparison with the current production model, and operational checks (latency, memory, footprint).
Auto-retrain + auto-promote causes silent quality oscillation: the model flips between versions every few hours, nobody can explain why, and the team stops trusting the pipeline. Gate the promotion; automate the candidate.
What is the difference between service monitoring and model monitoring?
Two layers, two different questions. You need both.
Service monitoring asks: "is the endpoint alive and fast?" — CPU, memory, latency, error rate, pod restarts. This is standard DevOps monitoring (Prometheus + Grafana). It catches a crashed pod or a memory leak.
Model monitoring asks: "are the predictions still correct?" — input drift, prediction distribution shift, and delayed ground-truth comparison once real labels arrive. This is MLOps-specific. It catches a model that is happily serving 200ms responses that are silently wrong because the input data shifted last month.
The trap: a model can be perfectly healthy as a service and wrong as a predictor. Service monitoring alone will never catch silent model degradation. Model monitoring alone will never catch a crashed pod. Run both; wire both to alerting.
As an MLOps lead, what do I look for when hiring an MLOps engineer?
Three signals, in priority order. Tool names are teachable in a week; judgement about ML in production takes months — hire for the judgement.
1. A built pipeline, not installed tools. "I set up MLflow" tells me nothing. "I built a pipeline that catches a bad model before it reaches users and rolls back when serving breaks" tells me everything. I ask: "walk me through what happens when a data contract breaks at 2am." If they can narrate detect → alert → rollback → investigate → fix without hand-waving, they have run production ML.
2. Failure reasoning. I ask about three scenarios: a candidate model regresses on the gate, the serving endpoint OOMs, ground-truth labels are delayed two weeks. Can they reason through each, or do they freeze?
3. ML literacy. Can they read a confusion matrix, explain precision vs. recall trade-off, and say why a model degrades in production? They do not need to build models — they need to reason about the ones they ship.
How long does it take to learn MLOps, and what should I learn after this roadmap?
Timeline: If you know Python, Git and basic ML — about 10–12 weeks of focused part-time work (8–10 hrs/week) to walk this roadmap end to end and ship a capstone pipeline. Coming from DevOps/SRE without ML — add 4–6 weeks for ML literacy first. Coming with no Python — add 6–8 weeks for Python and Git basics.
What to build: one capstone — a reproducible training pipeline, a containerised serving endpoint, a CI/CD pipeline with evaluation gates, a monitoring dashboard, and a runbook for three failure scenarios. That artefact is worth more on a CV than any certificate.
Where to go next:
• Deeper model development — feature engineering, validation, serving contracts → ML Engineer roadmap
• LLM-specific operations — prompt versioning, LLM evaluation, token cost → LLMOps roadmap
• AI-assisted incident detection → AIOps roadmap
• Structured practice with instructor feedback and a reviewed capstone → SCAI's MLOps course covers the full lifecycle live.