ROADMAP

MLOps Roadmap

From data contracts to production deployment, monitoring and rollback — build a repeatable ML delivery process.

MLOps is not a tool — it is the process that turns a notebook experiment into a production system you can deploy, monitor and recover from. This roadmap covers the full lifecycle: data versioning and contracts, reproducible training with experiment tracking, model packaging, serving patterns, pipeline orchestration, CI/CD with rollback, monitoring for both service health and model quality, and governance with retraining. Each stage has a build task, acceptance checks and a measurable outcome. By the end, you can take a model from experiment to production and keep it running.

For:ML engineers, data scientists, backend engineers and DevOps moving into AI infrastructure.

What is the right MLOps roadmap?

The right MLOps roadmap does not start with tools — it starts with a question: can you take a model from a data scientist's notebook to a production service, keep it running and recover when it breaks? That means versioning your data so training is reproducible, packaging the model with its preprocessing so inference matches training, choosing a serving pattern that fits your latency budget, automating releases with gates that block bad models, monitoring both service health and prediction quality, and treating retraining as a candidate — not an automatic promotion. If you cannot reproduce your promoted model from recorded inputs, you do not have MLOps — you have a deployment script.

Written byAshutosh· AI InstructorVerified byVivek· AIOps and Generative AI InstructorPublishedUpdated

Sources and methodology · This roadmap is reviewed when production practices, tools or platform patterns materially change.

Stages

9

Last reviewed

16 September 2026

Stage 1: ML lifecycle and operational foundations

Before you touch any MLOps tool, understand how a model travels from a data scientist's notebook to a production service that real users depend on. This stage maps the full lifecycle — training artifacts, inference artifacts, the people who own each stage, and the failure points where things go wrong.

Engineers who skip lifecycle understanding end up with tool-first thinking — they install MLflow before they know what artifacts they need to track, or set up Kubernetes before they know their serving latency requirements. Map the lifecycle first, then choose tools that fit it.

What you learn
  • Training versus inference.
  • Artifacts and owners.
  • Git, Linux and identity.
  • Failure points in the lifecycle.
What you should build
Take one ML model — yours or an open-source one — and map its full lifecycle: what artifacts are produced during training, what gets deployed for inference, who owns each stage, and where the top 3 failure points are. Write this as a one-page document.
Ready when
You can draw the lifecycle of one ML model from training to serving to monitoring, name the artifacts at each stage, identify who owns each transition, and point to where the model would fail if data changed, code changed or infrastructure changed.
Common mistake
Jumping straight into tools — installing Airflow, MLflow or Kubernetes — before you understand what you are orchestrating, tracking or deploying. The tool is not the lifecycle.
Acceptance checks
  • Map one model's training, release and operational dependencies.
Related resources

From roadmap to production

Build production MLOps systems with instructor feedback

You have the framework. The View the MLOps syllabus adds what self-study cannot: live instruction, instructor-reviewed labs, production deployment drills and a capstone that proves you can ship and operate — not just understand.

Build the core project from this roadmap with instructor review
Debug production failure modes hands-on with guided feedback
Produce a reviewed portfolio artifact by the end of the track

Fees, schedules and enrolment details are on the course page. No placement, salary or outcome is guaranteed.

Capstone

Build a reproducible ML release pipeline end to end

DataDVC versioning + schema contracts
TrainMLflow tracking + reproducible runs
PackageDocker container + preprocessing
ServeFastAPI endpoint + health check
ReleaseCI/CD gates + canary rollout
MonitorService health + model quality
Broken data contract

A column is renamed in the source — your schema validation rejects it before training runs on wrong data.

Degraded candidate model

A retrained model scores worse than the current one — your evaluation gate blocks it from production.

Service outage

The serving endpoint goes down — your monitoring detects it, alerts fire and rollback restores the previous version.

What you deliver

A reproducible training pipeline, a containerized serving endpoint, a CI/CD pipeline with evaluation gates, a monitoring dashboard, and a runbook covering the three failure scenarios. This is what an MLOps interviewer wants to see — not "I installed MLflow," but "I built a pipeline that catches bad models before they reach users and recovers when things break."

Take MLOps further

Go from understanding the framework to building production systems with feedback

StageThis roadmap (free)MLOps course adds
Data + TrainingSelf-guided DVC + MLflow setupGuided pipeline labs with instructor-reviewed experiment tracking
Packaging + ServingBuild Docker + FastAPI yourselfAWS SageMaker deployment with instructor feedback on your container
CI/CD + MonitoringSet up gates and dashboards independentlyDeployment drills with failure injection and rollback practice
CapstoneNo feedback on your workReviewed capstone with instructor feedback on your pipeline and runbook

You have the nine-stage framework — data contracts, reproducible training, packaging, serving, orchestration, CI/CD, monitoring and retraining. SCAI's 5-month live MLOps course helps you build it for real: data pipelines, distributed training, model release, drift detection and an AWS SageMaker deployment with instructor feedback at every step.

Ready to build production ML systems with instructor feedback?

Explore the MLOps course

FAQ

MLOps Roadmap — Frequently Asked Questions

Direct answers for engineers and hiring leads on building, transitioning into and operating ML systems in production.

I am a data scientist — what do I need to learn to move into MLOps?

You are roughly 40% of the way there — you can build the model. The 60% you are missing is everything that happens after model.fit(). Notebooks do not survive contact with production: they run on your laptop, depend on your local Python, and nobody can reproduce the run that produced your "best" model. Learn four things in this order:

1. Docker — ship the model + preprocessing together so it runs identically on your laptop, CI and production.
2. MLflow — log every run's params, metrics and artefact so any result is reproducible and traceable.
3. CI/CD (GitHub Actions) — automate training, tests and deployment so a merge does not silently change behaviour.
4. Monitoring — track drift and prediction quality after release, because a model that scored 0.93 at training can be 0.71 in production two months later and you will not know unless you measure.

You do not need to become a Kubernetes admin. This roadmap is built around exactly that 60%.

I am a DevOps or SRE engineer — can I move into MLOps, and what do I need to add?

Yes — you are roughly 60% of the way there. CI/CD, containers, IaC, monitoring, rollback, alerting: already in your toolkit. The 40% you are missing is ML literacy, and it is the part that will quietly break your pipelines if you skip it.

The three things that trip up DevOps engineers:
1. A model is not deterministic code. The same build can produce different behaviour because the data changed. Your standard "green build = safe to deploy" mental model does not hold.
2. You need to read a confusion matrix, understand precision vs. recall, and know what data leakage looks like — not to build models, but to reason about why a deployed model degraded.
3. Evaluation gates are not unit tests. A candidate model must beat the current production model on held-out data, not just pass a correctness check.

Add: data validation (schema + distribution checks), MLflow for experiment tracking, and model evaluation gates. The MLOps layer on top of your existing skills is thin — but it is the thin layer that makes or breaks production ML.

I am a software engineer with no ML background — can I learn MLOps directly?

Yes, but not by skipping ML. You do not need to design neural networks — you need enough ML literacy to treat a model as a deployable artefact that behaves differently from code.

Spend 2–3 weeks on five concepts: (1) training vs. inference (a model learns weights, then uses them — it is not "run the function"); (2) metrics — accuracy, precision, recall, AUC, and why accuracy alone lies on imbalanced data; (3) overfitting — a model that memorises training data and fails on new data; (4) data leakage — when test data sneaks into training and inflates your score; (5) why a model that scored 0.95 in the notebook can score 0.70 in production (distribution shift).

Then this roadmap is directly walkable. The MLOps skills themselves — pipelines, serving, monitoring — are closer to DevOps than to ML. Skip the literacy step and you will build pipelines that look correct but miss every real failure mode.

What does an MLOps engineer actually do day to day?

Closer to platform engineering for ML than to model research. A typical week breaks into four buckets:

Pipeline work (~40%) — maintain data validation, training, packaging and deployment. A scientist hands you a notebook; you turn it into a reproducible, tested pipeline that runs on schedule and produces a registered model.
Incidents (~25%) — a drift alert fires at 2am: investigate, decide rollback vs. retrain, write the runbook so the next on-call does not debug from scratch.
Collaboration (~20%) — review data scientists' code for leakage, help them log experiments properly, ship their model behind a versioned API.
Reliability improvements (~15%) — add evaluation gates, improve monitoring coverage, tighten the rollback path.

If you want to build novel models, that is an ML Engineer role. MLOps is the engineering that makes those models reliably reach and stay in production.

How is MLOps different from DevOps — and from ML Engineering?

Three roles, one artefact, different halves of its life:

DevOps ships deterministic code. Same build → same behaviour. CI/CD, IaC, rollback. The artefact does not change after deploy.
ML Engineering builds the model — features, validation strategy, serving contracts, model architecture. The artefact is the model itself.
MLOps ships a model that depends on data, and the data drifts. So MLOps adds three things DevOps does not need: data validation (schema + distribution checks before training), evaluation gates (candidate must beat the current model on held-out data, not just pass tests), and model monitoring (track prediction quality, not just service health).

In small teams one person does all three. In larger teams they split: ML Engineers build models, MLOps engineers make them reach and stay in production, DevOps owns the underlying platform.

What tools should I learn first in MLOps?

Five, in this order — each maps to one lifecycle stage:

1. DVC — version your datasets and preprocessing the way Git versions code. Without this, "retrain on the latest data" is not reproducible.
2. MLflow — log params, metrics and the model artefact for every run, and keep a model registry. This is how you prove which run was best.
3. Docker — package the model + preprocessing together so it runs identically in dev, CI and prod.
4. FastAPI — serve the containerised model behind a REST endpoint with a /health check.
5. GitHub Actions — wire training, tests, evaluation gates and deployment into a pipeline that runs on every merge.

Add Prometheus + Grafana for monitoring once something is in production. Do not start with Kubernetes — it is the most over-prescribed tool in MLOps and will cost you weeks before you need it. Airflow and feature stores come later, only when a real problem demands them.

Is Kubernetes required for MLOps?

No. Containers — yes, because the model and its preprocessing must ship together. Kubernetes — no, unless you have a specific reason.

Most production ML systems run on a managed service (AWS SageMaker, GCP Vertex AI, Azure ML) or even a single beefy host. Reach for Kubernetes only when you hit one of: distributed training that one machine cannot handle, multi-tenant serving with autoscaling, or shared GPU infrastructure across teams.

Rule of thumb: if your model trains in under an hour on a single machine and serves under 100 QPS, Kubernetes is overhead you do not need. Adding it before you need it will slow your delivery more than it helps — you will spend more time on cluster YAML than on the model.

Do I need a feature store?

Only when you have a specific problem it solves. A feature store adds real infrastructure cost — do not adopt one on FOMO.

You need one when:
• Multiple teams share feature definitions and you are tired of three teams recomputing the same feature three different ways.
• You serve in real time and need online/offline consistency (the feature value at serving time must match what was used in training).
• You need point-in-time correctness for training (avoiding future leakage into historical features).

For a single model or a small team, versioned datasets (DVC) plus a preprocessing pipeline packaged inside the model container are enough. Adopt a feature store when the pain of not having one is concrete, not because a conference talk made it sound modern.

What is data drift and how do I detect it?

Data drift = the input distribution changed between training and production. The world moved; your model did not. A churn model trained on 2024 customer behaviour starts mispredicting in 2026 because usage patterns shifted — the code is fine, the data is not.

Detect it three ways:
1. Feature-level statistical tests — KS test or Population Stability Index (PSI) on each feature; flag when a feature's distribution shifts beyond a threshold.
2. Prediction distribution shift — if the model's output distribution changes (e.g. suddenly 30% more "churn" predictions), inputs likely shifted.
3. Embedding distance — for text/image inputs, track distance between production embeddings and a reference set.

Tools: Evidently and WhyLabs compute these automatically and generate reports. Drift is a signal to investigate and retrain a candidate — not to auto-promote a new model.

Should drift automatically trigger a new production model?

No. This is the single most expensive mistake in MLOps.

Drift means the data changed. It does not mean a model retrained on the new data will be better — it might overfit to a temporary shift, or the new data might be noisy.

Split it:
• Auto-trigger candidate retraining — yes. Run the pipeline, produce a candidate, log it.
• Auto-promote to production — no. The candidate must pass evaluation gates first: held-out metrics against a fixed test set, head-to-head comparison with the current production model, and operational checks (latency, memory, footprint).

Auto-retrain + auto-promote causes silent quality oscillation: the model flips between versions every few hours, nobody can explain why, and the team stops trusting the pipeline. Gate the promotion; automate the candidate.

What is the difference between service monitoring and model monitoring?

Two layers, two different questions. You need both.

Service monitoring asks: "is the endpoint alive and fast?" — CPU, memory, latency, error rate, pod restarts. This is standard DevOps monitoring (Prometheus + Grafana). It catches a crashed pod or a memory leak.

Model monitoring asks: "are the predictions still correct?" — input drift, prediction distribution shift, and delayed ground-truth comparison once real labels arrive. This is MLOps-specific. It catches a model that is happily serving 200ms responses that are silently wrong because the input data shifted last month.

The trap: a model can be perfectly healthy as a service and wrong as a predictor. Service monitoring alone will never catch silent model degradation. Model monitoring alone will never catch a crashed pod. Run both; wire both to alerting.

As an MLOps lead, what do I look for when hiring an MLOps engineer?

Three signals, in priority order. Tool names are teachable in a week; judgement about ML in production takes months — hire for the judgement.

1. A built pipeline, not installed tools. "I set up MLflow" tells me nothing. "I built a pipeline that catches a bad model before it reaches users and rolls back when serving breaks" tells me everything. I ask: "walk me through what happens when a data contract breaks at 2am." If they can narrate detect → alert → rollback → investigate → fix without hand-waving, they have run production ML.

2. Failure reasoning. I ask about three scenarios: a candidate model regresses on the gate, the serving endpoint OOMs, ground-truth labels are delayed two weeks. Can they reason through each, or do they freeze?

3. ML literacy. Can they read a confusion matrix, explain precision vs. recall trade-off, and say why a model degrades in production? They do not need to build models — they need to reason about the ones they ship.

How long does it take to learn MLOps, and what should I learn after this roadmap?

Timeline: If you know Python, Git and basic ML — about 10–12 weeks of focused part-time work (8–10 hrs/week) to walk this roadmap end to end and ship a capstone pipeline. Coming from DevOps/SRE without ML — add 4–6 weeks for ML literacy first. Coming with no Python — add 6–8 weeks for Python and Git basics.

What to build: one capstone — a reproducible training pipeline, a containerised serving endpoint, a CI/CD pipeline with evaluation gates, a monitoring dashboard, and a runbook for three failure scenarios. That artefact is worth more on a CV than any certificate.

Where to go next:
• Deeper model development — feature engineering, validation, serving contracts → ML Engineer roadmap
• LLM-specific operations — prompt versioning, LLM evaluation, token cost → LLMOps roadmap
• AI-assisted incident detection → AIOps roadmap
• Structured practice with instructor feedback and a reviewed capstone → SCAI's MLOps course covers the full lifecycle live.