LLMOps Roadmap
Operate LLM applications with measurable quality, failure recovery, latency, cost and capacity limits.
A practical LLMOps roadmap for engineers operating LLM and generative AI systems in production. Learn service contracts, managed or self-hosted serving, failure recognition, recovery, latency, cost and caching, concurrency and routing, evaluation and release, retrieval operations, and security with SLOs.
What is the right roadmap for learning LLMOps?
LLMOps covers the lifecycle of LLM applications and inference services. Version the application, prompts, models and retrieval assets; evaluate changes before release; and observe production requests. Learn failure recovery, latency and cost control. Follow the managed-API branch for provider-backed applications and the self-hosted branch when you operate model-serving infrastructure.
Sources and methodology · This roadmap is reviewed when production practices, tools or platform patterns materially change.
Stages
10
Last reviewed
16 September 2026
Stage 1: Service contracts, tracing and a quality baseline
Request/output contracts, configuration, versioning, traces and task evaluation.
Without a baseline, changes cannot be evaluated.
- What you learn
- Request and output contracts.
- Configuration and versioning.
- Traces and task evaluation.
- Quality baseline.
- What you should build
- Record a baseline request's quality, latency and usage.
- Ready when
- You can record a baseline request's quality, latency and token usage.
- Common mistake
- Operating without a quality baseline or request contract.
- Acceptance checks
- Record a baseline request's quality, latency and usage.
- Related resources
- FastAPI tutorial — API contracts
- MLflow LLM documentation — Evaluation
- OpenTelemetry signals — Tracing
Stage 2: Managed APIs or self-hosted inference
Choose a branch: provider quotas and model versions, or hardware, model loading and serving.
The two branches have different operational concerns.
- What you learn
- Managed API branch.
- Self-hosted inference branch.
- Provider quotas and model versions.
- Hardware, model loading and serving.
- What you should build
- Document the chosen serving contract and resource constraints.
- Ready when
- You can document the chosen serving branch and its resource constraints.
- Common mistake
- Mixing managed and self-hosted concerns without choosing a branch.
- Acceptance checks
- Document the chosen serving contract and resource constraints.
- Related resources
- vLLM Quickstart — Self-hosted serving branch
Stage 3: Recognize service and generation failures
Timeouts, throttling, overload, invalid output, refusals, context limits and broken streams.
Different failures need different responses.
- What you learn
- Timeouts and throttling.
- Overload and invalid output.
- Refusals and context limits.
- Broken streams.
- What you should build
- Classify injected failures and identify the affected request phase.
- Ready when
- You can classify injected failures by type and identify the affected request phase.
- Common mistake
- Treating all errors as transient and retrying everything.
- Acceptance checks
- Classify injected failures and identify the affected request phase.
- Related resources
- FastAPI tutorial — Error handling
- OpenTelemetry signals — Tracing failures
Stage 4: Recovery, deadlines and cancellation
Bounded retries, backoff/jitter, circuit breakers, idempotency and degraded behaviour.
Unbounded retries amplify overload.
- What you learn
- Bounded retries with backoff and jitter.
- Circuit breakers.
- Idempotency and degraded behaviour.
- Deadlines and cancellation.
- What you should build
- Recover within a deadline without retry storms or duplicate actions.
- Ready when
- You can recover from a failure within a deadline without retry storms.
- Common mistake
- Retrying without bounds or jitter.
- Acceptance checks
- Recover within a deadline without retry storms or duplicate actions.
- Related resources
- AWS: control and limit retries — Bounded retry and backoff
Stage 5: Measure and reduce latency
Queueing, TTFT, output generation, retrieval, tools and network time.
Latency percentiles reveal tail behaviour that affects user experience.
- What you learn
- Queueing and TTFT.
- Output generation time.
- Retrieval and tool latency.
- Network time and streaming limits.
- What you should build
- Compare latency percentiles on the same workload; explain streaming limits.
- Ready when
- You can compare latency percentiles and explain what contributes to total latency.
- Common mistake
- Measuring only mean latency without percentiles.
- Acceptance checks
- Compare latency percentiles on the same workload.
- Related resources
- OpenTelemetry signals — Latency tracing
- vLLM Quickstart — Serving latency
Stage 6: Control token cost and caching
Token budgets, prompt-prefix caching, response caching and optional semantic caching.
Cost control determines sustainability of LLM applications.
- What you learn
- Token budgets.
- Prompt-prefix caching.
- Response caching.
- Optional semantic caching.
- What you should build
- Measure cost per successful task; test freshness and permission boundaries.
- Ready when
- You can measure cost per successful task and test cache freshness and permissions.
- Common mistake
- Confusing prompt caching with response caching.
- Acceptance checks
- Measure cost per successful task; test freshness and permission boundaries.
- Related resources
- Claude prompt caching — Provider-specific caching
Stage 7: Scale concurrency and route requests
Backpressure, admission, quotas, queues, routing and tested fallback.
Scaling decisions affect quality, cost and access control.
- What you learn
- Backpressure and admission.
- Quotas and queues.
- Request routing.
- Tested fallback.
- What you should build
- Load-test overload and fallback without silent quality regressions.
- Ready when
- Your load test shows how the system behaves under overload and fallback.
- Common mistake
- Routing to a cheaper model without evaluating quality impact.
- Acceptance checks
- Load-test overload and fallback without silent quality regressions.
- Related resources
- vLLM Quickstart — Serving concurrency
- AWS: control and limit retries — Backpressure and retries
Stage 8: Evaluate and release changes
Version application, prompts, model settings and datasets; canary and rollback.
Unversioned changes make regressions untraceable.
- What you learn
- Version application, prompts and model settings.
- Dataset versioning.
- Canary releases.
- Rollback.
- What you should build
- Block a regression and restore a previous complete release.
- Ready when
- You can block a regression and restore a previous release.
- Common mistake
- Changing prompts or models without versioning or evaluation.
- Acceptance checks
- Block a regression and restore a previous complete release.
- Related resources
- MLflow LLM documentation — Evaluation and versioning
Stage 9: Operate retrieval when the application uses it
Optional ingestion, index versions, permissions, freshness, deletion and embedding changes.
Stale indexes cause grounding failures.
- What you learn
- Ingestion and index versions.
- Permissions and freshness.
- Deletion propagation.
- Embedding changes.
- What you should build
- Demonstrate an index rollback and document deletion.
- Ready when
- Your index rollback preserves access rules and retrieval regression is detected.
- Common mistake
- Not versioning indexes or propagating deletions.
- Acceptance checks
- Demonstrate an index rollback and document deletion.
- Related resources
- LangSmith RAG evaluation tutorial — Retrieval operations and evaluation
Stage 10: Security, SLOs and incident response
Authorization, redaction, injection boundaries, operational targets and runbooks.
Operational targets and security boundaries protect users.
- What you learn
- Authorization and redaction.
- Injection boundaries.
- Operational targets and SLOs.
- Incident runbooks.
- What you should build
- Investigate and recover from a quality or availability incident.
- Ready when
- You can investigate and recover from a quality or availability incident.
- Common mistake
- Not redacting prompts in logs, exposing secrets.
- Acceptance checks
- Investigate and recover from a quality or availability incident.
- Related resources
- MCP security guidance — Security boundaries
- Google SRE monitoring — SLOs and incidents
- Google SRE: alerting on SLOs — SLO-based alerting
Stage 1: Service contracts, tracing and a quality baseline
Request/output contracts, configuration, versioning, traces and task evaluation.
Without a baseline, changes cannot be evaluated.
- What you learn
- Request and output contracts.
- Configuration and versioning.
- Traces and task evaluation.
- Quality baseline.
- What you should build
- Record a baseline request's quality, latency and usage.
- Ready when
- You can record a baseline request's quality, latency and token usage.
- Common mistake
- Operating without a quality baseline or request contract.
- Acceptance checks
- Record a baseline request's quality, latency and usage.
- Related resources
- FastAPI tutorial — API contracts
- MLflow LLM documentation — Evaluation
- OpenTelemetry signals — Tracing
Stage 2: Managed APIs or self-hosted inference
Choose a branch: provider quotas and model versions, or hardware, model loading and serving.
The two branches have different operational concerns.
- What you learn
- Managed API branch.
- Self-hosted inference branch.
- Provider quotas and model versions.
- Hardware, model loading and serving.
- What you should build
- Document the chosen serving contract and resource constraints.
- Ready when
- You can document the chosen serving branch and its resource constraints.
- Common mistake
- Mixing managed and self-hosted concerns without choosing a branch.
- Acceptance checks
- Document the chosen serving contract and resource constraints.
- Related resources
- vLLM Quickstart — Self-hosted serving branch
Stage 3: Recognize service and generation failures
Timeouts, throttling, overload, invalid output, refusals, context limits and broken streams.
Different failures need different responses.
- What you learn
- Timeouts and throttling.
- Overload and invalid output.
- Refusals and context limits.
- Broken streams.
- What you should build
- Classify injected failures and identify the affected request phase.
- Ready when
- You can classify injected failures by type and identify the affected request phase.
- Common mistake
- Treating all errors as transient and retrying everything.
- Acceptance checks
- Classify injected failures and identify the affected request phase.
- Related resources
- FastAPI tutorial — Error handling
- OpenTelemetry signals — Tracing failures
Stage 4: Recovery, deadlines and cancellation
Bounded retries, backoff/jitter, circuit breakers, idempotency and degraded behaviour.
Unbounded retries amplify overload.
- What you learn
- Bounded retries with backoff and jitter.
- Circuit breakers.
- Idempotency and degraded behaviour.
- Deadlines and cancellation.
- What you should build
- Recover within a deadline without retry storms or duplicate actions.
- Ready when
- You can recover from a failure within a deadline without retry storms.
- Common mistake
- Retrying without bounds or jitter.
- Acceptance checks
- Recover within a deadline without retry storms or duplicate actions.
- Related resources
- AWS: control and limit retries — Bounded retry and backoff
Stage 5: Measure and reduce latency
Queueing, TTFT, output generation, retrieval, tools and network time.
Latency percentiles reveal tail behaviour that affects user experience.
- What you learn
- Queueing and TTFT.
- Output generation time.
- Retrieval and tool latency.
- Network time and streaming limits.
- What you should build
- Compare latency percentiles on the same workload; explain streaming limits.
- Ready when
- You can compare latency percentiles and explain what contributes to total latency.
- Common mistake
- Measuring only mean latency without percentiles.
- Acceptance checks
- Compare latency percentiles on the same workload.
- Related resources
- OpenTelemetry signals — Latency tracing
- vLLM Quickstart — Serving latency
Stage 6: Control token cost and caching
Token budgets, prompt-prefix caching, response caching and optional semantic caching.
Cost control determines sustainability of LLM applications.
- What you learn
- Token budgets.
- Prompt-prefix caching.
- Response caching.
- Optional semantic caching.
- What you should build
- Measure cost per successful task; test freshness and permission boundaries.
- Ready when
- You can measure cost per successful task and test cache freshness and permissions.
- Common mistake
- Confusing prompt caching with response caching.
- Acceptance checks
- Measure cost per successful task; test freshness and permission boundaries.
- Related resources
- Claude prompt caching — Provider-specific caching
Stage 7: Scale concurrency and route requests
Backpressure, admission, quotas, queues, routing and tested fallback.
Scaling decisions affect quality, cost and access control.
- What you learn
- Backpressure and admission.
- Quotas and queues.
- Request routing.
- Tested fallback.
- What you should build
- Load-test overload and fallback without silent quality regressions.
- Ready when
- Your load test shows how the system behaves under overload and fallback.
- Common mistake
- Routing to a cheaper model without evaluating quality impact.
- Acceptance checks
- Load-test overload and fallback without silent quality regressions.
- Related resources
- vLLM Quickstart — Serving concurrency
- AWS: control and limit retries — Backpressure and retries
Stage 8: Evaluate and release changes
Version application, prompts, model settings and datasets; canary and rollback.
Unversioned changes make regressions untraceable.
- What you learn
- Version application, prompts and model settings.
- Dataset versioning.
- Canary releases.
- Rollback.
- What you should build
- Block a regression and restore a previous complete release.
- Ready when
- You can block a regression and restore a previous release.
- Common mistake
- Changing prompts or models without versioning or evaluation.
- Acceptance checks
- Block a regression and restore a previous complete release.
- Related resources
- MLflow LLM documentation — Evaluation and versioning
Stage 9: Operate retrieval when the application uses it
Optional ingestion, index versions, permissions, freshness, deletion and embedding changes.
Stale indexes cause grounding failures.
- What you learn
- Ingestion and index versions.
- Permissions and freshness.
- Deletion propagation.
- Embedding changes.
- What you should build
- Demonstrate an index rollback and document deletion.
- Ready when
- Your index rollback preserves access rules and retrieval regression is detected.
- Common mistake
- Not versioning indexes or propagating deletions.
- Acceptance checks
- Demonstrate an index rollback and document deletion.
- Related resources
- LangSmith RAG evaluation tutorial — Retrieval operations and evaluation
Stage 10: Security, SLOs and incident response
Authorization, redaction, injection boundaries, operational targets and runbooks.
Operational targets and security boundaries protect users.
- What you learn
- Authorization and redaction.
- Injection boundaries.
- Operational targets and SLOs.
- Incident runbooks.
- What you should build
- Investigate and recover from a quality or availability incident.
- Ready when
- You can investigate and recover from a quality or availability incident.
- Common mistake
- Not redacting prompts in logs, exposing secrets.
- Acceptance checks
- Investigate and recover from a quality or availability incident.
- Related resources
- MCP security guidance — Security boundaries
- Google SRE monitoring — SLOs and incidents
- Google SRE: alerting on SLOs — SLO-based alerting
From roadmap to production
Build production LLMOps systems with instructor feedback
You have the framework. The View the LLMOps syllabus adds what self-study cannot: live instruction, instructor-reviewed labs, production deployment drills and a capstone that proves you can ship and operate — not just understand.
Fees, schedules and enrolment details are on the course page. No placement, salary or outcome is guaranteed.
Capstone
Operate an LLM service under failure and load
Use one API-backed or self-hosted service. Demonstrate provider/worker failure, throttling, cancellation, load, routing and rollback. Report quality, TTFT, end-to-end latency and cost under a specified workload.
Training alignment
How this roadmap aligns with SCAI's LLMOps course
This roadmap is free and self-paced. SCAI's LLMOps course covers serving, reliability, evaluation and cost control with live instruction and guided labs.
The course adds what the roadmap cannot: instructor review of your evaluation strategy and serving configuration, plus deployment drills with provider failure injection and latency optimization practice. If you prefer independent study, this roadmap gives you the full framework.
What to read next
What to read next
For application construction — APIs, data and model integration — see the AI Developer roadmap. For generative AI model concepts, see the Generative AI roadmap. For broader ML lifecycle operations, see the MLOps roadmap.
Related learning
- Continue to the Generative AI roadmapFor the LLM and RAG fundamentals LLMOps operates.
- Continue to the MLOps roadmapFor the broader ML operations foundation.
- Continue to the AIOps roadmapTo broaden across all production AI operations.
- Understand LLMOps in depthDefinition, lifecycle, platform architecture and operational responsibility.
- Compare MLOps and LLMOps and AIOpsWhere each operations track starts and ends.
FAQ
LLMOps Roadmap — Frequently Asked Questions
Direct answers for engineers operating LLM and generative AI systems in production.
Can I learn LLMOps with hosted APIs?
Yes. The managed-API branch covers provider quotas, model versions, failure handling, cost and evaluation. Self-hosted serving is a separate branch for teams that operate inference infrastructure.
Does streaming make generation faster?
Streaming improves perceived latency by showing tokens early, but total generation time is not necessarily shorter. Measure TTFT and end-to-end latency separately to understand what streaming changes for users.
How do prompt, response and semantic caching differ?
Prompt-prefix caching reduces input token cost for repeated prefixes; response caching returns identical answers for identical requests; semantic caching returns answers for similar but not identical requests. Each has different freshness and permission risks.
Which failures should I retry?
Retry transient failures such as throttling and timeouts with bounded backoff and jitter. Do not retry invalid outputs, refusals or context-limit errors; handle them with fallback or degraded behaviour.
When is model fallback safe to use?
Fallback is safe when the fallback model has been evaluated on the same task and its quality impact is measured. Routing to a cheaper model without evaluation can cause silent quality regressions.