AIOps Roadmap
Detect, investigate and respond to IT incidents using telemetry, anomaly detection and AI-assisted analysis.
AIOps is not a buzzword — it is the practice of applying AI and analytics to IT operations to reduce alert noise, accelerate investigation and control remediation. This roadmap takes you from service reliability fundamentals through telemetry instrumentation, anomaly detection, event correlation, evidence-based investigation and controlled remediation. Each stage has a build task, acceptance checks and a measurable outcome. By the end, you can instrument a service, detect incidents earlier than static thresholds, investigate with structured evidence and remediate with approval, verification and rollback.
What is the right AIOps roadmap?
Start with IT operations and service reliability — SLOs, incidents, runbooks. Then instrument services with metrics, logs and traces. Add anomaly detection and event correlation to reduce alert noise without missing incidents. Use AI-assisted investigation to gather evidence and rank hypotheses. Automate remediation only with permissions, approval, verification and rollback. Measure improvement with incident recall, not just reduced alert counts.
Sources and methodology · This roadmap is reviewed when production practices, tools or platform patterns materially change.
Stages
9
Last reviewed
16 September 2026
Stage 1: IT operations and service reliability
Understand services, dependencies, incidents, ownership and user-facing reliability measures before adding AI.
AIOps augments IT operations. Without operational foundations, AI-assisted methods have no context to act on.
- What you learn
- Services, dependencies and failure domains.
- Incident lifecycle, severity levels and escalation.
- SLOs, SLAs, error budgets and runbooks.
- Service ownership and dependency mapping.
- What you should build
- Describe the impact and owner of one service failure in a written runbook.
- Ready when
- You can describe one service's failure impact and ownership, and identify its critical dependencies.
- Common mistake
- Trying to apply AI to operations before understanding how incidents are detected and handled manually.
- Acceptance checks
- Describe the impact and owner of one service failure.
- Identify the service's critical dependencies and their failure modes.
- Related resources
- Google SRE monitoring — Foundational monitoring concepts for IT operations
- Google SRE: alerting on SLOs — Reliability targets tied to user impact
Stage 2: Metrics, logs and traces
Collect complementary signals with service and request context.
AI-assisted analysis depends on telemetry quality. No telemetry means no signal to correlate.
- What you learn
- Metrics: counters, gauges, histograms and percentiles.
- Structured logs with correlation IDs and service identity.
- Distributed traces and span context.
- Deployment events and configuration changes.
- What you should build
- Trace a failed request across a small service using metrics, logs and traces together.
- Ready when
- You can instrument a service and correlate a single request across all three telemetry signals.
- Common mistake
- Collecting telemetry without service identity, deployment context or event labels.
- Acceptance checks
- Instrument a small service with metrics, logs and traces.
- Correlate one request across all three telemetry signals.
- Related resources
- OpenTelemetry signals — Canonical reference for metrics, logs and traces
Stage 3: Telemetry quality and service context
Detect and handle missing samples, time alignment, cardinality and deployment changes.
Poor telemetry quality produces false anomalies and missed incidents. AI methods amplify data quality problems.
- What you learn
- Time alignment, clock skew and timestamp normalization.
- Missing samples, gaps and dropout detection.
- High-cardinality labels and cardinality limits.
- Changing topology and deployment context.
- What you should build
- Detect a telemetry gap (missing samples or cardinality spike) without calling the service healthy.
- Ready when
- You can distinguish a telemetry instrumentation failure from a service failure.
- Common mistake
- Treating missing telemetry data as normal service health.
- Acceptance checks
- Detect a telemetry gap and explain whether it is instrumentation failure or service failure.
- Identify a high-cardinality label causing storage or query problems.
- Related resources
- OpenTelemetry signals — Telemetry data quality concepts
- Google SRE monitoring — Service-level monitoring foundations
Stage 4: Baselines and actionable alerts
Establish thresholds, seasonality, maintenance windows and operational impact.
Anomaly detection is only useful when compared against a baseline that accounts for expected patterns.
- What you learn
- Static and dynamic thresholds.
- Seasonality, daily/weekly patterns and traffic changes.
- Expected maintenance windows and deployment changes.
- Alert actionability and ownership.
- What you should build
- Compare alert behaviour during normal traffic and a labelled incident for one service.
- Ready when
- Your baseline alert strategy has explicit false-alert and missed-incident definitions.
- Common mistake
- Setting static thresholds without accounting for seasonality, traffic patterns or maintenance windows.
- Acceptance checks
- Compare alert behaviour during normal traffic and a labelled incident.
- Explain why a static threshold would produce false alerts for a seasonal service.
- Related resources
- Google SRE: alerting on SLOs — Alert design tied to user impact
Stage 5: Anomaly detection
Apply statistical and ML anomaly scoring with temporal evaluation and false positives.
Anomaly detection extends baselines with pattern-based detection, but must be evaluated against labelled incidents.
- What you learn
- Statistical anomaly scoring: z-scores, IQR, EWMA.
- ML-based anomaly detection: isolation forest, autoencoder.
- Temporal evaluation and changing traffic patterns.
- False positives and comparison with baseline.
- What you should build
- Compare a statistical or ML anomaly detector with the simple baseline on held-out incident scenarios.
- Ready when
- Your detector is evaluated against labelled incidents with documented recall, false-alert rate and detection delay.
- Common mistake
- Deploying an anomaly detector without evaluating it against labelled incidents or comparing it with the existing baseline.
- Acceptance checks
- Compare a detector with the simple baseline on held-out incident scenarios.
- Document recall, false-alert rate and detection delay.
- Related resources
- scikit-learn common pitfalls — Avoiding evaluation mistakes in ML methods
- scikit-learn cross-validation — Sound evaluation splits for temporal data
- SCAI's AIOps programme — The course covers ML-based anomaly detection with guided labs and instructor-reviewed evaluation
Stage 6: Event correlation and alert grouping
Deduplicate, group and correlate events using dependency context while preserving distinct incidents.
Correlation reduces alert noise, but grouping can hide a second incident if it is too aggressive.
- What you learn
- Deduplication of repeated alerts.
- Event windows and time-based grouping.
- Dependency context and topology-aware grouping.
- Correlation limits and preserving distinct incident evidence.
- What you should build
- Group related alerts while preserving separate incidents in a simulated event stream.
- Ready when
- Your correlation groups related events without hiding a concurrent distinct incident.
- Common mistake
- Over-grouping events so that a second concurrent incident is hidden inside the first.
- Acceptance checks
- Group related alerts while preserving separate incidents.
- Demonstrate that a concurrent incident is not hidden by the grouping.
- Related resources
- IBM: AIOps — Industry definition of AIOps correlation and investigation
Stage 7: Evidence-based incident investigation
Retrieve runbooks and logs, gather evidence, rank hypotheses and verify; LLM assistance is optional.
AI-assisted investigation can accelerate evidence gathering, but conclusions need uncertainty and verification.
- What you learn
- Retrieve runbooks, logs and relevant context.
- Gather and structure evidence.
- Rank hypotheses and express uncertainty.
- Optional LLM-assisted summarization of incident context.
- What you should build
- Produce an investigation report that distinguishes evidence from suspected causes for a simulated incident.
- Ready when
- Your investigation report separates observations, hypotheses and confirmed causes.
- Common mistake
- Presenting a plausible hypothesis as a confirmed root cause without additional evidence.
- Acceptance checks
- Produce an investigation report that distinguishes evidence from suspected causes.
- Express uncertainty for unconfirmed hypotheses.
- Related resources
- Google SRE monitoring — Service reliability investigation practices
- Agentic AI roadmap — Agent patterns for investigation assistants with bounded tool access
- SCAI's AIOps programme — The course includes simulated incident scenarios with instructor-reviewed investigation reports
Stage 8: Controlled remediation and rollback
Run an approved recovery action and verify whether it restored the service.
Automated remediation can reduce incident time, but uncontrolled actions can cause new incidents.
- What you learn
- Action allowlists and least privilege.
- Approval boundaries and dry runs.
- Idempotency and safe retry of side-effecting actions.
- Recovery verification and rollback.
- What you should build
- Run an approved simulated action and demonstrate rollback.
- Ready when
- Your simulated remediation requires approval and verifies service recovery after the action.
- Common mistake
- Treating a model's recommendation as permission to execute it.
- Acceptance checks
- Replaying the same request does not restart the worker twice.
- An unauthorized action is rejected by the application.
- Failed recovery produces an escalation and a usable rollback record.
- Related resources
- AWS: control and limit retries — Bounded retry and backoff for transient failures
- MCP security guidance — Action boundaries and authorization for tool-using systems
- SCAI's AIOps programme — The course covers production AI operations with deployment drills and failure injection
Stage 9: Measure operational improvement
Track missed incidents, false alerts, detection delay and investigation effort.
Evaluation prevents noise reduction from hiding missed incidents behind lower alert counts.
- What you learn
- Incident recall and missed-incident tracking.
- False-alert rate and alert fatigue measurement.
- Detection delay and investigation time.
- Analyst feedback and qualitative review.
- What you should build
- Report improvements and regressions against the original baseline.
- Ready when
- Your evaluation reports incident recall alongside noise reduction, not alert count alone.
- Common mistake
- Reporting reduced alert counts without measuring whether missed incidents increased.
- Acceptance checks
- Report improvements and regressions against the original baseline.
- Do not report alert count reduction alone without missed-incident data.
- Related resources
- Google SRE: alerting on SLOs — User-impact-based measurement
- IBM: AIOps — AIOps evaluation context
Stage 1: IT operations and service reliability
Understand services, dependencies, incidents, ownership and user-facing reliability measures before adding AI.
AIOps augments IT operations. Without operational foundations, AI-assisted methods have no context to act on.
- What you learn
- Services, dependencies and failure domains.
- Incident lifecycle, severity levels and escalation.
- SLOs, SLAs, error budgets and runbooks.
- Service ownership and dependency mapping.
- What you should build
- Describe the impact and owner of one service failure in a written runbook.
- Ready when
- You can describe one service's failure impact and ownership, and identify its critical dependencies.
- Common mistake
- Trying to apply AI to operations before understanding how incidents are detected and handled manually.
- Acceptance checks
- Describe the impact and owner of one service failure.
- Identify the service's critical dependencies and their failure modes.
- Related resources
- Google SRE monitoring — Foundational monitoring concepts for IT operations
- Google SRE: alerting on SLOs — Reliability targets tied to user impact
Stage 2: Metrics, logs and traces
Collect complementary signals with service and request context.
AI-assisted analysis depends on telemetry quality. No telemetry means no signal to correlate.
- What you learn
- Metrics: counters, gauges, histograms and percentiles.
- Structured logs with correlation IDs and service identity.
- Distributed traces and span context.
- Deployment events and configuration changes.
- What you should build
- Trace a failed request across a small service using metrics, logs and traces together.
- Ready when
- You can instrument a service and correlate a single request across all three telemetry signals.
- Common mistake
- Collecting telemetry without service identity, deployment context or event labels.
- Acceptance checks
- Instrument a small service with metrics, logs and traces.
- Correlate one request across all three telemetry signals.
- Related resources
- OpenTelemetry signals — Canonical reference for metrics, logs and traces
Stage 3: Telemetry quality and service context
Detect and handle missing samples, time alignment, cardinality and deployment changes.
Poor telemetry quality produces false anomalies and missed incidents. AI methods amplify data quality problems.
- What you learn
- Time alignment, clock skew and timestamp normalization.
- Missing samples, gaps and dropout detection.
- High-cardinality labels and cardinality limits.
- Changing topology and deployment context.
- What you should build
- Detect a telemetry gap (missing samples or cardinality spike) without calling the service healthy.
- Ready when
- You can distinguish a telemetry instrumentation failure from a service failure.
- Common mistake
- Treating missing telemetry data as normal service health.
- Acceptance checks
- Detect a telemetry gap and explain whether it is instrumentation failure or service failure.
- Identify a high-cardinality label causing storage or query problems.
- Related resources
- OpenTelemetry signals — Telemetry data quality concepts
- Google SRE monitoring — Service-level monitoring foundations
Stage 4: Baselines and actionable alerts
Establish thresholds, seasonality, maintenance windows and operational impact.
Anomaly detection is only useful when compared against a baseline that accounts for expected patterns.
- What you learn
- Static and dynamic thresholds.
- Seasonality, daily/weekly patterns and traffic changes.
- Expected maintenance windows and deployment changes.
- Alert actionability and ownership.
- What you should build
- Compare alert behaviour during normal traffic and a labelled incident for one service.
- Ready when
- Your baseline alert strategy has explicit false-alert and missed-incident definitions.
- Common mistake
- Setting static thresholds without accounting for seasonality, traffic patterns or maintenance windows.
- Acceptance checks
- Compare alert behaviour during normal traffic and a labelled incident.
- Explain why a static threshold would produce false alerts for a seasonal service.
- Related resources
- Google SRE: alerting on SLOs — Alert design tied to user impact
Stage 5: Anomaly detection
Apply statistical and ML anomaly scoring with temporal evaluation and false positives.
Anomaly detection extends baselines with pattern-based detection, but must be evaluated against labelled incidents.
- What you learn
- Statistical anomaly scoring: z-scores, IQR, EWMA.
- ML-based anomaly detection: isolation forest, autoencoder.
- Temporal evaluation and changing traffic patterns.
- False positives and comparison with baseline.
- What you should build
- Compare a statistical or ML anomaly detector with the simple baseline on held-out incident scenarios.
- Ready when
- Your detector is evaluated against labelled incidents with documented recall, false-alert rate and detection delay.
- Common mistake
- Deploying an anomaly detector without evaluating it against labelled incidents or comparing it with the existing baseline.
- Acceptance checks
- Compare a detector with the simple baseline on held-out incident scenarios.
- Document recall, false-alert rate and detection delay.
- Related resources
- scikit-learn common pitfalls — Avoiding evaluation mistakes in ML methods
- scikit-learn cross-validation — Sound evaluation splits for temporal data
- SCAI's AIOps programme — The course covers ML-based anomaly detection with guided labs and instructor-reviewed evaluation
Stage 6: Event correlation and alert grouping
Deduplicate, group and correlate events using dependency context while preserving distinct incidents.
Correlation reduces alert noise, but grouping can hide a second incident if it is too aggressive.
- What you learn
- Deduplication of repeated alerts.
- Event windows and time-based grouping.
- Dependency context and topology-aware grouping.
- Correlation limits and preserving distinct incident evidence.
- What you should build
- Group related alerts while preserving separate incidents in a simulated event stream.
- Ready when
- Your correlation groups related events without hiding a concurrent distinct incident.
- Common mistake
- Over-grouping events so that a second concurrent incident is hidden inside the first.
- Acceptance checks
- Group related alerts while preserving separate incidents.
- Demonstrate that a concurrent incident is not hidden by the grouping.
- Related resources
- IBM: AIOps — Industry definition of AIOps correlation and investigation
Stage 7: Evidence-based incident investigation
Retrieve runbooks and logs, gather evidence, rank hypotheses and verify; LLM assistance is optional.
AI-assisted investigation can accelerate evidence gathering, but conclusions need uncertainty and verification.
- What you learn
- Retrieve runbooks, logs and relevant context.
- Gather and structure evidence.
- Rank hypotheses and express uncertainty.
- Optional LLM-assisted summarization of incident context.
- What you should build
- Produce an investigation report that distinguishes evidence from suspected causes for a simulated incident.
- Ready when
- Your investigation report separates observations, hypotheses and confirmed causes.
- Common mistake
- Presenting a plausible hypothesis as a confirmed root cause without additional evidence.
- Acceptance checks
- Produce an investigation report that distinguishes evidence from suspected causes.
- Express uncertainty for unconfirmed hypotheses.
- Related resources
- Google SRE monitoring — Service reliability investigation practices
- Agentic AI roadmap — Agent patterns for investigation assistants with bounded tool access
- SCAI's AIOps programme — The course includes simulated incident scenarios with instructor-reviewed investigation reports
Stage 8: Controlled remediation and rollback
Run an approved recovery action and verify whether it restored the service.
Automated remediation can reduce incident time, but uncontrolled actions can cause new incidents.
- What you learn
- Action allowlists and least privilege.
- Approval boundaries and dry runs.
- Idempotency and safe retry of side-effecting actions.
- Recovery verification and rollback.
- What you should build
- Run an approved simulated action and demonstrate rollback.
- Ready when
- Your simulated remediation requires approval and verifies service recovery after the action.
- Common mistake
- Treating a model's recommendation as permission to execute it.
- Acceptance checks
- Replaying the same request does not restart the worker twice.
- An unauthorized action is rejected by the application.
- Failed recovery produces an escalation and a usable rollback record.
- Related resources
- AWS: control and limit retries — Bounded retry and backoff for transient failures
- MCP security guidance — Action boundaries and authorization for tool-using systems
- SCAI's AIOps programme — The course covers production AI operations with deployment drills and failure injection
Stage 9: Measure operational improvement
Track missed incidents, false alerts, detection delay and investigation effort.
Evaluation prevents noise reduction from hiding missed incidents behind lower alert counts.
- What you learn
- Incident recall and missed-incident tracking.
- False-alert rate and alert fatigue measurement.
- Detection delay and investigation time.
- Analyst feedback and qualitative review.
- What you should build
- Report improvements and regressions against the original baseline.
- Ready when
- Your evaluation reports incident recall alongside noise reduction, not alert count alone.
- Common mistake
- Reporting reduced alert counts without measuring whether missed incidents increased.
- Acceptance checks
- Report improvements and regressions against the original baseline.
- Do not report alert count reduction alone without missed-incident data.
- Related resources
- Google SRE: alerting on SLOs — User-impact-based measurement
- IBM: AIOps — AIOps evaluation context
From roadmap to production
Build production AIOps systems with instructor feedback
You have the framework. The Compare SCAI's production AI operations programme adds what self-study cannot: live instruction, instructor-reviewed labs, production deployment drills and a capstone that proves you can ship and operate — not just understand.
Fees, schedules and enrolment details are on the course page. No placement, salary or outcome is guaranteed.
Capstone
Build an incident investigation system
Database latency spike
PostgreSQL
Queries go from 50ms to 5 seconds
Anomaly detector fires before SLO breach
Traces slowdown to missing index after schema migration
Rollback migration with approval + verify recovery
External API rate limiting
Third-party CRM API
429 responses cascade into downstream timeouts
Event correlation groups 47 alerts into 1 incident
Identifies the rate limit as root cause
Throttle requests with circuit breaker
Deployment regression
Canary deployment
Memory leak introduced by new release
Telemetry detects rising memory in 10 min — before users affected
Traces regression to the new deployment version
Rollback to previous version + verify memory stabilizes
Telemetry gap
OpenTelemetry collector
Collector crashes — telemetry stops for 3 minutes
Telemetry-quality check flags as instrumentation failure, not service outage
Confirms service stayed healthy during the gap
Restart collector + verify telemetry resumes
What you deliver
Traces for each scenario, labelled incident labels (true positive, false positive, missed incident), a baseline comparison showing improvement over static thresholds, and a runbook covering all four scenarios. Use a simulator — do not create live destructive infrastructure actions.
Take AIOps further
Go from understanding AIOps to mastering production AI operations
AI for IT operations: telemetry, anomaly detection, event correlation, incident investigation, controlled remediation — the nine stages you just completed
CompletedReproducible training pipelines, model deployment, drift monitoring and rollback — operate ML systems end to end
Build nextLLM serving, prompt versioning, evaluation, cost control and guardrails — operate LLM applications in production
Build nextAgent deployment, tool permissions, execution monitoring and supervision — operate AI agents safely
Build nextYou have the framework — nine stages from IT operations foundations through controlled remediation. SCAI's production AI programme helps you go further: actually building, deploying and operating AI systems in production with instructor feedback at every step.
The programme covers three layers that build on this roadmap's foundations: MLOps — reproducible training pipelines, model deployment, drift monitoring and rollback. LLMOps — LLM serving, prompt versioning, evaluation, cost control and guardrails. AgentOps — agent deployment, tool permissions, execution monitoring and supervision.
Ready to build production AI systems with instructor feedback?
Explore the production AI programmeWhat to read next
What to read next
MLOps Roadmap
Reproducible training pipelines, model releases, drift monitoring and rollback — the ML-specific operations that AIOps detects but MLOps prevents.
LLMOps Roadmap
LLM serving, prompt versioning, evaluation, latency and cost control — the LLM-specific operations that generate a different class of incidents.
Agentic AI Roadmap
Controlled agent patterns with tool permissions, state management and failure recovery — the patterns that AI-assisted investigation tools use.
SCAI's AIOps Programme
All three — MLOps, LLMOps and AgentOps — in one structured programme with instructor review, guided labs and a reviewed capstone.
Explore programmeRelated learning
- Continue to the MLOps roadmapFor the ML-specific operations core that AIOps augments.
- Continue to the LLMOps roadmapFor LLM-specific operations.
- Continue to the Agentic AI roadmapFor the agent patterns used in AI-assisted investigation with bounded tool access.
- Compare MLOps and LLMOps and AIOpsWhere each operations track starts and ends.
- Compare the MLOps and AIOps coursesCourse-level decision between ML operations and AI for IT operations.
FAQ
AIOps questions
Direct answers about AIOps as AI for IT operations.
How is AIOps different from MLOps and LLMOps?
AIOps applies AI to IT operations — detecting, investigating and responding to incidents in production systems. MLOps manages the ML model lifecycle: training, validation, deployment and monitoring of predictive models. LLMOps manages LLM application lifecycle: prompt versioning, evaluation, serving and cost control. They overlap in production environments but answer different questions. AIOps asks 'what is breaking and why?' MLOps asks 'is my model still accurate?' LLMOps asks 'is my LLM application fast, correct and affordable?'
Can a DevOps or SRE engineer start here?
Yes — your operational experience is the right starting point. You already understand services, incidents, monitoring and runbooks. What you need to add is telemetry instrumentation with OpenTelemetry, statistical evaluation for anomaly detection, and evidence-based investigation methods. You do not need to become a data scientist first; you need to learn enough statistics to evaluate whether a detector actually improves on your existing alerts.
Do I need an LLM for AIOps?
No. Statistical anomaly detection (z-scores, EWMA, isolation forests) and event correlation work without any language model. An LLM becomes useful when you need to summarize long log streams, retrieve relevant runbooks from a knowledge base, or help an on-call engineer rank hypotheses during a stressful incident. Start with statistics and correlation. Add an LLM only when text interpretation would save meaningful investigation time.
Does an anomaly identify the root cause?
No. An anomaly tells you something unusual is happening. It does not tell you why. A latency spike could be a database issue, a deployment regression, a traffic surge or a dependency failure. Root cause requires service context (what changed recently?), dependency knowledge (what is downstream?) and verification (does rolling back the change fix it?). An anomaly is the starting point of an investigation, not the conclusion.
When should remediation be automated?
Only after you have defined four things: which actions are allowed (an allowlist, not arbitrary commands), who approves (human approval for impactful actions, automatic only for safe and reversible ones), how you verify recovery (health checks after the action, not just 'the command succeeded'), and how you roll back if the remediation makes things worse. Start with detection and recommendations. Automate only the actions you would confidently take manually at 3am.
Is this roadmap the same as SCAI's AIOps course?
This roadmap gives you the complete AIOps framework — nine stages with build tasks, acceptance checks and resource links. SCAI's production AI programme takes you further: you build real systems with MLOps, LLMOps and AgentOps, get instructor feedback on your evaluation reports and runbooks, and complete a reviewed capstone. The roadmap is the foundation; the programme is the guided practice that turns understanding into production-ready skill.
What should I learn after this roadmap?
If your incidents often involve ML model quality degradation, the MLOps roadmap covers the prevention side — reproducible training, drift monitoring and model rollback. If your incidents involve LLM applications (token rate limits, context overflow, provider outages), the LLMOps roadmap covers that specific operations layer. If you want to build AI-assisted investigation tools with bounded tool access, the Agentic AI roadmap covers agent patterns, permissions and failure recovery.