AI Engineer Roadmap
Select, adapt, evaluate and integrate models into a working AI system.
AI engineering combines software, data and model evaluation. Learn to build a reliable baseline, understand the models you use, and connect them to an application. Develop depth in one area such as language or vision, then practise deployment and monitoring. The required depth depends on whether the role builds applications, adapts models or develops training systems.
Quick answer
What is the right AI engineer roadmap?
AI engineering combines software, data and model evaluation. Learn to build a reliable baseline, understand the models you use, and connect them to an application. Develop depth in one area such as language or vision, then practise deployment and monitoring. The required depth depends on whether the role builds applications, adapts models or develops training systems.
Sources and methodology · This roadmap is reviewed when production practices, tools or platform patterns materially change.
Stages
9
Last reviewed
16 September 2026
Stage 1: Software engineering and applied maths
Python, version control, testing, linear algebra, calculus and probability for ML.
Models live inside software; you cannot evaluate or serve them without solid engineering and maths.
- What you learn
- Python and tooling.
- Git and testing.
- Linear algebra.
- Probability and calculus.
- What you should build
- Ship a tested Python module that loads data, computes summary statistics and runs under CI.
- Ready when
- You can ship a tested Python module with CI and explain the linear-algebra operations it uses.
- Common mistake
- Skipping testing and version control while chasing model topics.
- Acceptance checks
- Ship a tested Python module with CI and explain the linear-algebra operations it uses.
- Related resources
- Google ML Crash Course — Applied maths and ML foundations
Stage 2: Data, labels and validation design
Collect, label and split data so that evaluation reflects real use.
A model evaluated on leaked or mislabelled data looks good and fails in production.
- What you learn
- Data collection.
- Labelling and annotation.
- Split strategies.
- Validation design.
- What you should build
- Build a labelled dataset with a documented split strategy and validation checks.
- Ready when
- You can produce a labelled dataset with a defensible split and validation rationale.
- Common mistake
- Splitting data randomly when groups or time dependencies exist.
- Acceptance checks
- Produce a labelled dataset with a defensible split and validation rationale.
- Related resources
- scikit-learn common pitfalls — Data leakage and split design
Stage 3: Build classical ML baselines
Train and evaluate regression and classification baselines with sound metrics.
A strong classical baseline tells you whether deep learning is even necessary.
- What you learn
- Regression baselines.
- Classification baselines.
- Metrics and error analysis.
- What you should build
- Train a classical model and compare it with a naive baseline on held-out data.
- Ready when
- You can train a classical model that beats a naive baseline and explain its errors.
- Common mistake
- Jumping to deep learning before establishing a classical baseline.
- Acceptance checks
- Train a classical model that beats a naive baseline and explain its errors.
- Related resources
- scikit-learn user guide — Classical models and evaluation
Stage 4: Deep learning and transfer learning
Train neural networks and fine-tune pretrained models with reproducible runs.
Transfer learning is how most production models are built; training from scratch is rare.
- What you learn
- Neural network basics.
- Transfer learning.
- Fine-tuning.
- Experiment tracking.
- What you should build
- Fine-tune a pretrained model and log metrics across reproducible runs.
- Ready when
- You can fine-tune a pretrained model and reproduce its metrics from a clean run.
- Common mistake
- Training large models from scratch without compute or data to justify it.
- Acceptance checks
- Fine-tune a pretrained model and reproduce its metrics from a clean run.
- Related resources
- Hugging Face Transformers documentation — Pretrained models and fine-tuning
Stage 5: Choose language, vision or another specialization
Develop depth in one modality based on the problem you want to solve.
AI engineers are hired for depth in a modality, not shallow coverage of all of them.
- What you learn
- Language specialization.
- Vision specialization.
- Other modalities.
- Choosing a specialization.
- What you should build
- Complete a project in your chosen specialization with evaluation.
- Ready when
- You can ship a small project in your specialization with a defensible evaluation.
- Common mistake
- Trying to learn language, vision and audio simultaneously.
- Acceptance checks
- Ship a small project in your specialization with a defensible evaluation.
- Related resources
- Hugging Face course — Language and vision model courses
Stage 6: Choose retrieval, adaptation or a simpler approach
Decide whether to ground with retrieval, adapt with fine-tuning, or keep the model unchanged.
Choosing the wrong method wastes compute and can make the system worse.
- What you learn
- Retrieval option.
- Adaptation option.
- Simpler approach.
- Decision framework.
- What you should build
- Compare a retrieval-only, an adaptation-only and a baseline approach on the same task.
- Ready when
- You can compare retrieval, adaptation and a baseline and justify your choice.
- Common mistake
- Fine-tuning when retrieval would solve the problem at lower cost.
- Acceptance checks
- Compare retrieval, adaptation and a baseline and justify your choice.
- Related resources
- LangChain retrieval documentation — Retrieval patterns and trade-offs
Stage 7: Integrate the model into a system
Wrap the model behind an API with validation, caching and error handling.
A model is not a system until it has an interface, validation and recovery behaviour.
- What you learn
- Model serving API.
- Input and output validation.
- Caching and latency.
- Error handling.
- What you should build
- Serve your model behind an API with input validation and a documented failure mode.
- Ready when
- You can serve a model behind an API that validates input and handles a failure mode.
- Common mistake
- Exposing model internals directly without an API contract.
- Acceptance checks
- Serve a model behind an API that validates input and handles a failure mode.
- Related resources
- FastAPI documentation — API design and validation
Stage 8: Evaluate quality, robustness and trade-offs
Measure accuracy, latency, cost and failure behaviour under realistic conditions.
Production systems live on trade-offs, not accuracy alone.
- What you learn
- Quality metrics.
- Robustness testing.
- Trade-off analysis.
- What you should build
- Produce an evaluation report covering quality, latency, cost and a failure case.
- Ready when
- You can produce an evaluation report that reports quality, latency, cost and a failure case.
- Common mistake
- Optimizing accuracy while ignoring latency or cost budgets.
- Acceptance checks
- Produce an evaluation report covering quality, latency, cost and a failure case.
- Related resources
- Hugging Face evaluation guide — Metrics and evaluation suites
Stage 9: Deploy and monitor the AI service
Deploy with observability, drift detection and a rollback plan.
Models degrade silently; monitoring is how you detect and recover.
- What you learn
- Deployment.
- Monitoring and drift.
- Rollback.
- Operational runbook.
- What you should build
- Deploy the service with metrics, drift detection and a documented rollback.
- Ready when
- You can deploy with metrics, drift detection and a rollback procedure.
- Common mistake
- Deploying without monitoring or a rollback plan for model-side failures.
- Acceptance checks
- Deploy with metrics, drift detection and a rollback procedure.
- Related resources
- OpenTelemetry documentation — Observability for production services
Stage 1: Software engineering and applied maths
Python, version control, testing, linear algebra, calculus and probability for ML.
Models live inside software; you cannot evaluate or serve them without solid engineering and maths.
- What you learn
- Python and tooling.
- Git and testing.
- Linear algebra.
- Probability and calculus.
- What you should build
- Ship a tested Python module that loads data, computes summary statistics and runs under CI.
- Ready when
- You can ship a tested Python module with CI and explain the linear-algebra operations it uses.
- Common mistake
- Skipping testing and version control while chasing model topics.
- Acceptance checks
- Ship a tested Python module with CI and explain the linear-algebra operations it uses.
- Related resources
- Google ML Crash Course — Applied maths and ML foundations
Stage 2: Data, labels and validation design
Collect, label and split data so that evaluation reflects real use.
A model evaluated on leaked or mislabelled data looks good and fails in production.
- What you learn
- Data collection.
- Labelling and annotation.
- Split strategies.
- Validation design.
- What you should build
- Build a labelled dataset with a documented split strategy and validation checks.
- Ready when
- You can produce a labelled dataset with a defensible split and validation rationale.
- Common mistake
- Splitting data randomly when groups or time dependencies exist.
- Acceptance checks
- Produce a labelled dataset with a defensible split and validation rationale.
- Related resources
- scikit-learn common pitfalls — Data leakage and split design
Stage 3: Build classical ML baselines
Train and evaluate regression and classification baselines with sound metrics.
A strong classical baseline tells you whether deep learning is even necessary.
- What you learn
- Regression baselines.
- Classification baselines.
- Metrics and error analysis.
- What you should build
- Train a classical model and compare it with a naive baseline on held-out data.
- Ready when
- You can train a classical model that beats a naive baseline and explain its errors.
- Common mistake
- Jumping to deep learning before establishing a classical baseline.
- Acceptance checks
- Train a classical model that beats a naive baseline and explain its errors.
- Related resources
- scikit-learn user guide — Classical models and evaluation
Stage 4: Deep learning and transfer learning
Train neural networks and fine-tune pretrained models with reproducible runs.
Transfer learning is how most production models are built; training from scratch is rare.
- What you learn
- Neural network basics.
- Transfer learning.
- Fine-tuning.
- Experiment tracking.
- What you should build
- Fine-tune a pretrained model and log metrics across reproducible runs.
- Ready when
- You can fine-tune a pretrained model and reproduce its metrics from a clean run.
- Common mistake
- Training large models from scratch without compute or data to justify it.
- Acceptance checks
- Fine-tune a pretrained model and reproduce its metrics from a clean run.
- Related resources
- Hugging Face Transformers documentation — Pretrained models and fine-tuning
Stage 5: Choose language, vision or another specialization
Develop depth in one modality based on the problem you want to solve.
AI engineers are hired for depth in a modality, not shallow coverage of all of them.
- What you learn
- Language specialization.
- Vision specialization.
- Other modalities.
- Choosing a specialization.
- What you should build
- Complete a project in your chosen specialization with evaluation.
- Ready when
- You can ship a small project in your specialization with a defensible evaluation.
- Common mistake
- Trying to learn language, vision and audio simultaneously.
- Acceptance checks
- Ship a small project in your specialization with a defensible evaluation.
- Related resources
- Hugging Face course — Language and vision model courses
Stage 6: Choose retrieval, adaptation or a simpler approach
Decide whether to ground with retrieval, adapt with fine-tuning, or keep the model unchanged.
Choosing the wrong method wastes compute and can make the system worse.
- What you learn
- Retrieval option.
- Adaptation option.
- Simpler approach.
- Decision framework.
- What you should build
- Compare a retrieval-only, an adaptation-only and a baseline approach on the same task.
- Ready when
- You can compare retrieval, adaptation and a baseline and justify your choice.
- Common mistake
- Fine-tuning when retrieval would solve the problem at lower cost.
- Acceptance checks
- Compare retrieval, adaptation and a baseline and justify your choice.
- Related resources
- LangChain retrieval documentation — Retrieval patterns and trade-offs
Stage 7: Integrate the model into a system
Wrap the model behind an API with validation, caching and error handling.
A model is not a system until it has an interface, validation and recovery behaviour.
- What you learn
- Model serving API.
- Input and output validation.
- Caching and latency.
- Error handling.
- What you should build
- Serve your model behind an API with input validation and a documented failure mode.
- Ready when
- You can serve a model behind an API that validates input and handles a failure mode.
- Common mistake
- Exposing model internals directly without an API contract.
- Acceptance checks
- Serve a model behind an API that validates input and handles a failure mode.
- Related resources
- FastAPI documentation — API design and validation
Stage 8: Evaluate quality, robustness and trade-offs
Measure accuracy, latency, cost and failure behaviour under realistic conditions.
Production systems live on trade-offs, not accuracy alone.
- What you learn
- Quality metrics.
- Robustness testing.
- Trade-off analysis.
- What you should build
- Produce an evaluation report covering quality, latency, cost and a failure case.
- Ready when
- You can produce an evaluation report that reports quality, latency, cost and a failure case.
- Common mistake
- Optimizing accuracy while ignoring latency or cost budgets.
- Acceptance checks
- Produce an evaluation report covering quality, latency, cost and a failure case.
- Related resources
- Hugging Face evaluation guide — Metrics and evaluation suites
Stage 9: Deploy and monitor the AI service
Deploy with observability, drift detection and a rollback plan.
Models degrade silently; monitoring is how you detect and recover.
- What you learn
- Deployment.
- Monitoring and drift.
- Rollback.
- Operational runbook.
- What you should build
- Deploy the service with metrics, drift detection and a documented rollback.
- Ready when
- You can deploy with metrics, drift detection and a rollback procedure.
- Common mistake
- Deploying without monitoring or a rollback plan for model-side failures.
- Acceptance checks
- Deploy with metrics, drift detection and a rollback procedure.
- Related resources
- OpenTelemetry documentation — Observability for production services
From roadmap to production
Build production AI Engineer systems with instructor feedback
You have the framework. The View the AI Engineering curriculum adds what self-study cannot: live instruction, instructor-reviewed labs, production deployment drills and a capstone that proves you can ship and operate — not just understand.
Fees, schedules and enrolment details are on the course page. No placement, salary or outcome is guaranteed.
Capstone
Build and evaluate a document-processing system
Build a system that processes a small document set: extraction, a classical or transfer-learning model, an evaluation report covering quality, latency and cost, and a deployed service with monitoring and rollback.
Training alignment
How this roadmap aligns with SCAI's AI Engineering course
This roadmap is free and self-paced. SCAI's AI Engineering course covers model selection, evaluation and system integration with live instruction and guided projects.
The course adds what the roadmap cannot: instructor review of your model choices and evaluation strategy, plus a reviewed capstone that connects model selection to deployment. If you prefer independent study, this roadmap gives you the full framework.
What to read next
What to read next
For predictive model systems with reproducible training and serving, see the ML Engineer roadmap. For generative AI techniques — LLMs, RAG and evaluation — see the Generative AI roadmap. For operating AI systems in production, see the MLOps roadmap and the LLMOps roadmap.
Related learning
- Continue to the AI Developer roadmapFor a software-engineering-first application focus.
- Continue to the ML Engineer roadmapTo go model-centric rather than broad-system.
- Continue to the AIOps roadmapTo specialize in production AI operations and reliability.
- Compare AI Engineer and ML Engineer pathsBroad engineering versus model-centric development.
- Compare the AI Developer and AI Engineer pathsWhere the application track and the broad engineering track differ.
FAQ
AI Engineer Roadmap — Frequently Asked Questions
Direct answers for engineers building end-to-end AI systems.
Do I need to learn deep learning before classical ML?
No. Build classical baselines first; they often beat deep models and always tell you whether deep learning is justified.
Should I specialize in language or vision?
Pick one based on the problems you want to solve. Depth in one modality is more valuable than shallow coverage of both.
When should I fine-tune versus use retrieval?
Fine-tune when retrieval cannot encode the behaviour you need; otherwise prefer retrieval because it is cheaper and easier to update.
What does an AI engineer do that an AI developer does not?
An AI engineer selects, adapts and evaluates models; an AI developer integrates existing models into applications.
How much deployment and monitoring do I need?
Enough to detect drift, recover from failures and roll back a bad model. Without that, production deployments are unsafe.