AI Developer Roadmap
Build and deploy an application that uses models, data and tools reliably.
An AI developer builds software that uses models to perform useful tasks. Start with APIs, data storage and testing, then integrate a model and define how its outputs will be checked. Add retrieval or tools when the application requires them. Deploy the application with authentication, observability, usage limits and recovery behaviour.
Quick answer
What is the right AI developer roadmap?
An AI developer builds software that uses models to perform useful tasks. Start with APIs, data storage and testing, then integrate a model and define how its outputs will be checked. Add retrieval or tools when the application requires them. Deploy the application with authentication, observability, usage limits and recovery behaviour.
Sources and methodology · This roadmap is reviewed when production practices, tools or platform patterns materially change.
Stages
8
Last reviewed
16 September 2026
Stage 1: Application foundations: APIs, data and tests
Design APIs, persist data and write tests that catch real defects.
A model bolted onto an unstable application produces failures that are hard to attribute.
- What you learn
- REST API design.
- Data persistence.
- Input validation.
- Automated tests.
- What you should build
- Build a small CRUD API with a database, input validation and automated tests.
- Ready when
- You can ship a tested API that persists data and rejects invalid input.
- Common mistake
- Skipping tests and validation because the model will "handle it".
- Acceptance checks
- Ship a tested API that persists data and rejects invalid input.
- Related resources
- FastAPI documentation — API design, validation and testing patterns
Stage 2: Integrate a model through an SDK or API
Call a hosted model, handle retries and timeouts, and log requests.
Model calls fail in different ways than databases; treating them like simple functions causes outages.
- What you learn
- Model SDKs and clients.
- Retries and timeouts.
- Structured logging.
- Token cost tracking.
- What you should build
- Add a model call to your API with retry, timeout and structured logging.
- Ready when
- You can call a model SDK with retry, timeout and logged request metadata.
- Common mistake
- Calling a model SDK synchronously without timeouts or error handling.
- Acceptance checks
- Call a model SDK with retry, timeout and logged request metadata.
- Related resources
- OpenAI API documentation — SDK usage, retries and error codes
Stage 3: Validate structured outputs
Parse model outputs into typed contracts and reject malformed responses.
Downstream code assumes structure; unvalidated outputs cause silent data corruption.
- What you learn
- Output schemas.
- Typed parsers.
- Malformed-response handling.
- What you should build
- Validate a model response against a schema and return a typed result or a controlled error.
- Ready when
- You can validate a model response against a schema and handle failures without crashing.
- Common mistake
- Trusting JSON from a model without schema validation.
- Acceptance checks
- Validate a model response against a schema and handle failures without crashing.
- Related resources
- Pydantic documentation — Schema validation and typed parsing
Stage 4: Manage conversations and streaming
Maintain conversation state, stream tokens and isolate user sessions.
Stateless model calls leak context between users and break multi-turn experiences.
- What you learn
- Conversation state.
- Streaming responses.
- Session isolation.
- Context window management.
- What you should build
- Add a streaming chat endpoint with per-session conversation history.
- Ready when
- You can stream a model response while keeping session history isolated.
- Common mistake
- Storing all conversation history in a single global buffer.
- Acceptance checks
- Add a streaming chat endpoint with per-session conversation history.
- Related resources
- OpenAI streaming guide — Server-sent events and token streaming
Stage 5: Add retrieval when external knowledge is needed
Index documents, retrieve relevant passages and ground responses with citations.
Models hallucinate when they lack grounding; retrieval reduces that risk for knowledge-bound tasks.
- What you learn
- Embeddings and indexing.
- Vector store.
- Grounding and citations.
- Ranking and reranking.
- What you should build
- Add a retrieval step that supplies cited passages to the model for a small document set.
- Ready when
- You can retrieve passages and return an answer with citations to source documents.
- Common mistake
- Dumping raw chunks into the prompt without ranking or citations.
- Acceptance checks
- Retrieve passages and return an answer with citations to source documents.
- Related resources
- LangChain retrieval documentation — Embeddings, vector stores and retrieval patterns
Stage 6: Add tools when the application needs actions
Expose bounded tools, confirm destructive actions and audit tool calls.
Unbounded tool access lets a model take actions the user never approved.
- What you learn
- Tool definitions.
- Boundaries and permissions.
- Confirmation for writes.
- Audit logs.
- What you should build
- Add a read-only tool and a confirmed write tool to your application.
- Ready when
- You can expose a bounded tool and gate a write action behind confirmation.
- Common mistake
- Giving the model direct database access without a sandboxed interface.
- Acceptance checks
- Expose a bounded tool and gate a write action behind confirmation.
- Related resources
- OpenAI function calling guide — Tool definitions, schemas and call patterns
Stage 7: Evaluate behaviour and failure cases
Build an evaluation set, measure outputs and track regressions over changes.
Without evaluation, model changes that look fine locally break behaviour users depend on.
- What you learn
- Evaluation set.
- Metrics and rubrics.
- Regression tracking.
- What you should build
- Create an evaluation set and a script that scores outputs before each deploy.
- Ready when
- You can run an evaluation script and report a score that blocks a bad deploy.
- Common mistake
- Shipping model or prompt changes with no evaluation baseline.
- Acceptance checks
- Run an evaluation script and report a score that blocks a bad deploy.
- Related resources
- LangSmith evaluation guide — Datasets, evaluators and regression testing
Stage 8: Deploy, observe and recover
Add authentication, observability, usage limits and rollback behaviour.
An AI application in production fails differently than a CRUD app; you need recovery, not just uptime.
- What you learn
- Authentication and authorization.
- Metrics, logs and traces.
- Usage limits and quotas.
- Rollback and recovery.
- What you should build
- Deploy the application with auth, metrics, a usage cap and a documented rollback.
- Ready when
- You can deploy with auth, observability, a usage limit and a rollback procedure.
- Common mistake
- Deploying without rate limits or a rollback plan for model-side failures.
- Acceptance checks
- Deploy with auth, observability, a usage limit and a rollback procedure.
- Related resources
- OpenTelemetry documentation — Traces, metrics and logs for production services
Stage 1: Application foundations: APIs, data and tests
Design APIs, persist data and write tests that catch real defects.
A model bolted onto an unstable application produces failures that are hard to attribute.
- What you learn
- REST API design.
- Data persistence.
- Input validation.
- Automated tests.
- What you should build
- Build a small CRUD API with a database, input validation and automated tests.
- Ready when
- You can ship a tested API that persists data and rejects invalid input.
- Common mistake
- Skipping tests and validation because the model will "handle it".
- Acceptance checks
- Ship a tested API that persists data and rejects invalid input.
- Related resources
- FastAPI documentation — API design, validation and testing patterns
Stage 2: Integrate a model through an SDK or API
Call a hosted model, handle retries and timeouts, and log requests.
Model calls fail in different ways than databases; treating them like simple functions causes outages.
- What you learn
- Model SDKs and clients.
- Retries and timeouts.
- Structured logging.
- Token cost tracking.
- What you should build
- Add a model call to your API with retry, timeout and structured logging.
- Ready when
- You can call a model SDK with retry, timeout and logged request metadata.
- Common mistake
- Calling a model SDK synchronously without timeouts or error handling.
- Acceptance checks
- Call a model SDK with retry, timeout and logged request metadata.
- Related resources
- OpenAI API documentation — SDK usage, retries and error codes
Stage 3: Validate structured outputs
Parse model outputs into typed contracts and reject malformed responses.
Downstream code assumes structure; unvalidated outputs cause silent data corruption.
- What you learn
- Output schemas.
- Typed parsers.
- Malformed-response handling.
- What you should build
- Validate a model response against a schema and return a typed result or a controlled error.
- Ready when
- You can validate a model response against a schema and handle failures without crashing.
- Common mistake
- Trusting JSON from a model without schema validation.
- Acceptance checks
- Validate a model response against a schema and handle failures without crashing.
- Related resources
- Pydantic documentation — Schema validation and typed parsing
Stage 4: Manage conversations and streaming
Maintain conversation state, stream tokens and isolate user sessions.
Stateless model calls leak context between users and break multi-turn experiences.
- What you learn
- Conversation state.
- Streaming responses.
- Session isolation.
- Context window management.
- What you should build
- Add a streaming chat endpoint with per-session conversation history.
- Ready when
- You can stream a model response while keeping session history isolated.
- Common mistake
- Storing all conversation history in a single global buffer.
- Acceptance checks
- Add a streaming chat endpoint with per-session conversation history.
- Related resources
- OpenAI streaming guide — Server-sent events and token streaming
Stage 5: Add retrieval when external knowledge is needed
Index documents, retrieve relevant passages and ground responses with citations.
Models hallucinate when they lack grounding; retrieval reduces that risk for knowledge-bound tasks.
- What you learn
- Embeddings and indexing.
- Vector store.
- Grounding and citations.
- Ranking and reranking.
- What you should build
- Add a retrieval step that supplies cited passages to the model for a small document set.
- Ready when
- You can retrieve passages and return an answer with citations to source documents.
- Common mistake
- Dumping raw chunks into the prompt without ranking or citations.
- Acceptance checks
- Retrieve passages and return an answer with citations to source documents.
- Related resources
- LangChain retrieval documentation — Embeddings, vector stores and retrieval patterns
Stage 6: Add tools when the application needs actions
Expose bounded tools, confirm destructive actions and audit tool calls.
Unbounded tool access lets a model take actions the user never approved.
- What you learn
- Tool definitions.
- Boundaries and permissions.
- Confirmation for writes.
- Audit logs.
- What you should build
- Add a read-only tool and a confirmed write tool to your application.
- Ready when
- You can expose a bounded tool and gate a write action behind confirmation.
- Common mistake
- Giving the model direct database access without a sandboxed interface.
- Acceptance checks
- Expose a bounded tool and gate a write action behind confirmation.
- Related resources
- OpenAI function calling guide — Tool definitions, schemas and call patterns
Stage 7: Evaluate behaviour and failure cases
Build an evaluation set, measure outputs and track regressions over changes.
Without evaluation, model changes that look fine locally break behaviour users depend on.
- What you learn
- Evaluation set.
- Metrics and rubrics.
- Regression tracking.
- What you should build
- Create an evaluation set and a script that scores outputs before each deploy.
- Ready when
- You can run an evaluation script and report a score that blocks a bad deploy.
- Common mistake
- Shipping model or prompt changes with no evaluation baseline.
- Acceptance checks
- Run an evaluation script and report a score that blocks a bad deploy.
- Related resources
- LangSmith evaluation guide — Datasets, evaluators and regression testing
Stage 8: Deploy, observe and recover
Add authentication, observability, usage limits and rollback behaviour.
An AI application in production fails differently than a CRUD app; you need recovery, not just uptime.
- What you learn
- Authentication and authorization.
- Metrics, logs and traces.
- Usage limits and quotas.
- Rollback and recovery.
- What you should build
- Deploy the application with auth, metrics, a usage cap and a documented rollback.
- Ready when
- You can deploy with auth, observability, a usage limit and a rollback procedure.
- Common mistake
- Deploying without rate limits or a rollback plan for model-side failures.
- Acceptance checks
- Deploy with auth, observability, a usage limit and a rollback procedure.
- Related resources
- OpenTelemetry documentation — Traces, metrics and logs for production services
From roadmap to production
Build production AI Developer systems with instructor feedback
You have the framework. The View the AI Developer syllabus adds what self-study cannot: live instruction, instructor-reviewed labs, production deployment drills and a capstone that proves you can ship and operate — not just understand.
Fees, schedules and enrolment details are on the course page. No placement, salary or outcome is guaranteed.
Capstone
Build a documentation and account-support app
Documentation search plus a read-only account-status tool. Include authentication, grounded responses, isolated sessions, evaluations and operational recovery.
Training alignment
How this roadmap aligns with SCAI's AI Developer course
This roadmap is free and self-paced. SCAI's AI Developer course covers application construction, model integration, retrieval and deployment with live instruction and guided labs.
The course adds what the roadmap cannot: instructor review of your API design, evaluation datasets and deployment strategy, plus simulated dependency failures for recovery practice. If you prefer independent study, this roadmap gives you the full framework.
What to read next
What to read next
For model techniques — retrieval, prompting, evaluation and adaptation — see the Generative AI roadmap. For controlled agent patterns with tool permissions and failure recovery, see the Agentic AI roadmap. For operating your application in production, see the LLMOps roadmap.
Related learning
- Continue to the Generative AI roadmapTo deepen LLM, prompt, embedding and RAG fundamentals.
- Continue to the Agentic AI roadmapWhen you are ready for tool calling, state and multi-agent patterns.
- Continue to the AI Engineer roadmapFor the broader engineering track spanning DL, serving and production.
- Compare the AI Developer and Agentic AI coursesDecide between application-first AI and production agent engineering.
- Compare AI Developer and AI Engineer pathsUnderstand where software-engineering-first AI diverges from the broad engineering track.
- Compare AI Developer and Data Scientist pathsApplication engineering versus analysis and modelling.
FAQ
AI Developer Roadmap — Frequently Asked Questions
Direct answers for software engineers building model-powered applications.
Do I need ML theory to be an AI developer?
No. You need software engineering plus model API integration, output validation and evaluation.
What is the difference between an AI developer and an AI engineer?
An AI developer builds applications that call models; an AI engineer also selects, adapts and integrates models into systems.
Should I start with RAG or tools?
Start with neither. Build a tested API first, add retrieval only when the task needs external knowledge, and add tools only when actions are required.
How do I evaluate an AI application?
Build an evaluation set of expected behaviours, score outputs with rubrics or model-based judges, and block deploys on regressions.
What fails first in production AI apps?
Model-side errors, cost spikes and unbounded tool actions. Auth, rate limits, observability and rollback mitigate all three.