ROADMAP

Data Science Roadmap

Turn a question and dataset into a reproducible analysis and a defensible decision.

Start with a question that data can help answer. Learn Python, SQL, statistics and exploratory analysis, then practise experimental reasoning and predictive modelling. Evaluate uncertainty and data leakage before presenting results. A useful data-science project explains its assumptions, validation and business implications; it does not always need a deployed model.

For:Aspiring data scientists, analysts and career switchers.

Quick answer

What is the right roadmap to learn data science?

Start with a question that data can help answer. Learn Python, SQL, statistics and exploratory analysis, then practise experimental reasoning and predictive modelling. Evaluate uncertainty and data leakage before presenting results. A useful data-science project explains its assumptions, validation and business implications; it does not always need a deployed model.

Written byAshutosh· AI InstructorVerified byVivek· AIOps and Generative AI InstructorPublishedUpdated

Sources and methodology · This roadmap is reviewed when production practices, tools or platform patterns materially change.

Stages

9

Last reviewed

16 September 2026

Stage 1: Frame the question and decision

Translate a business question into a measurable data question and a decision it would inform.

A precise question prevents analysis that is technically correct but useless to the decision maker.

What you learn
  • Question framing.
  • Decision mapping.
  • Metric definition.
  • Stakeholders and constraints.
What you should build
Write a question brief with the decision, metric and stakeholders.
Ready when
You can write a question brief that names the decision, metric and stakeholders.
Common mistake
Starting to model before agreeing on the question and the decision it informs.
Acceptance checks
  • Write a question brief that names the decision, metric and stakeholders.
Related resources

From roadmap to production

Build production Data Science systems with instructor feedback

You have the framework. The View the Data Science syllabus adds what self-study cannot: live instruction, instructor-reviewed labs, production deployment drills and a capstone that proves you can ship and operate — not just understand.

Build the core project from this roadmap with instructor review
Debug production failure modes hands-on with guided feedback
Produce a reviewed portfolio artifact by the end of the track

Fees, schedules and enrolment details are on the course page. No placement, salary or outcome is guaranteed.

Capstone

Analyse retention and propose an experiment

Analyse a retention dataset, build a predictive baseline, produce an error analysis and propose an A/B test with a pre-registered hypothesis and sample size. Deliver a report a decision maker can act on, with limitations stated.

Training alignment

How this roadmap aligns with SCAI's Data Science course

This roadmap is free and self-paced. SCAI's Data Science course covers statistics, SQL, analysis and modelling with live instruction and guided projects.

The course adds what the roadmap cannot: instructor review of your analysis decisions, experimental design and communication, plus structured progression through statistics and ML fundamentals. If you prefer independent study, this roadmap gives you the full framework.

What to read next

What to read next

For repeatable predictive systems with training and serving pipelines, see the ML Engineer roadmap. For programming prerequisites, see the AI Roadmap for Beginners. For an optional AI-assisted analysis extension, see the Generative AI roadmap.

FAQ

Data Science Roadmap — Frequently Asked Questions

Direct answers for aspiring data scientists.

Do I need a maths degree to learn data science?

No. You need applied statistics and probability taught through examples, not an advanced maths curriculum before your first analysis.

Is SQL still relevant for data science?

Yes. Most business data lives in databases; SQL is the fastest way to read and join it.

Does every data-science project need a deployed model?

No. A defensible analysis with assumptions and limitations is often more useful than a deployed model.

What is the difference between data science and ML engineering?

Data science starts with a question and a decision; ML engineering starts with a model and a system.

How much statistics do I really need?

Enough to report estimates with intervals, design experiments and avoid claiming causation from correlation.