Data Science Roadmap
Turn a question and dataset into a reproducible analysis and a defensible decision.
Start with a question that data can help answer. Learn Python, SQL, statistics and exploratory analysis, then practise experimental reasoning and predictive modelling. Evaluate uncertainty and data leakage before presenting results. A useful data-science project explains its assumptions, validation and business implications; it does not always need a deployed model.
Quick answer
What is the right roadmap to learn data science?
Start with a question that data can help answer. Learn Python, SQL, statistics and exploratory analysis, then practise experimental reasoning and predictive modelling. Evaluate uncertainty and data leakage before presenting results. A useful data-science project explains its assumptions, validation and business implications; it does not always need a deployed model.
Sources and methodology · This roadmap is reviewed when production practices, tools or platform patterns materially change.
Stages
9
Last reviewed
16 September 2026
Stage 1: Frame the question and decision
Translate a business question into a measurable data question and a decision it would inform.
A precise question prevents analysis that is technically correct but useless to the decision maker.
- What you learn
- Question framing.
- Decision mapping.
- Metric definition.
- Stakeholders and constraints.
- What you should build
- Write a question brief with the decision, metric and stakeholders.
- Ready when
- You can write a question brief that names the decision, metric and stakeholders.
- Common mistake
- Starting to model before agreeing on the question and the decision it informs.
- Acceptance checks
- Write a question brief that names the decision, metric and stakeholders.
- Related resources
- Google ML Crash Course — Problem framing and ML metrics
Stage 2: Python, SQL and reproducible analysis
Use Python and SQL to read, transform and summarize data in a reproducible workflow.
Reproducibility is what separates analysis from one-off notebooks nobody can rerun.
- What you learn
- Python for analysis.
- SQL queries.
- Reproducible pipelines.
- Environment management.
- What you should build
- Build a reproducible analysis pipeline that reads from SQL and outputs a summary report.
- Ready when
- You can rerun the pipeline from a clean environment and get the same summary.
- Common mistake
- Hardcoding file paths and intermediate results so the analysis cannot be rerun.
- Acceptance checks
- Rerun the pipeline from a clean environment and get the same summary.
- Related resources
- pandas introductory tutorials — Read, transform and summarize data
- PostgreSQL tutorial — SQL queries and joins
Stage 3: Clean and validate data
Detect missing values, outliers, duplicates and schema issues without destroying signal.
Cleaning decisions are analysis decisions; silent drops bias results.
- What you learn
- Missing values.
- Outliers.
- Schema validation.
- Duplicates.
- What you should build
- Produce a data-quality report with row counts and reasons for each cleaning step.
- Ready when
- You can produce a data-quality report that justifies each cleaning step.
- Common mistake
- Dropping missing values or outliers without recording why.
- Acceptance checks
- Produce a data-quality report that justifies each cleaning step.
- Related resources
- pandas documentation — Missing data handling
Stage 4: Explore patterns and communicate uncertainty
Explore distributions and relationships, and communicate uncertainty honestly.
Exploration shapes hypotheses; overconfident summaries mislead decisions.
- What you learn
- Distributions.
- Relationships and correlations.
- Uncertainty communication.
- Effective plots.
- What you should build
- Create an exploratory analysis with plots and a summary that states uncertainty.
- Ready when
- You can present an exploratory analysis that shows uncertainty, not just point estimates.
- Common mistake
- Presenting a single summary statistic without distribution or uncertainty.
- Acceptance checks
- Present an exploratory analysis that shows uncertainty, not just point estimates.
- Related resources
- Plotly Python documentation — Interactive exploratory plots
Stage 5: Statistics and estimation
Use estimation, confidence intervals and hypothesis tests appropriate to the data.
Estimates without intervals are opinions; tests without assumptions are noise.
- What you learn
- Point estimation.
- Confidence intervals.
- Hypothesis tests.
- Assumptions and diagnostics.
- What you should build
- Report an estimate with a confidence interval and the assumptions behind it.
- Ready when
- You can report an estimate with a confidence interval and state its assumptions.
- Common mistake
- Reporting a point estimate with no interval or assumption check.
- Acceptance checks
- Report an estimate with a confidence interval and state its assumptions.
- Related resources
- scipy statistics documentation — Distributions, tests and intervals
Stage 6: Experiments and causal reasoning
Design experiments and reason about causality rather than correlation alone.
Decisions require causal claims; correlations do not justify action.
- What you learn
- A/B testing.
- Pre-registration.
- Causal reasoning.
- Confounders and controls.
- What you should build
- Design an A/B test with a pre-registered hypothesis and sample size.
- Ready when
- You can design an A/B test with a pre-registered hypothesis and a sample-size calculation.
- Common mistake
- Claiming causation from observational data without controls or a design.
- Acceptance checks
- Design an A/B test with a pre-registered hypothesis and a sample-size calculation.
- Related resources
- Google ML Crash Course — Fairness, confounding and experiment design
Stage 7: Build predictive baselines
Train simple predictive models with a clear baseline and a leakage-free evaluation.
A predictive model is only useful if it beats a baseline on valid evaluation data.
- What you learn
- Features and labels.
- Baselines.
- Splits and leakage.
- Metrics.
- What you should build
- Train a model and compare it with a naive baseline on held-out data.
- Ready when
- You can train a model that beats a naive baseline and explain its errors.
- Common mistake
- Reporting accuracy without checking the baseline or class balance.
- Acceptance checks
- Train a model that beats a naive baseline and explain its errors.
- Related resources
- scikit-learn common pitfalls — Data leakage and evaluation
Stage 8: Validate models and investigate errors
Validate models honestly and investigate where and why they fail.
A model you cannot challenge is a model you cannot trust.
- What you learn
- Validation strategies.
- Error analysis.
- Segment-level evaluation.
- Fairness checks.
- What you should build
- Produce an error analysis that identifies the segments where the model fails worst.
- Ready when
- You can produce an error analysis identifying the worst-performing segments.
- Common mistake
- Reporting a single metric without segment-level error analysis.
- Acceptance checks
- Produce an error analysis identifying the worst-performing segments.
- Related resources
- scikit-learn user guide — Validation and metrics
Stage 9: Communicate results and deliver the analysis
Present results with assumptions, limitations and business implications.
An analysis that cannot be understood or acted on has no value.
- What you learn
- Report writing.
- Decision-focused visuals.
- Limitations and assumptions.
- Recommended action.
- What you should build
- Deliver a report with the question, method, results, limitations and a recommended action.
- Ready when
- You can deliver a report that a decision maker can act on, with limitations stated.
- Common mistake
- Delivering a notebook instead of a report aimed at the decision maker.
- Acceptance checks
- Deliver a report that a decision maker can act on, with limitations stated.
- Related resources
- Plotly Python documentation — Decision-focused dashboards and plots
Stage 1: Frame the question and decision
Translate a business question into a measurable data question and a decision it would inform.
A precise question prevents analysis that is technically correct but useless to the decision maker.
- What you learn
- Question framing.
- Decision mapping.
- Metric definition.
- Stakeholders and constraints.
- What you should build
- Write a question brief with the decision, metric and stakeholders.
- Ready when
- You can write a question brief that names the decision, metric and stakeholders.
- Common mistake
- Starting to model before agreeing on the question and the decision it informs.
- Acceptance checks
- Write a question brief that names the decision, metric and stakeholders.
- Related resources
- Google ML Crash Course — Problem framing and ML metrics
Stage 2: Python, SQL and reproducible analysis
Use Python and SQL to read, transform and summarize data in a reproducible workflow.
Reproducibility is what separates analysis from one-off notebooks nobody can rerun.
- What you learn
- Python for analysis.
- SQL queries.
- Reproducible pipelines.
- Environment management.
- What you should build
- Build a reproducible analysis pipeline that reads from SQL and outputs a summary report.
- Ready when
- You can rerun the pipeline from a clean environment and get the same summary.
- Common mistake
- Hardcoding file paths and intermediate results so the analysis cannot be rerun.
- Acceptance checks
- Rerun the pipeline from a clean environment and get the same summary.
- Related resources
- pandas introductory tutorials — Read, transform and summarize data
- PostgreSQL tutorial — SQL queries and joins
Stage 3: Clean and validate data
Detect missing values, outliers, duplicates and schema issues without destroying signal.
Cleaning decisions are analysis decisions; silent drops bias results.
- What you learn
- Missing values.
- Outliers.
- Schema validation.
- Duplicates.
- What you should build
- Produce a data-quality report with row counts and reasons for each cleaning step.
- Ready when
- You can produce a data-quality report that justifies each cleaning step.
- Common mistake
- Dropping missing values or outliers without recording why.
- Acceptance checks
- Produce a data-quality report that justifies each cleaning step.
- Related resources
- pandas documentation — Missing data handling
Stage 4: Explore patterns and communicate uncertainty
Explore distributions and relationships, and communicate uncertainty honestly.
Exploration shapes hypotheses; overconfident summaries mislead decisions.
- What you learn
- Distributions.
- Relationships and correlations.
- Uncertainty communication.
- Effective plots.
- What you should build
- Create an exploratory analysis with plots and a summary that states uncertainty.
- Ready when
- You can present an exploratory analysis that shows uncertainty, not just point estimates.
- Common mistake
- Presenting a single summary statistic without distribution or uncertainty.
- Acceptance checks
- Present an exploratory analysis that shows uncertainty, not just point estimates.
- Related resources
- Plotly Python documentation — Interactive exploratory plots
Stage 5: Statistics and estimation
Use estimation, confidence intervals and hypothesis tests appropriate to the data.
Estimates without intervals are opinions; tests without assumptions are noise.
- What you learn
- Point estimation.
- Confidence intervals.
- Hypothesis tests.
- Assumptions and diagnostics.
- What you should build
- Report an estimate with a confidence interval and the assumptions behind it.
- Ready when
- You can report an estimate with a confidence interval and state its assumptions.
- Common mistake
- Reporting a point estimate with no interval or assumption check.
- Acceptance checks
- Report an estimate with a confidence interval and state its assumptions.
- Related resources
- scipy statistics documentation — Distributions, tests and intervals
Stage 6: Experiments and causal reasoning
Design experiments and reason about causality rather than correlation alone.
Decisions require causal claims; correlations do not justify action.
- What you learn
- A/B testing.
- Pre-registration.
- Causal reasoning.
- Confounders and controls.
- What you should build
- Design an A/B test with a pre-registered hypothesis and sample size.
- Ready when
- You can design an A/B test with a pre-registered hypothesis and a sample-size calculation.
- Common mistake
- Claiming causation from observational data without controls or a design.
- Acceptance checks
- Design an A/B test with a pre-registered hypothesis and a sample-size calculation.
- Related resources
- Google ML Crash Course — Fairness, confounding and experiment design
Stage 7: Build predictive baselines
Train simple predictive models with a clear baseline and a leakage-free evaluation.
A predictive model is only useful if it beats a baseline on valid evaluation data.
- What you learn
- Features and labels.
- Baselines.
- Splits and leakage.
- Metrics.
- What you should build
- Train a model and compare it with a naive baseline on held-out data.
- Ready when
- You can train a model that beats a naive baseline and explain its errors.
- Common mistake
- Reporting accuracy without checking the baseline or class balance.
- Acceptance checks
- Train a model that beats a naive baseline and explain its errors.
- Related resources
- scikit-learn common pitfalls — Data leakage and evaluation
Stage 8: Validate models and investigate errors
Validate models honestly and investigate where and why they fail.
A model you cannot challenge is a model you cannot trust.
- What you learn
- Validation strategies.
- Error analysis.
- Segment-level evaluation.
- Fairness checks.
- What you should build
- Produce an error analysis that identifies the segments where the model fails worst.
- Ready when
- You can produce an error analysis identifying the worst-performing segments.
- Common mistake
- Reporting a single metric without segment-level error analysis.
- Acceptance checks
- Produce an error analysis identifying the worst-performing segments.
- Related resources
- scikit-learn user guide — Validation and metrics
Stage 9: Communicate results and deliver the analysis
Present results with assumptions, limitations and business implications.
An analysis that cannot be understood or acted on has no value.
- What you learn
- Report writing.
- Decision-focused visuals.
- Limitations and assumptions.
- Recommended action.
- What you should build
- Deliver a report with the question, method, results, limitations and a recommended action.
- Ready when
- You can deliver a report that a decision maker can act on, with limitations stated.
- Common mistake
- Delivering a notebook instead of a report aimed at the decision maker.
- Acceptance checks
- Deliver a report that a decision maker can act on, with limitations stated.
- Related resources
- Plotly Python documentation — Decision-focused dashboards and plots
From roadmap to production
Build production Data Science systems with instructor feedback
You have the framework. The View the Data Science syllabus adds what self-study cannot: live instruction, instructor-reviewed labs, production deployment drills and a capstone that proves you can ship and operate — not just understand.
Fees, schedules and enrolment details are on the course page. No placement, salary or outcome is guaranteed.
Capstone
Analyse retention and propose an experiment
Analyse a retention dataset, build a predictive baseline, produce an error analysis and propose an A/B test with a pre-registered hypothesis and sample size. Deliver a report a decision maker can act on, with limitations stated.
Training alignment
How this roadmap aligns with SCAI's Data Science course
This roadmap is free and self-paced. SCAI's Data Science course covers statistics, SQL, analysis and modelling with live instruction and guided projects.
The course adds what the roadmap cannot: instructor review of your analysis decisions, experimental design and communication, plus structured progression through statistics and ML fundamentals. If you prefer independent study, this roadmap gives you the full framework.
What to read next
What to read next
For repeatable predictive systems with training and serving pipelines, see the ML Engineer roadmap. For programming prerequisites, see the AI Roadmap for Beginners. For an optional AI-assisted analysis extension, see the Generative AI roadmap.
Related learning
- Continue to the ML Engineer roadmapWhen you want model development over analysis.
- Continue to the Generative AI roadmapTo add LLMs and RAG as modern extensions.
- Continue to the MLOps roadmapTo operationalize models you build.
- Compare AI Developer and Data Scientist pathsApplication engineering versus data analysis and modelling.
FAQ
Data Science Roadmap — Frequently Asked Questions
Direct answers for aspiring data scientists.
Do I need a maths degree to learn data science?
No. You need applied statistics and probability taught through examples, not an advanced maths curriculum before your first analysis.
Is SQL still relevant for data science?
Yes. Most business data lives in databases; SQL is the fastest way to read and join it.
Does every data-science project need a deployed model?
No. A defensible analysis with assumptions and limitations is often more useful than a deployed model.
What is the difference between data science and ML engineering?
Data science starts with a question and a decision; ML engineering starts with a model and a system.
How much statistics do I really need?
Enough to report estimates with intervals, design experiments and avoid claiming causation from correlation.