CURRICULUM EVALUATION GUIDE
What Should a Generative AI Course Actually Teach?
A useful Generative AI course should help you explain how a model behaves, choose a suitable approach, prepare data, test the result and improve a system when it fails. A list containing transformers, RAG and fine-tuning does not show whether those skills are taught. Look for the decisions learners make, the work they submit and the feedback they receive.
A useful Generative AI course should help you explain how a model behaves, choose a suitable approach, prepare data, test the result and improve a system when it fails. A syllabus listing transformers, RAG and fine-tuning does not show whether those skills are actually taught. Look for the decisions learners make, the work they submit and the feedback they receive — not just the tools covered.
Start with the work you want to do
Before evaluating a syllabus, identify what you actually want to build. Different learning goals require different depths of understanding. Someone who wants to use AI tools effectively in everyday work needs a different programme from someone who wants to understand and improve model behaviour. A backend developer adding an AI feature to an existing product needs retrieval, evaluation and integration skills. An ML practitioner investigating whether adaptation improves a defined task needs deeper data preparation and fine-tuning knowledge.
Four common learning goals exist, and they require different depths. First, using AI tools effectively in everyday work — this is awareness-level, not engineering. Second, building AI-powered applications — this requires API integration, retrieval, structured outputs and evaluation. Third, understanding and improving model and data behaviour — this requires foundations, adaptation, evaluation and failure analysis. Fourth, deploying and operating AI systems reliably — this is the domain of MLOps and LLMOps, not a general GenAI course.
This guide focuses on engineering-oriented GenAI learning — goals two and three. If you are comparing broader course paths, the existing course-path comparison covers AI Developer, Generative AI and LLMOps side by side.
Four learning goals and what each requires
Foundations should explain model behaviour
A serious syllabus does not just name transformers — it explains what the model learns from data and why that matters for engineering decisions. Without this understanding, you cannot diagnose why a model produces a particular output, when to trust it, or when to adapt it.
What the model learns from data
Models learn representations — ways of encoding input text, images or audio into numerical forms that capture meaningful relationships. Tokenisation determines how text is split into units the model processes. Training objectives define what the model is optimised to predict. Model parameters are the learned weights that shape behaviour, and generalisation is the ability to perform well on inputs the model has not seen before.
The training dataset affects behaviour because the model learns patterns from it. If the data over-represents certain styles, domains or perspectives, the model will reflect that bias. Lower training loss does not automatically mean a better application — a model can fit the training data more closely while becoming worse at the specific task you care about. A good course teaches you to reason about this gap, not just to report a loss number.
What changes between training and inference
Training changes the model's parameters. Inference uses those fixed parameters with a new input and context to produce output. Sampling strategies — temperature, top-p, top-k — affect which tokens are selected and therefore what the output looks like. Context window limits determine how much text the model can consider at once. Memory and compute constraints shape practical choices: a model that works in a notebook may be too slow or expensive to serve in production.
A small teaching experiment can reveal these behaviours without requiring frontier-model training. A good course uses such experiments to build intuition, not just to demonstrate that a tool runs.
Where language, vision and multimodal models differ
Generative AI extends beyond text chat. A syllabus may cover language generation, vision-language understanding (models that read images and describe them), image or audio generation, and multimodal document processing (extracting information from documents that contain both text and images). A programme can have a legitimate focus without teaching every modality — but it should make its scope explicit. If a course claims to cover multimodal AI, check whether it teaches evaluation for visual grounding and abstention, not just whether it runs a vision-language model.
A serious curriculum teaches data and adaptation decisions
Data preparation is an engineering responsibility, not a preprocessing afterthought. A good course teaches you to evaluate data suitability, provenance, example quality, train-development-test separation, leakage, representativeness and error categories. Without these skills, you cannot honestly evaluate whether an adaptation improved the model.
Choosing between prompting, retrieval and fine-tuning
Use this table to check whether a syllabus teaches the decision, not just the technique.
| Problem observed | Approach worth testing | Evidence needed | Limitation to remember |
|---|---|---|---|
| Model returns inconsistent output format | Prompting with structured outputs | Schema adherence rate, task pass rate | Does not add knowledge the model lacks |
| Answer depends on private or current knowledge | RAG with citations | Retrieval recall, citation correctness, grounded answer quality | RAG does not guarantee correctness if retrieval fails |
| Model repeatedly fails a domain-specific behaviour | LoRA or QLoRA fine-tuning | Held-out task accuracy, regression suite results | Fine-tuning does not install a factual database |
| Poor quality with an unclear cause | Diagnose first: is it retrieval, generation or task definition? | Per-layer trace and failure category analysis | A complex approach is not automatically better |
| Task can be solved with rules or deterministic code | Use the simpler method first | Comparison of accuracy, cost and maintainability | GenAI is not always the right tool |
Understanding what LoRA and QLoRA change
There is a difference between calling a PEFT library and understanding what it does. LoRA (Low-Rank Adaptation) trains a small set of additional parameters while freezing the base model's weights. QLoRA extends this with quantisation to reduce memory requirements. A good course teaches which parameters are trained, the memory trade-offs involved, what training data and objective are used, how to detect overfitting, how to run held-out evaluation, and ultimately whether the adaptation was useful.
The key question is not 'did you fine-tune a model?' but 'did the fine-tuned model perform better than the baseline on the task you care about, without regressing on safety or general behaviour?' If a syllabus cannot answer how learners evaluate this, the fine-tuning module is surface-level.
Evaluation should shape the project from the start
Evaluation is not a final step — it should be designed before the first tool is selected. Define the task, establish a baseline, separate development from held-out evaluation, inspect failure categories, retest for regressions after every change, and include cost or latency when relevant.
Consider an illustrative structured-extraction example. A model is asked to extract fields from a purchase order and return JSON. The output is valid JSON — it passes schema validation. But the extracted total uses the wrong currency, or confuses a line item price with the order total. Schema validity misses this mistake because the structure is correct but the values are wrong. A reviewer would ask: 'Which fields were extracted incorrectly, and what test would catch this?' The test to add is field-level value comparison, not just schema validation. This example is illustrative — it has not been executed as a measured benchmark.
What a meaningful project submission and review look like
A meaningful project submission includes: a problem statement, a data boundary (what data is used and where it came from), a baseline (what a simpler approach achieves), design choices (why this approach was selected), reproducible work (someone else can run it), evaluation cases (specific tests with expected outcomes), failed examples (cases where the system does not work) and limitations (what the system does not handle).
An illustrative reviewer comment might be: 'Which held-out examples improved after your change, and what became worse?' To answer this, the learner needs a before-and-after comparison on a held-out set, not just a demo that produces output. This is the difference between a project that teaches tool usage and one that teaches engineering judgement. For complete project briefs, see the Generative AI projects guide.
Use this checklist when comparing a GenAI syllabus
Print or save this table. Use it to evaluate any Generative AI programme before you pay.
| Capability | What the learner should explain or do | Evidence to ask the provider for | Follow-up question when the syllabus is vague |
|---|---|---|---|
| Model foundations | Explain how tokenisation, training objectives and parameters affect output | Lab exercises that inspect model behaviour, not just call APIs | Do learners analyse why a model produces a specific output? |
| Data preparation | Prepare, split and evaluate training data honestly | Data preparation assignments with evaluation criteria | Is data quality taught as an engineering responsibility? |
| Method selection | Choose between prompting, RAG and fine-tuning with evidence | Decision frameworks or comparison exercises | Can learners explain when not to use GenAI? |
| Adaptation | Run LoRA/QLoRA and evaluate against a baseline | Fine-tuning assignment with held-out evaluation and regression check | Do learners prove the adaptation was useful? |
| Multimodal scope | Build and evaluate text-plus-image workflows if covered | Multimodal project with field accuracy and abstention tests | Is multimodal evaluation taught, or just inference? |
| Evaluation | Define a task, build a baseline, run held-out tests, inspect failures | Evaluation sets and failure analysis as graded deliverables | Is evaluation designed before tools are selected? |
| Project review | Submit work and receive feedback on design and evaluation | Description of review process and example feedback | Who reviews, and what do they check? |
| Delivery and resources | Know the time commitment, format, compute access and total cost | Clear schedule, recording policy and cost breakdown | Are additional cloud or API costs disclosed? |
When guided learning is worth considering
Independent learning may work well if you have strong prerequisites, can scope and evaluate your own projects, have access to useful feedback from a community or colleagues, and maintain consistent study habits. Many practitioners build successful GenAI skills through documentation, experiments and open-source projects without paying for a course.
Live guidance may help when you have difficulty choosing a learning sequence, are uncertain whether your project approach is correct, keep making the same technical mistakes, or need structured feedback from someone who can inspect your work. Paying for a course does not fix weak prerequisites or guarantee a career outcome — but it can accelerate progress for someone who is ready and needs direction and review.
If this is the depth you want to compare, review the syllabus and project expectations in SCAI's live Generative AI course. The roadmap provides a free learning sequence if you prefer to start independently.