PROJECT GUIDE

Generative AI Projects: From Demo to Evaluated System

A useful GenAI project answers a defined problem and shows how you tested the result. Choose a project whose data, baseline and failure cases you can explain. The examples below show both what to build and what evidence makes the work worth reviewing.

Cluster
AI Operations
Owner Course
AIOps Course
Updated
Type
Core Guide
Direct Answer

A useful Generative AI project answers a defined problem and shows how you tested the result. The evidence that makes a project credible includes a clear problem statement, a data boundary, a baseline, design decisions, test cases, failed examples, reproduction steps and documented limitations. A working demo alone does not prove engineering ability — the evaluation and failure analysis do.

Choose a project for the skill you want to develop

These are independent learning ideas, not additional SCAI course projects.

ProblemSkill testedUsable dataBaselineEvaluation evidence
Grounded knowledge answersRetrieval, grounding, abstention3–5 policy or FAQ documentsPrompt-only without retrievalCitation accuracy, abstention on missing facts
Multimodal document extractionVision-language, field accuracy, abstentionSynthetic invoices or purchase ordersText extraction plus rulesField-level correctness, schema validity, unreadable handling
Controlled model adaptationLoRA/QLoRA, evaluation, regression100–500 task-specific examplesPrompt-only baseline on same taskHeld-out accuracy, regression suite, comparison report
Domain summarisation with evidenceGeneration quality, grounding, length controlTechnical articles or papersGeneric prompt without source constraintsCoverage, faithfulness, hallucination check
Model-behaviour comparison labTokenisation, context, sampling, structured outputSame input set across configurationsSingle default configurationFailure categorisation, config comparison table

What every credible project should contain

Regardless of which project you choose, a credible GenAI project submission should contain eight elements. A problem and scope statement — what the system does and does not do. A data boundary — what data is used, where it came from and what it represents. A baseline — what a simpler approach achieves. Design decisions — why you chose this approach over alternatives. Tests — specific cases with expected outcomes. Failed examples — cases where the system does not work and why. Reproduction steps — how someone else can run your work. Limitations — what the system does not handle.

A project missing these elements is a demo, not engineering evidence. A demo shows that a tool runs. Evidence shows that you understand when it fails and why.

Worked project — multimodal document intelligence

This walkthrough uses a synthetic purchase-order example. The documents are fictional — created for teaching, not drawn from real transactions. The goal is to show how a project moves from problem definition through evaluation, not to present a measured benchmark.

Define the contract and data boundary

Input types: scanned or digital purchase orders containing supplier name, line items, totals, currency and document date. Expected fields: supplier (string), line items (array of {description, quantity, unit_price, total}), total (number), currency (ISO 4217 code), document date (ISO 8601). Missing or ambiguous fields: some documents may lack a currency field — the system should detect this rather than guessing. Unreadable documents: scans with blurry or overlapping text — the system should abstain rather than fabricating. Data provenance: synthetic documents created for this exercise, not real commercial records.

Establish a baseline

A viable baseline is text extraction plus rules: use OCR or text extraction to get the document text, then apply regular expressions or keyword matching to find totals, dates and supplier names. This baseline works well on clean digital documents with consistent formatting. It fails on scanned documents with variable layouts, handwritten annotations or non-standard field labels. Acknowledge where the baseline applies and where it does not — a fair baseline comparison requires testing on the same document set.

Identify a failure before changing the system

Consider a failure case where the model returns valid JSON with correct structure but wrong values. The input is a purchase order in EUR. The output extracts the total correctly but labels the currency as USD. Another failure: a line item description is extracted as the total because the model confused a description field with a numeric field. A third failure: the model fills in a delivery date that does not appear anywhere in the document.

For each failure, identify the category: wrong currency is a field-level value error. Confusing a line item with the total is a semantic misunderstanding. Filling a field unsupported by the document is a hallucination. Each category requires a different intervention. Schema validity does not catch any of these — the JSON is well-formed in all three cases.

Improve the relevant part

Possible interventions include: better parsing (preprocessing the document to separate fields before model input), different input representation (sending the image directly to a vision-language model instead of OCR text), prompt or schema changes (explicitly asking the model to extract currency from the document and abstain if missing), model choice (switching to a model with better vision-language capability), multimodal input (providing both text and image), or adaptation (fine-tuning on a small set of labelled purchase orders).

The key principle: choose the intervention that addresses the observed failure. If the failure is hallucinated fields, the fix is better abstention design, not fine-tuning. If the failure is semantic misunderstanding, a vision-language model may help. Justify why the chosen intervention should address the specific failure before implementing it.

Evaluate on held-out cases

Define evaluation metrics: field-level correctness (each field compared to ground truth), schema validity (does the output parse), unsupported values (did the model fill a field with no evidence), appropriate abstention (did the system decline on unreadable documents), and resource measurements if actually run (latency, token cost, API calls). Distinguish an evaluation plan from a measured result — if you have not run the evaluation, state that clearly. Do not present planned metrics as measured outcomes.

Present the result and limitations

An evaluation table template should show: metric name, baseline result, improved result, difference, and notes. A failed-example log should show: input description, expected output, actual output, failure category and proposed fix. Reproduction steps should include: environment setup, data access, command to run, and expected output. Scope limitations should state: what document types were tested, what was excluded, and what generalisation claims are not supported. Do not invent improvement percentages — if you ran the evaluation, report the actual numbers. If you did not, label the table as a template.

Two shorter project briefs

Brief A — Grounded answering: Build a Q&A system over 3–5 policy documents. The system must answer with supporting evidence or say it does not know. Cover: source coverage (are all documents indexed), supporting evidence (does the answer cite the right passage), missing-answer cases (does the system abstain when evidence is absent), and retrieval versus generation errors (separate 'could not find the passage' from 'found it but generated incorrectly'). The free evaluation lesson walks through this pattern in detail.

Brief B — Controlled adaptation: Fine-tune a small open-weight model on 100–500 task-specific examples. Cover: data quality (are examples clean and representative), baseline (what does prompt-only achieve), train/development/test separation (no leakage), behavioural objective (what should the model do differently), regressions (did general behaviour degrade), and why adaptation is justified (what evidence shows prompting and retrieval are insufficient). The fine-tuning comparison guide covers the technical decision in depth.

What an instructor or reviewer should check

Use this rubric to review your own work before submitting it. If you cannot answer these questions, the project is not complete.

Review questionWhat to look forAction if the answer is unclear
Is the baseline fair?Baseline tested on the same data as the improved systemRun the baseline and re-measure
Do tests leak answers?No test example appears in training or retrieval dataCheck for overlap between train and test sets
Are improvements supported?Held-out evaluation with before-and-after comparisonRun evaluation on unseen cases
Were regressions checked?General behaviour tested alongside target taskRun a regression suite
Are trade-offs explained?Cost, latency or accuracy trade-offs documentedMeasure and report the trade-off
Can another person reproduce the work?README with setup, data and run instructionsAsk someone else to run it

How to present the work in a portfolio

A practical README structure for a GenAI project: Problem (what the system does), Data (what data is used and where it came from), Baseline (what a simpler approach achieves), Approach (what you built and why), Evaluation (metrics, test cases and results), Reproduction (how to run it), Limitations (what it does not handle).

An illustrative reviewer comment: 'The JSON output is valid, but the currency field is wrong on 3 of 10 test documents. Which documents, and what in the input causes the error?' The revision this requests: add a field-level error log showing which documents have wrong currency, inspect the input to find the pattern, and either fix the extraction or document the limitation. For deployment and operations extensions, see the production AI projects guide.

LIVE COURSE

Build guided projects with review

If you want guided projects with review of your design and evaluation work, see what the live GenAI course asks learners to build and submit.

6 guided projectsMentor reviewsEvaluation criteriaCapstone

Each project includes a defined task, baseline, evaluation set and failure analysis.

Last reviewed: 2026-09-28
Technical review: School of Core AI editorial team