FREE APPLIED LESSON
Free GenAI Lesson: Build and Evaluate a Grounded Answering System
An answer can sound convincing while using the wrong document or claiming something the source never says. In this lesson you will define a small task, inspect a baseline, add evidence and test when the system should answer or say it does not know.
In this free lesson, you work through a small grounded answering task using synthetic policy documents. You define the task, inspect a baseline, add retrieval and evidence, test six failure modes, diagnose a mistake and revise it, and check your answers against worked solutions. The lesson separates retrieval failure (the right evidence was not found) from generation failure (the right evidence was found but used incorrectly). No paid API or GPU is required — the walkthrough uses preselected passages and illustrative outputs.
Lesson contract
Prerequisites: Basic Python knowledge and understanding of what an LLM API call does. You do not need to have built a RAG system before.
Learning outcomes: By the end of this lesson, you can define a grounded answering task with an answer contract, inspect a baseline for failure cases, add evidence and test whether the system uses it correctly, distinguish retrieval failure from generation failure, diagnose a specific mistake and make a targeted revision, and evaluate when a system should abstain rather than answer.
Required materials: This lesson is a reading and reasoning exercise. No code execution is required. If you want to run the optional code version, you will need a Python environment and an API key for an LLM service — that is disclosed separately and is not required to complete the lesson.
What you will complete: A set of test cases with expected outcomes, a diagnosis of one failure, a targeted revision, and a self-check against worked answers.
The task and what a correct answer requires
You are building a Q&A system for a fictional company's service policies. The corpus contains four documents:
DOC-001 (v1.0, 2026-01-15): 'Return Policy — Items can be returned within 30 days of purchase with original receipt. Refunds are processed to the original payment method within 5–7 business days.'
DOC-002 (v1.0, 2026-01-15): 'Shipping Policy — Standard shipping takes 3–5 business days. Express shipping takes 1–2 business days. Free shipping is available for orders above ₹2,000.'
DOC-003 (v2.0, 2026-06-01): 'Return Policy Update — Effective June 2026, the return window is extended to 45 days for all customers. Refunds now process within 3–5 business days. Items without receipt are eligible for store credit only.'
DOC-004 (v1.0, 2026-01-15): 'Warranty Policy — Electronic items carry a 12-month manufacturer warranty. Warranty claims require proof of purchase and the original product packaging.'
The answer contract: every answer must include the answer text, supporting document IDs, and an explicit insufficient-evidence outcome when the system cannot find supporting evidence. Adding a citation label does not prove the answer is supported — the cited document must actually contain the claim.
Establish a baseline
An illustrative ungrounded answer to the question 'What is the return window?' might be: 'The return window is 30 days from the date of purchase.' This is accurate for DOC-001 but outdated — DOC-003 extends it to 45 days. Without retrieval, the model may produce the older answer because it was trained on or has seen more data about 30-day policies.
If a baseline supplies the whole small corpus in the prompt, it may work well — with only four documents, the model can read all of them and find the current policy. This is a valid baseline. Acknowledge that retrieval is not automatically better on tiny inputs — its value appears with larger corpora where the model cannot read everything at once. This walkthrough uses preselected passages rather than a measured retrieval result.
Add retrieval and evidence
The retrieval step selects documents or passages relevant to the question and provides them as evidence to the model. For this walkthrough, passages are preselected — this is a guided exercise, not a measured retrieval run. In a real system, a retriever would use embeddings and vector search to find relevant passages automatically.
For the question 'What is the return window?', the relevant passage is from DOC-003: 'Effective June 2026, the return window is extended to 45 days for all customers.' The system provides this passage as evidence and asks the model to answer based only on the provided evidence. The answer should be: 'The return window is 45 days for all customers (DOC-003).' The citation is accurate because DOC-003 actually contains this claim.
Test different failure modes
These six cases cover the most common failure patterns in grounded answering. Development cases are used for iteration; held-out cases test generalisation.
| Case | Question | Expected evidence or abstention | Likely failure category |
|---|---|---|---|
| 1. Directly answerable (dev) | What is the return window? | DOC-003 — 45 days | Model may use outdated DOC-001 (30 days) |
| 2. Paraphrased (dev) | How many days do I have to return something? | DOC-003 — 45 days | Retrieval may miss the paraphrase if embeddings differ |
| 3. Ambiguous (dev) | How long does shipping take? | DOC-002 — standard 3–5 days, express 1–2 days | Model may give one number without specifying which service |
| 4. Missing fact (held-out) | What is the restocking fee? | Insufficient evidence — no document mentions restocking fees | Model may hallucinate a fee instead of abstaining |
| 5. Outdated document (held-out) | How long do refunds take to process? | DOC-003 — 3–5 business days (v2.0 supersedes DOC-001) | Model may cite DOC-001 (5–7 days) without checking version dates |
| 6. Partially supported (held-out) | Can I return an item without a receipt? | DOC-003 — store credit only, without receipt | Model may say 'no' based on DOC-001 which does not mention this case |
Diagnose and revise one mistake
Consider Case 5: 'How long do refunds take to process?' The system returns: 'Refunds are processed within 5–7 business days (DOC-001).' The failure: DOC-003 updates this to 3–5 business days, but the system cited the older document. The failure category is outdated-document retrieval — the retriever found DOC-001 but not DOC-003, or the model did not check version dates.
A targeted change: add date filtering to the retrieval step — when multiple versions of a policy exist, retrieve only the latest version. Or: provide both documents and instruct the model to use the most recent version. The reason for the change: the failure is caused by the system not recognising that DOC-003 supersedes DOC-001. After making this change, retest Cases 1, 5 and 6 — all involve the return policy and could be affected by the version handling.
Do not claim that one prompt change eliminates hallucination. The fix addresses one specific failure: outdated-document retrieval. Other failure modes (missing-fact hallucination, paraphrase mismatch) require different fixes.
Worked answers and self-check
Self-check question 1: Identify an unsupported claim. Look at the baseline answer 'The return window is 30 days.' Which document supports this, and which document supersedes it? Answer: DOC-001 supports 30 days, but DOC-003 supersedes it with 45 days. The baseline is technically supported by one document but is outdated.
Self-check question 2: Select supporting evidence. For the question 'Can I return an item without a receipt?', which document provides the answer? Answer: DOC-003 states 'Items without receipt are eligible for store credit only.' DOC-001 does not mention this case.
Self-check question 3: Decide when to abstain. For the question 'What is the restocking fee?', what should the system do? Answer: Return an insufficient-evidence outcome. No document mentions restocking fees. The system should say it cannot find supporting evidence rather than inventing a number.
Self-check question 4: Choose the next useful change. After fixing the outdated-document issue, what is the next most impactful improvement? Answer: Add abstention handling for missing facts (Case 4). The system should be explicitly instructed to say 'I cannot find information about this in the available documents' when retrieval returns no relevant passages, rather than generating an answer from general knowledge.
What this lesson leaves for larger systems
This lesson uses four documents and preselected passages. Real systems face additional challenges: more diverse data (hundreds or thousands of documents with varying quality and format), scale (retrieval must be fast and accurate across a large corpus), permissions (access control over which documents each user can query), broader evaluation (hundreds of test cases with category labels and regression tracking), and operations (serving, monitoring, drift detection and cost control).
These challenges are covered in the roadmap and the production AI projects guide. This lesson is one slice of GenAI engineering — it teaches the evaluation mindset that the rest of the system depends on.
Continue with guided projects
You can extend these tests on your own. If you want to build more complex systems with guided project reviews, explore the live Generative AI course.
The course adds instructor review of your evaluation datasets, retrieval design and adaptation experiments.