Build the core skills · 07 / 15
Make answers traceable to evidence
Retrieval-augmented generation, or RAG, supplies relevant documents to a model before it answers. Use it when the answer depends on a document collection that must be traceable or updated. First test whether simple search or placing a few short documents directly in context already solves the task.
Build a small, inspectable version
- Choose 20–50 short public or synthetic documents in one domain. Record where each came from and whether you can use it.
- Write 20 questions before tuning search. Include answerable, ambiguous and unanswerable questions. Keep a separate set for final testing.
- Start with keyword search; then add embeddings if the baseline misses meaningfully different wording. Keep stable document and chunk IDs.
- Inspect the actual text returned for each question. Fix parsing and chunking before buying a more complex retrieval system.
- Ask the model to answer from the retrieved evidence and return source IDs. Check that the cited passage actually supports the claim.
- If evidence is missing or contradictory, return a clear limitation or ask for clarification. Never invent a citation.
What to measure
For questions with labeled sources, recall@k asks whether the expected source appears in the first k results. Separately judge whether the final answer is correct, supported and useful. Good retrieval does not automatically produce a good answer. Test exact identifiers as well as paraphrases; hybrid search can help when each search method has different strengths.
Common failure → first investigation
- Correct document absent → inspect ingestion, permissions and query matching.
- Correct document present but answer wrong → inspect context length, contradictions and instructions.
- Citation exists but is irrelevant → verify support at the passage level.
- Stale answers → version the corpus and define when indexes are refreshed.
Watch Full Stack LLM Bootcamp: Augmented Language Models ↗ for the concept. If you want a guided exercise, consider DeepLearning.AI: Building and Evaluating Advanced RAG ↗ after your baseline works. Its named framework APIs may have changed.
You can show retrieved passages for a failure, compare search variants on fixed questions and demonstrate a refusal when the documents do not support an answer.
Your lesson resources
Download these files to follow along and put the lesson into practice.