Build the core skills · 07 / 15

Make answers traceable to evidence

AI Engineer Path15 min guide

Retrieval-augmented generation, or RAG, supplies relevant documents to a model before it answers. Use it when the answer depends on a document collection that must be traceable or updated. First test whether simple search or placing a few short documents directly in context already solves the task.

Question
→
Search permitted documents
→
Generate with evidence
→
Answer + source IDs
Check access before retrieval. Keep source identifiers through every step. Evaluate search and answer quality separately.

Build a small, inspectable version

  1. Choose 20–50 short public or synthetic documents in one domain. Record where each came from and whether you can use it.
  2. Write 20 questions before tuning search. Include answerable, ambiguous and unanswerable questions. Keep a separate set for final testing.
  3. Start with keyword search; then add embeddings if the baseline misses meaningfully different wording. Keep stable document and chunk IDs.
  4. Inspect the actual text returned for each question. Fix parsing and chunking before buying a more complex retrieval system.
  5. Ask the model to answer from the retrieved evidence and return source IDs. Check that the cited passage actually supports the claim.
  6. If evidence is missing or contradictory, return a clear limitation or ask for clarification. Never invent a citation.

What to measure

For questions with labeled sources, recall@k asks whether the expected source appears in the first k results. Separately judge whether the final answer is correct, supported and useful. Good retrieval does not automatically produce a good answer. Test exact identifiers as well as paraphrases; hybrid search can help when each search method has different strengths.

Common failure → first investigation

  • Correct document absent → inspect ingestion, permissions and query matching.
  • Correct document present but answer wrong → inspect context length, contradictions and instructions.
  • Citation exists but is irrelevant → verify support at the passage level.
  • Stale answers → version the corpus and define when indexes are refreshed.

Watch Full Stack LLM Bootcamp: Augmented Language Models ↗ for the concept. If you want a guided exercise, consider DeepLearning.AI: Building and Evaluating Advanced RAG ↗ after your baseline works. Its named framework APIs may have changed.

Move on when

You can show retrieved passages for a failure, compare search variants on fixed questions and demonstrate a refusal when the documents do not support an answer.

Your lesson resources

Download these files to follow along and put the lesson into practice.

Saved in this browser.