# Evaluation plan and result sheet
Original worksheet by Codex | September 28, 2026

## Define success first
System task:
Intended users:
Allowed inputs:
Expected output:
What is unacceptable even if the overall score is high?

## Dataset
Start with 20 development and 10 held-out cases; expand as failures appear. These are suggested practice sizes, not production validation requirements.
Source of cases:
Permission to use them:
Ordinary / ambiguous / unanswerable / malformed / adversarial cases:
How labels were produced and checked:
How held-out data is kept out of prompt tuning:

## Case schema
case_id; split; input; expected_behavior; relevant_source_ids; objective_checks; human_rubric; risk_category.
The supplied JSONL and synthetic corpus are tiny teaching fixtures. Add domain-specific cases before making quality claims. Do not pass answer labels or expected source IDs into your app at inference time.

## Run record
run_id:
Date:
Code commit:
Model identifier:
Prompt version:
Corpus/index version:
Retrieval configuration:
Number of trials per case:

## Result comparison
| Measure | Baseline | New version | Sample size / notes |
|---|---|---|---|
| Task success (count/total) | | | |
| Supported answers (count/total) | | | |
| Correct abstentions (count/total) | | | |
| Recall@k on labeled cases | | | |
| Median / p95 latency | | | |
| Estimated cost per successful task | | | |
| Critical permission/tool failures | | | |

## Failure review
Case ID:
Observed behavior:
Expected behavior:
Failure layer: ingestion / retrieval / model / validation / tool / product.
Hypothesis:
Smallest experiment:
Result:
Regression case added:

Use deterministic checks for objective conditions. Calibrate model graders against human labels. Report failures, denominators and uncertainty. A test set repeatedly used for tuning is no longer held out.
