Ten real tickets, scored on every change: version 1 hits 60%, version 2 80%, version 3 100%. Keep the tickets as a regression gate and every release is measured.
python3 --version.git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
cd methodologies/05-evals-loop-engineering python3 src/evals.py
from tickets import CASES, route, VERSIONS
for rules in VERSIONS:
ok = [route(t, rules) == want
for t, want in CASES]
rate = sum(ok) / len(ok)
print(f"{rate:.0%} pass")Your AI aced the demo, then got four out of ten tickets wrong. The fix is a loop: run, score, fix, rerun. Ten real tickets with right answers: a practice test for your AI. Each version of the router runs every ticket.
The pass rate is the share it got right. Version one: 60%. Read the failures, fix one thing, rerun. 80%.
Then 100%. Keep those tickets. Every future change must pass them too. Run, score, fix, rerun.
Read the lesson on GitHub →