Run, score, fix, rerun on 10 real support tickets: the ticket router goes from 60% to 80% to 100%. Then the eval set stays as a regression gate, so a prompt change can't quietly break production.
python3 --version.git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
cd methodologies/05-evals-loop-engineering python3 src/evals.py
from cases import CASES
def route(text, rules):
for team, words in rules.items():
if any(w in text for w in words):
return team
return "other"
def score(v, rules):
fails = [t for t, want in CASES
if route(t, rules) != want]
rate = 1 - len(fails) / len(CASES)
print(v, "pass rate", f"{rate:.0%}")
rules = {"billing": ["refund", "charge"],
"tech": ["crash", "error"]}
score("v1", rules)
rules["account"] = ["password", "log in"]
score("v2", rules)
rules["billing"].append("invoice")
score("v3", rules)Your AI aced the demo. In production, it gets four out of ten tickets wrong. The fix is a loop: run, score, fix, rerun. Like a teacher grading practice tests with an answer key.
Our system routes support tickets to billing, tech, or account. Does it work? Don't guess. Measure.
First, the eval set: ten real tickets, each labeled with the right team. Route is the system we're testing. Here it's keyword rules: a stand-in for an LLM prompt. It returns the first team whose words match, or other.
Score runs every ticket and keeps the ones it got wrong. The pass rate is the share it got right. Version one knows two teams: billing and tech. Run it.
60%. Four tickets failed. Read the failures. Password and log-in tickets get labeled other.
There's no account team. So the fix: add one. Rerun. 80%.
Two still fail. Both mention an invoice. So billing learns the word invoice. Rerun.
100%. Every ticket routed right. Run the file: 60, 80, 100. In real life, route is an LLM prompt, an agent, or a RAG pipeline.
Keep the eval set as a regression test. Change the model or prompt. If the pass rate drops, it doesn't ship. Teams watch that number like uptime.
Run. Score. Fix. Rerun.
That's loop engineering.
Read the lesson on GitHub →