Split the docs into chunks, turn them into vectors and find the closest match to the question. "How many vacation days do new hires get?" The best chunk scores 0.91 and the answer, 15 days, comes with its source.
python3 --version.git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
python3 -m venv .venv source .venv/bin/activate
cd ai-engineering/06-rag python3 src/rag.py
import math, re
from docs import DOCS, VOCAB
def embed(text):
words = re.findall(r"[a-z]+", text.lower())
return [words.count(w) for w in VOCAB]
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
return dot / (math.hypot(*a) * math.hypot(*b))
ask = "How many vacation days do new hires get?"
q = embed(ask)
scores = {d: cosine(q, embed(t))
for d, t in DOCS.items()}
top = sorted(scores, key=scores.get,
reverse=True)[:2]
context = "\n".join(DOCS[d] for d in top)
prompt = f"{context}\n\nQuestion: {ask}"
for d in top:
print(d, f"{scores[d]:.2f}", DOCS[d])
print("sources:", *top)Your company's AI never read your HR handbook. Ask about vacation days, and it guesses. The fix: look it up first. That's RAG: retrieval-augmented generation.
Like an open-book exam: find the right page, then answer. An employee asks: how many vacation days do new hires get? First, cut the handbook into eight chunks. Embed turns a chunk into numbers, called a vector.
It counts each vocabulary word in the text. Real systems use a learned model here. Every chunk is now a point. Similar text lands close.
Cosine asks: do two vectors point the same way? Multiply matching counts, divide by length. One is a perfect match. The question is embedded the same way.
We score every chunk against it, then keep the top two. Their text goes into the prompt, above the question. Then print what came back. Let's run it.
Top match: hr-1, 0.91. New hires get 15 vacation days. Second, the laptop rule: same words, different meaning.
A model reading that prompt can answer: 15 days, citing hr-1. Now swap these chunks for your wiki pages, tickets and contracts. Edit a doc, and the answer changes. No retraining.
Every answer cites a source, and you control what gets retrieved. Find the right page, then answer. That's RAG.
Read the lesson on GitHub →