Every AI agent you've heard of runs on one tiny loop: think, act, observe, repeat. Strip away the frameworks and the
marketing, and what's left is a model in a for loop with a few tools and a list it writes into. This article
builds that loop in 16 lines of standard-library Python, runs it, and shows the one line that keeps it from running
forever.
What you’ll needPython 3.10+
What you’ll need
- python.org/downloads →
Python 3.10 or newerRuns the code. Check yours with
python3 --version. - code.visualstudio.com →
VS Code or CursorEither works; so does any editor you like.
- git-scm.com/downloads →
gitGets the code from GitHub.
- No packages needed: it uses only the Python standard library.
- 1Get the code
git clone https://github.com/DayanEbrar0X/data-anatomy.ai.git cd data-anatomy.ai
- 2Make a virtual environmentIt keeps this project's packages separate from the rest of your computer.
python3 -m venv .venv source .venv/bin/activate
- 3Run it
cd ai-engineering/03-ai-agent-loop python3 src/agent.py
Project structure03-ai-agent-loop/
Project structure
- data/
- warehouse.jsonthe numbers the search tool can find
- src/
- agent.pymainthe loop from the video
- model.pythink(): stand-in for the LLM
- tools.pysearch and calculator
- README.mdthe idea, how to run it, a line-by-line walkthrough, things to try
The loop
An agent is a language model in a loop. Each step, it reads what it knows so far, picks an action, runs it, and writes the result back into memory. That's the whole trick.
- Think: the model reads everything in memory and picks an action, like "search for Q3 revenue".
- Act: your code runs the tool it asked for.
- Observe: the result goes back into memory, so the next thought knows more.
The loop stops when the model decides it can answer, or when it hits a step limit.
Tools and memory
Say finance asks: what's 15% of Q3 revenue? The model alone can't know that number. It needs tools. So we hand it two: a search over company data, and a calculator. In a company these would be real APIs. Here they're two small functions, and the "company data" is a JSON file:
# Two tools the agent can call. In a company these would be real APIs.
import json
from pathlib import Path
DATA = Path(__file__).resolve().parents[1] / "data" / "warehouse.json"
WAREHOUSE = json.loads(DATA.read_text())
def search(query):
return f"{query}: {WAREHOUSE[query]}"
def calculator(expr):
a, b = expr.split(" * ")
return float(a) * float(b)Memory starts with just the goal. Now the loop itself, the file you see in the video:
from tools import search, calculator
from model import think # the LLM call
goal = "What is 15% of Q3 revenue?"
memory = [goal]
for step in range(5):
action, arg = think(memory)
if action == "answer":
print("answer:", arg)
break
tool = {"search": search,
"calc": calculator}[action]
result = tool(arg)
memory.append(result)
print(step, action, "->", result)Read it top to bottom and you can see each part of the loop:
- Line 8 is think:
think(memory)reads memory and returns the next action and its argument. - Lines 9 to 11: if the model is ready, it prints the answer and stops.
- Lines 12 and 13 look up the tool it asked for in a dictionary. This is all "tool calling" is: a name the model chose, mapped to a function your code owns.
- Line 14 is act: run the tool.
- Line 15 is observe: write the result into memory, so the next thought knows more.
The model is a stand-in
think() is the model. Here it's a stand-in, so the example runs offline and gives the same answer every time. In
production, it's one LLM call that returns the same (action, argument) shape.
# Stand-in for the LLM: reads memory, returns the next action.
# In production this is one API call that returns the same (action, arg) shape.
def think(memory):
facts = " ".join(map(str, memory))
if "revenue:" not in facts:
return "search", "Q3 revenue"
if len(memory) < 3:
revenue = memory[1].split(": ")[1]
return "calc", f"0.15 * {revenue}"
return "answer", f"${memory[-1]:,.0f}"The rules are simple on purpose. If memory doesn't contain a revenue figure yet, search for it. If it has the figure but no calculation, multiply by 0.15. Otherwise, answer. A real model makes the same three decisions, just from the text of the question and the tool results instead of hard-coded rules.
The run
python3 src/agent.py0 search -> Q3 revenue: 4200000
1 calc -> 630000.0
answer: $630,000Step zero: it searches, and finds Q3 revenue: 4.2 million. Step one: it calls the calculator: 630,000. Step two: it has everything, so it answers.
answer: $630,000. Swap the stand-ins for real APIs and the same loop books meetings, files tickets or
queries your warehouse.
Notice what memory looks like at the end: the goal, the search result, the calculation. That list is the agent's whole state. Everything it "knows" at a step is what earlier steps wrote down, which is why the observe step matters as much as the model.
The guardrail
range(5) on line 7 is the step limit. Without it, an agent can loop forever. A real model can get confused and keep
calling tools, burning time and money on every turn. Production agents always have limits: on steps, on tokens, on
cost, and on which tools they may call.
You can see why with one experiment. Make think() never return "answer". With range(5), the loop gives up after
five steps. With while True, it never stops.
From the stand-in to a real agent
Nothing in agent.py changes when you move to a real model. Only think() does:
- Send the memory to a language model, along with a description of the two tools.
- Ask it to reply with JSON like
{"action": "search", "arg": "Q3 revenue"}. - Parse that reply and return it as
(action, arg).
The tools stay plain functions, the dictionary still maps names to functions, and the guardrail still caps the run. That's the part frameworks wrap for you, and it's worth having written once yourself, because when an agent misbehaves, the bug is almost always in one of these places: what went into memory, which tool was picked, or a loop with no limit.
Try this
- Ask for Q2 instead: change the goal and update
model.pyso the search uses "Q2 revenue". - Add a third tool, for example
format_currency, and makethink()use it before answering. - Break the guardrail: make
think()never return"answer". What happens with and withoutrange(5)? - Replace
think()with a real LLM call that returns JSON like{"action": "search", "arg": "Q3 revenue"}.
Think. Act. Observe. Every agent, one loop. The full lesson, with warehouse.json and a line-by-line table, is in
the repo under ai-engineering/03-ai-agent-loop.