
Welcome to AIEdTalks’ Newsletter!
This is Part 1 of 4 in “The Agent Loop, From Scratch.”
In today's edition:
Why I want you to write the agent loop yourself once
What a tool actually is (three things, no magic)
Your first API call with tools attached
What we're building over the next four weeks
Let’s dive in.
I put a lot of care into each edition, and your rating would mean a lot. It would also help me make the next one even better :)
Rate today's Newsletter
Stop typing AI prompts. Start talking.
You think 4x faster than you type. So why are you typing prompts? Wispr Flow turns your voice into ready-to-paste text inside any AI tool. Speak naturally, tangents and all, and Flow cleans it up. Available on Mac, Windows, iPhone, and Android.
Today’s Edition
Build an agent from scratch

Part 1 of 4. A tool is a name, a schema, and a function.
Every agent framework you'll be asked to use — LangGraph, the OpenAI Agents SDK, smolagents, Anthropic's Tool Runner — is built on the same twenty lines: call the model, check whether it wants a tool, run the tool, send the result back, repeat.
Most engineers moving into AI skip those twenty lines. They pip install a framework, get a demo running by lunch, then spend a month debugging behaviour they can't see.
I've been building production systems for 18+ years. The most useful thing I tell backend engineers making this move is: write the loop yourself once. Not because frameworks are bad — I use them. Because you can't operate what you don't understand, and an agent loop is just a distributed system where one node is non-deterministic and bills you per token.
This is Part 1 of 4 in “The Agent Loop, From Scratch.” We're building one production-shaped agent loop, one piece per week, in raw Python with no framework.
Here's the plan:
Part 1 (today) — Tools and your first tool-calling request. What a tool actually is.
Part 2 — The executor and the loop. The 40 lines frameworks hide.
Part 3 — Token and cost accounting. Why steps are the cost driver.
Part 4 — Timeouts, retries, idempotency, tests. The loop you'd let page you.
By the end you'll have a repo you could defend in a design review. Today we start with the part everyone skips.
Part 1: What a tool actually is
The industry finally agreed on a definition. Simon Willison's is the one that stuck: an LLM agent runs tools in a loop to achieve a goal.
That's it. The “agent” is not a product or a personality. It's a loop with tools in it.
So start with the tool. A tool is three things:

A name, a JSON Schema describing what the model should fill in, and a function of yours that runs. No decorators, no base classes, no framework.
The part engineers underrate: the description is an API contract. The model reads it to decide whether to call your tool. A vague description is a flaky agent, and you'll blame the model.
Part 2: Build it
Versions this was written against: Python 3.10+, anthropic 1.11.0, model claude-sonnet-5-5. The SDK went 1.0 in August 2026, so if you're pinned to 0.x, check MIGRATION.md.
python -m venv .venv && source .venv/bin/activate
pip install "anthropic>=1.11,<2" pytest
export ANTHROPIC_API_KEY=sk-ant-...
mkdir -p agent tests && touch agent/__init__.pyTwo tools — one read, one write. The write one matters later, when we talk about what happens if the model calls it twice.
# agent/tools.py
from dataclasses import dataclass
from typing import Any, Callable
@dataclass(frozen=True)
class Tool:
name: str
description: str
input_schema: dict
handler: Callable[..., Any]
timeout_s: float = 5.0
# Fake backends -- swap for real clients later.
ORDERS = {
"A-1001": {"status": "shipped", "carrier": "UPS", "eta": "2026-10-06"},
"A-1002": {"status": "processing", "eta": None},
}
TICKETS: dict[str, dict] = {}
def get_order_status(order_id: str) -> dict:
if order_id not in ORDERS:
# Error text is written FOR THE MODEL: what went wrong, and what valid input looks like.
raise LookupError(f"No order {order_id!r}. Order IDs look like 'A-1001'.")
return ORDERS[order_id]
def create_ticket(order_id: str, summary: str) -> dict:
if order_id not in TICKETS:
TICKETS[order_id] = {"ticket_id": f"T-{order_id}", "order_id": order_id, "summary": summary}
return TICKETS[order_id]
TOOLS = {t.name: t for t in [
Tool("get_order_status",
"Look up the current status of ONE order by ID. Use before answering any question about an order.",
{"type": "object",
"properties": {"order_id": {"type": "string", "description": "e.g. 'A-1001'"}},
"required": ["order_id"]},
get_order_status),
Tool("create_ticket",
"Open a support ticket for a human. Only use when the user reports a problem you cannot resolve.",
{"type": "object",
"properties": {"order_id": {"type": "string"}, "summary": {"type": "string"}},
"required": ["order_id", "summary"]},
create_ticket),
]}
def api_tools() -> list[dict]:
"""The exact shape the Messages API expects in `tools=`."""
return [{"name": t.name, "description": t.description, "input_schema": t.input_schema}
for t in TOOLS.values()]Now one request. Not a loop — one call, so you can see exactly what comes back.
# run.py
import anthropic
from agent.tools import api_tools
client = anthropic.Anthropic(timeout=60.0) # never rely on the 10-minute default
resp = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=1024,
system="You are an order-support assistant. Use tools for any order facts; never guess.",
tools=api_tools(),
messages=[{"role": "user", "content": "Where is order A-1001?"}],
)
print(resp.stop_reason)
for block in resp.content:
print(block.type, getattr(block, "name", ""), getattr(block, "input", ""))Run it. You get:
tool_use
tool_use get_order_status {'order_id': 'A-1001'}Stop and look at that, because it's the whole mental model. The model did not fetch anything. It can't. It returned a structured request — “please call get_order_status with this argument” — and stopped.
stop_reason == “tool_use” is the model raising its hand. Nothing happens until your code does something about it.
What this means for you
If you've built HTTP services, you already know how to think about this:
The model is a client calling your API. Its arguments are untrusted input — validate them.
stop_reason is a status code. You branch on it.
The model has no memory between calls. You own the conversation state. Every step resends it, which is why Part 3 is about cost.
That framing is the thing frameworks take away from you. It's also the thing that makes agents debuggable.
Next week — Part 2
We run the tool and close the loop: the executor that validates arguments, times out, and turns failures into something the model can recover from — plus the forty lines that every framework hides.
Your turn
Hit reply and tell me: what's the first tool you'd give an agent on your system? Read-only, or does it write? I read every reply, and I'll use the best ones in Part 4.
If you prefer watching to reading, subscribe to the AIEdTalks YouTube channel. https://www.youtube.com/@AIEdTalks
AI is easy to demo. Hard to ship.
Sources
The agent definition and the API behaviour come from the sources above. The code structure, the tool design, and the “model is a client calling your API” framing are my own.
Who's writing this?
I'm a research engineer with 18+ years building production systems, 35+ patents, and 17+ publications. I work on infrastructure for multi-agent systems. AIEdTalks is my field notes: real problems, what I tried, and what I learned.
Views are my own.
👋 Before you go
💬 Hit reply. I read every reply.
▶️ Watch it instead. This week's breakdown is on YouTube. Subscribe to the channel for more.
📨 Know an engineer who'd find this useful? Share your referral link and earn rewards.
💡 Topic ideas or sponsorships: Reply to this email.
Until next time,
AIEdTalks team.
P.S. AI is easy to demo. Hard to ship. That's what this newsletter is about.

