
Welcome to AIEdTalks’ Newsletter!

In today's edition:
Why every step resends the whole conversation
Token and cost accounting in 30 lines
The 17× worked example
The context-window leak nobody budgets for
Let’s dive in.
After reading, please rate today’s edition.
Rate today's Newsletter

Most prompting advice stops at "be specific."
What’s the key to maintaining brand integrity at scale? A lot goes into it, but the teams that get it right have one thing in common: they know how to talk to AI models.
As teams speed up and the required creative volume balloons, communication matters more than ever— both with LLMs, and with the people around you. Creative teams need to focus on ideas, concepts, craft, while technology augments their manual execution demands.
Join our discussion on this evolution with Ari Murray, Chief Digital Officer at Salt and Stone, and get 5 tips for expert LLM prompting with ready to run prompts.
Today’s Edition
Steps are the cost driver

Part 3 of 4. Cost doesn't grow with steps. It grows with steps squared.
Our loop works. It also has a while True in it and no idea what it costs.
This is Part 3 of 4 in “The Agent Loop, From Scratch.” We're building one production-shaped agent loop, one piece per week, in raw Python with no framework.
Today: money.
Part 1: Why agent cost surprises people
A chat request has one bill: prompt in, answer out.
An agent loop resends the entire conversation on every step — system prompt, tool definitions, every previous tool call, every previous tool result. Step 5 pays for everything that happened in steps 1 through 4.
So cost doesn't grow with steps. It grows roughly with steps squared.

Here's the arithmetic (my numbers, not a benchmark — plug in your own). Take a 3,000-token starting prompt: system message, tool definitions, the question. Each step adds about 1,500 tokens of tool call and tool result.
Over 8 steps, total input = 8 × 3,000 + 1,500 × (0+1+…+7) = 66,000 tokens ≈ $0.13, plus ~2,400 output tokens ≈ $0.02.
Answering the same question in one direct call: about 3,000 tokens ≈ $0.009.
That's roughly 17× for the same question. Multiply by your daily request volume before you decide whether an agent is the right shape for the problem.
Part 2: Accounting, in 30 lines
Every response carries usage. Charge it, cap it, and make an unpriced model raise instead of silently costing $0.
# agent/budget.py
from dataclasses import dataclass
# USD per million tokens. Pin and date this -- prices change.
PRICES = {
"claude-sonnet-5-5": {"in": 2.00, "out": 10.00, "cache_write_5m": 2.50, "cache_read": 0.20},
"claude-haiku-4-5": {"in": 1.00, "out": 5.00, "cache_write_5m": 1.25, "cache_read": 0.10},
}
@dataclass
class Budget:
model: str
max_steps: int = 8
max_usd: float = 0.25
steps: int = 0
tokens_in: int = 0
tokens_out: int = 0
usd: float = 0.0
def charge(self, usage) -> None:
p = PRICES[self.model] # KeyError on an unpriced model is deliberate
cw = getattr(usage, "cache_creation_input_tokens", 0) or 0
cr = getattr(usage, "cache_read_input_tokens", 0) or 0
self.steps += 1
self.tokens_in += usage.input_tokens
self.tokens_out += usage.output_tokens
self.usd += (usage.input_tokens * p["in"] + usage.output_tokens * p["out"]
+ cw * p["cache_write_5m"] + cr * p["cache_read"]) / 1_000_000
def over_cost(self) -> bool:
return self.usd >= self.max_usdWire it into the loop — two gates instead of one:
def run_agent(user_msg, client=None, model=MODEL, budget=None):
budget = budget or Budget(model=model)
...
while True:
if budget.steps >= budget.max_steps:
return Result("max_steps", "", budget, messages)
resp = client.messages.create(...)
budget.charge(resp.usage)
log.info("step=%d stop=%s usd=%.4f", budget.steps, resp.stop_reason, budget.usd)
...
if resp.stop_reason == "tool_use":
messages.append({"role": "user", "content": results})
if budget.over_cost():
return Result("budget", "", budget, messages)
continueTwo caps, because they fail differently. max_steps catches a stuck agent. max_usd catches an agent making progress through an expensive context. You want both.
Here's what breaks: tool output eats the window
One SELECT *, or one chatty API response, and step 3 is dragging 50k tokens of JSON through every subsequent step. You get the steep part of that curve and worse answers, because the model is now hunting for the relevant fact inside a wall of noise.
This isn't a corner case. Braintrust analysed a typical agent conversation and found tool responses made up 67.6% of total tokens, with tool definitions another 10.7% — nearly 80% of what the agent sees. The system prompt everyone obsesses over? 3.4%.
That's why Part 2's executor truncates at 8,000 characters. The better fix is upstream: return summaries, not raw dumps. A tool that returns {"status": "shipped", "eta": "2026-10-06"} beats one that returns the whole order object with 40 fields the model will never use.
Design tool outputs like you'd design a mobile API response, not a database row.
What to log per step
One line per step, and you can answer almost any question later:
step=3 stop=tool_use tool=get_order_status in=7400 out=120 usd=0.0163Step number, stop reason, tool name, tokens in/out, cumulative USD. Cheap to add now, impossible to reconstruct after the incident.
Next week — Part 4 (final)
The last part is the one that decides whether this thing is shippable: timeouts, a single retry budget, idempotent tools, and tests that run with no API key. Plus two production failures I've watched happen, and the checklist I'd run before letting an agent near a customer.
Your turn
Hit reply: do you know what one agent run costs you today? Not your monthly bill — one run. If you don't, you're not alone, and that's the point of this issue.
If you prefer watching to reading, subscribe to the AIEdTalks YouTube channel. https://www.youtube.com/@AIEdTalks
AI is easy to demo. Hard to ship.
Sources
Prices and caching behaviour come from Anthropic's docs (checked October 2026 — re-check before relying on them). The 67.6% token-share figure is Braintrust's. The 17× example is my own arithmetic on assumed token counts, and the two-cap design and logging format are my recommendations.
Who's writing this?
I'm a research engineer with 18+ years building production systems, 35+ patents, and 17+ publications. I work on infrastructure for multi-agent systems. AIEdTalks is my field notes: real problems, what I tried, and what I learned.
Views are my own.
👋 Before you go
💬 Hit reply. I read every reply.
▶️ Watch it instead. This week's breakdown is on YouTube. Subscribe to the channel for more.
📨 Know an engineer who'd find this useful? Share your referral link and earn rewards.
💡 Topic ideas or sponsorships: Reply to this email.
Until next time,
AIEdTalks team.
P.S. AI is easy to demo. Hard to ship. That's what this newsletter is about.


