Welcome to AIEdTalks’ Newsletter!

This is Part 2 of 4 in “The Agent Loop, From Scratch.”

In today's edition:

  • The executor: where model output meets your systems

  • The loop itself, in full

  • Three API details that cause most “the model is broken” bugs

  • Why every exit should be a typed status, not an exception

Let’s dive in.

After reading, please rate today’s edition.

Rate today's Newsletter

Login or Subscribe to participate

Attio is the agentic CRM for modern teams. It’s your always-on revenue engine: agents and workflows build pipeline, chase every buying signal, and move deals forward alongside your team. Try Attio now.

Today’s Edition

The 40 lines every framework hides

Part 2 of 4. The loop is forty lines. The discipline around it is the job.

Last week we built two tools and made one request. The model came back with stop_reason == “tool_use” and a request to call get_order_status. Then we stopped.

This is Part 2 of 4 in “The Agent Loop, From Scratch.” We're building one production-shaped agent loop, one piece per week, in raw Python with no framework. Part 1 is in the archive if you're joining now — you'll want the tools file.

Today we close the loop.

Part 1: The executor, where the model meets your systems

Before the loop, the boring part that actually decides whether this thing survives contact with production.

The model's tool arguments are input from a client you don't control. So: validate them, time-box them, cap the output size, and — this is the part tutorials get wrong — return failures to the model instead of raising them.

# agent/executor.py
import json
from concurrent.futures import ThreadPoolExecutor, TimeoutError as FutureTimeout
from .tools import TOOLS

_POOL = ThreadPoolExecutor(max_workers=8)
MAX_RESULT_CHARS = 8_000   # tool output is most of your context bill; cap it

def _validate(schema: dict, args: dict) -> None:
    missing = [k for k in schema.get("required", []) if k not in args]
    unknown = sorted(set(args) - set(schema.get("properties", {})))
    if missing:
        raise ValueError(f"Missing required argument(s): {missing}")
    if unknown:
        raise ValueError(f"Unknown argument(s): {unknown}")

def run_tool(block) -> dict:
    """Turn one tool_use block into one tool_result block. Never raises."""
    def error(msg: str) -> dict:
        return {"type": "tool_result", "tool_use_id": block.id, "content": msg, "is_error": True}

    tool = TOOLS.get(block.name)
    if tool is None:
        return error(f"Unknown tool {block.name!r}. Available: {sorted(TOOLS)}")
    try:
        _validate(tool.input_schema, block.input)
        output = _POOL.submit(tool.handler, **block.input).result(timeout=tool.timeout_s)
    except FutureTimeout:
        return error(f"{tool.name} timed out after {tool.timeout_s}s. Retry at most once.")
    except Exception as exc:
        return error(f"{type(exc).__name__}: {exc}")

    text = json.dumps(output, default=str)
    if len(text) > MAX_RESULT_CHARS:
        text = text[:MAX_RESULT_CHARS] + f" ...[truncated {len(text) - MAX_RESULT_CHARS} chars]"
    return {"type": "tool_result", "tool_use_id": block.id, "content": text}

is_error: True is the API's documented way to say “that tool call failed.” The model reads it and can fix its own input — but only if your message says what valid input looks like. Write tool errors the way you'd write a 400 response body.

One caveat: result(timeout=...) stops waiting, it doesn't kill the thread. For real I/O, put a timeout on the underlying client too.

Part 2: The loop

Here it is. This is what you're paying a framework to hide.

# agent/loop.py
import logging
from dataclasses import dataclass, field
import anthropic
from .tools import api_tools
from .executor import run_tool

log = logging.getLogger("agent")
MODEL = "claude-sonnet-5-5"
SYSTEM = ("You are an order-support assistant. Use tools for any order facts; never guess. "
          "If a tool returns an error, fix your input once or tell the user what failed.")
MAX_STEPS = 8

@dataclass
class Result:
    status: str          # done | max_steps | truncated | model_unavailable | refusal | ...
    text: str
    steps: int
    messages: list = field(repr=False)

def _text(resp) -> str:
    return "".join(b.text for b in resp.content if b.type == "text")

def run_agent(user_msg: str, client=None, model: str = MODEL) -> Result:
    client = client or anthropic.Anthropic(timeout=60.0, max_retries=2)
    messages = [{"role": "user", "content": user_msg}]
    tools = api_tools()          # build once: a stable prefix is cache-friendly
    steps = 0

    while True:
        if steps >= MAX_STEPS:
            return Result("max_steps", "", steps, messages)
        try:
            resp = client.messages.create(model=model, max_tokens=2048, system=SYSTEM,
                                          tools=tools, messages=messages)
        except (anthropic.RateLimitError, anthropic.InternalServerError,
                anthropic.APITimeoutError, anthropic.APIConnectionError) as exc:
            # The SDK already retried these. Degrade; don't stack another retry loop on top.
            log.warning("model unavailable after SDK retries: %s", type(exc).__name__)
            return Result("model_unavailable", "", steps, messages)
        # Any other APIStatusError (400/401/404) is a bug in OUR request: let it raise.

        steps += 1
        log.info("step=%d stop=%s in=%d out=%d", steps, resp.stop_reason,
                 resp.usage.input_tokens, resp.usage.output_tokens)
        messages.append({"role": "assistant", "content": resp.content})   # verbatim

        if resp.stop_reason == "tool_use":
            results = [run_tool(b) for b in resp.content if b.type == "tool_use"]
            messages.append({"role": "user", "content": results})          # tool_results ONLY
            continue
        if resp.stop_reason == "end_turn":
            return Result("done", _text(resp), steps, messages)
        if resp.stop_reason == "max_tokens":
            return Result("truncated", _text(resp), steps, messages)
        return Result(resp.stop_reason, _text(resp), steps, messages)

Forty lines. Run it:

r = run_agent("Where is order A-1002? It's late -- open a ticket if needed.")
print(r.status, r.steps, "steps")
print(r.text)

Three details that cause most “the model is broken” bugs

1. Handle every tool_use block, not just the first. The model can request several tools in one turn, and each one needs a matching tool_result in your next message. Miss one and you get a 400.

2. The tool-result message contains only tool_result blocks. Add a chatty line of text after them and you'll get empty end_turn responses or errors. Anthropic's docs call this out specifically; I've still watched teams lose a day to it.

3. Append resp.content back verbatim. Don't filter it down to “just the useful bits.” On current models, reasoning between tool calls lives in blocks you don't want to drop.

Also: don't pass temperature. The v1 SDK dropped sampling params from messages.create(), and several current models reject non-default values.

Why exits are statuses, not exceptions

Notice there's no raise for normal endings. Every way out of the loop is a typed status: done, max_steps, truncated, refusal, model_unavailable.

That's deliberate, and it pays off twice. Your caller branches on status instead of catching a grab-bag of exceptions. And your dashboard can count them — the ratio of done to everything else is the single most useful health metric an agent has.

A loop that can only succeed or crash will do neither. It'll run forever.

Next week — Part 3

Which brings us to what that unbounded while True is quietly doing to your bill. We add token and cost accounting, and I'll show you why an 8-step loop costs roughly 17× a single answer to the same question.

Your turn

Hit reply with one number: what's the real max_steps on your production agent, and how did you pick it? If the honest answer is “there isn't one,” I especially want to hear from you — I'll share the anonymised spread in Part 4.

If you prefer watching to reading, subscribe to the AIEdTalks YouTube channel. https://www.youtube.com/@AIEdTalks

AI is easy to demo. Hard to ship.

Sources

API behaviour (is_error, tool-result message rules, stop reasons, SDK v1 changes) comes from the docs above. The executor design, the typed-status pattern, and the “ratio of done to everything else” metric are my own recommendations.

Who's writing this?

I'm a research engineer with 18+ years building production systems, 35+ patents, and 17+ publications. I work on infrastructure for multi-agent systems. AIEdTalks is my field notes: real problems, what I tried, and what I learned.

Views are my own.

👋 Before you go

💬 Hit reply. I read every reply.

▶️ Watch it instead. This week's breakdown is on YouTube. Subscribe to the channel for more.

📨 Know an engineer who'd find this useful? Share your referral link and earn rewards.

💡 Topic ideas or sponsorships: Reply to this email.

Until next time,

AIEdTalks team.

P.S. AI is easy to demo. Hard to ship. That's what this newsletter is about.

Recommended for you

View all
caret-right