An agent run is a long-lived distributed workflow, not a single request. Build it like one: with checkpoints and idempotency, not a hopeful while-loop.
Every MCP server you connect makes your agent more capable on paper — and worse in production. Here's the failure, and the ~50 lines of Python that fix it.