Skip to content

Module 5: Agent Engineering

Listen to this page14:09
Read the transcript

1. An Agent Is Not a Smarter Model

Host: Okay, module five, agent engineering. And I want to start by killing a phrase, because I think it’s actively getting in people’s way. An agent is not a smarter model. That’s not what we’re building here.

Guest: Right, and I think that’s the single most common misconception people bring into this. They think if the model gets good enough, it just becomes an agent on its own. But an agent is a control loop — observe, decide, act, observe again — and the model only owns one step of that, the deciding part. Everything else is structure you build around it.

Host: So when we say agent engineering is a discipline, we mean the discipline is really about that loop — making it safe, bounded, debuggable — not about coaxing more intelligence out of the model. And the twist is that every hard problem we’re about to cover — runaway execution, side effects, permissions, partial failure — is just a control-systems problem this handbook already has words for, now sitting at the seam between what the model asks to do and what your system actually lets happen.

2. Inside the Loop: A Model That Requests, a System That Decides

Host: So let’s get concrete about that seam. When the model decides to use a tool, what actually happens — it’s not like it’s reaching out and executing something on your machine, right?

Guest: Right, it never runs anything. It emits a structured request — a function name and some arguments — that matches a schema you gave it up front. Your system is the one that validates that request, decides whether it’s even well-formed, and only then executes it and feeds the result back in as the next observation.

Host: So a hallucinated tool name or garbage arguments never even gets a chance to run — it just bounces back as an observation the model has to react to. That’s a pretty clean trust boundary to design around.

3. Two Guards That Separate a Demo from a Production Loop

Host: Right, but the trust boundary isn’t just about catching malformed calls — the diagram shows two actual guard components doing the load-bearing work here. What are they?

Guest: The first is a step or cost budget check, and it runs before every single decision, not just once at the start — without it enforced continuously, an agent isn’t a reliability feature, it’s just a resource-usage risk waiting to happen. The second is a permission and schema check sitting between the model’s decision and actual execution, and that’s the one people skip in demos: the model requesting a tool call is not the same thing as the system agreeing to run it.

4. Four Levers to Stop a Loop from Running Forever

Host: So the budget check runs continuously, fine — but what’s actually in that check? You said step or cost, and I want to slow down on that, because it sounds like one thing and I suspect it’s not.

Guest: It’s four separate levers, and all four need to exist at once: a step budget capping total iterations, a cost budget capping estimated spend, a wall-clock timeout, and the model’s own done-signal. Each one catches a runaway shape the others miss — a step budget catches an infinite loop of cheap tool calls, a cost budget catches a handful of very expensive calls that never trip a step count, and a wall-clock timeout catches a loop that’s technically making progress but far too slowly to be useful. The done-signal is real too, it’s just not sufficient on its own — relying on the model to know when to stop isn’t a termination condition, it’s a hope, so it rides alongside the other three rather than replacing them.

Host: And you’d said context growth needs its own budget separately from that — why isn’t a wall-clock timeout enough to catch that too?

Guest: Because context growth is a cost and latency problem that compounds silently long before you’d hit a hard limit or a timeout — every loop iteration adds tokens to the running context, so cost and latency compound over a long task even if nothing’s gone wrong. That’s why summarizing or pruning older observations has to happen well before the context window’s actual ceiling, not at it. And separately, per-step latency needs its own budget too, because one slow tool call can quietly eat the entire task’s time allowance while every other check still looks green.

5. Plan-Then-Act vs. Interleaved Reasoning

Host: Let’s talk about how the loop actually structures its thinking. There seem to be two schools here — have the model draw up a full plan before touching anything, or have it think and act one step at a time. Walk me through that split.

Guest: Right, so the first approach is plan-then-act: the model produces a complete sequence of steps upfront, and only then does execution start. That’s great for review — a human can look at the whole plan before anything real happens, and it gives clearer visibility into what the agent intends to do. The trade-off is it’s brittle: it’s less adaptive when the plan doesn’t survive contact with real tool results.

Host: And the alternative is the ReAct pattern you mentioned back when we were talking about the loop itself — reason, act, observe, repeat.

Guest: Exactly, that interleaving is what makes it adaptive — each step’s reasoning incorporates the actual observation from the last action, so it responds to reality as it unfolds instead of assuming the plan still holds. The cost is reviewability: there’s no fixed plan sitting there for a human to inspect before execution starts, because the plan is being generated one step at a time as it goes. Neither one wins outright — it’s really a question of whether the task needs to be reviewable before it starts or adaptive while it’s running.

6. Memory Isn’t Just the Context Window

Host: Okay, next control-systems problem: memory. It’s tempting to think the context window just is the agent’s memory, but you’re saying that’s already a category error.

Guest: Right, the context window is short-term memory only — it’s bounded, and every token in it costs money and latency on every single call, so an agent that treats it as the only memory either runs out of room on long tasks or pays to re-send the whole history every turn. Long-term memory is retrieval from an external store, something Module 8 goes deep on, and it’s what lets the agent recall something from beyond what fits in context without carrying it in-window the whole time. Production agents need both — context for immediate continuity within the current step, retrieval for anything that has to persist beyond what fits or beyond what’s affordable to keep re-sending.

7. The Bounded Agent Loop, in Code

Host: Let’s make this concrete with actual code, because I think the two guards you described — the step budget and the permission check — sound abstract until you see where they live. Walk me through this BoundedAgentLoop.

Guest: Sure. The run method loops up to max_steps times, calling decide with the observations so far, and if the model returns a final answer, you’re done. But look at what happens on a tool call — it goes to an execute step, and every single failure path in there, an unknown tool name, a permission denial, a tool that literally throws an exception, all of them return a string observation instead of raising. Nothing crashes the loop.

Host: So a tool blowing up doesn’t end the task, it just becomes another line of text the model reads on the next step. That’s a deliberate design choice, not just error handling for its own sake.

Guest: Exactly, the model gets to see its own failures the same way it sees successes and react to them, maybe retry with different arguments, maybe try a different tool. And the permission check is the other guard made concrete — each tool carries a set of task types it’s allowed to run for, checked with a plain if statement in code, not a suggestion in a prompt. If the task type isn’t in that set, the tool never runs, full stop, no model judgment involved.

8. A Support Ticket, Including the Moment It Breaks

Host: Let’s walk through an actual ticket so this stops being abstract. What does one clean iteration look like end to end?

Guest: The model gets a ticket with some error message, decides to call search_kb with that error text, and the tool comes back with three matching articles as the observation. Now the model has real context, so on the next step it calls draft_reply, pointing at one of those articles as the source. Two tool calls, each result fed back in before the next decision — nothing exotic, just the loop doing its job.

Host: Okay, now break it. What happens when draft_reply actually fails?

Guest: Say the templating service is down and draft_reply raises. That exception doesn’t crash the task, it just becomes another observation — literally a string saying the tool failed and why — and the model sees that like it sees any other result. It might retry, it might escalate to a human tool instead, or it might burn through its step budget without resolving anything, and at that point the loop’s own fallback fires: incomplete, step budget exhausted, handing off to a human. That’s the whole point of the guard — a graceful, informative stop instead of a silent hang or the model hammering the same broken call forever.

9. Where Agents Actually Break

Host: So we’ve seen one failure mode up close. What does the fuller catalog look like — where do these loops actually go wrong in production?

Guest: Start with the boring one: no budget at all — that’s the same step and cost cap we already covered as the fix. Then there’s malformed tool calls — the model invents a tool that doesn’t exist, or gets the arguments wrong, and if you dynamically dispatch on that unchecked output instead of validating against a schema first, you’re executing garbage.

Host: And the side-effect one — that’s got to be where Module 2 comes back, right? Ambiguous timeouts on a real action like sending an email or charging a card.

Guest: Exactly the same problem, just with a model in the loop instead of a retry policy — the tool call times out, the agent can’t tell if it landed, and if it retries blind, that email or charge fires twice. Same fix too: idempotency keys, claimed before the action runs. And there’s one more subtler one — a model can shrug off a failed step and hand you a confident answer built on top of it instead of surfacing that something never actually succeeded.

10. Guardrails and the Autonomy Trade-off

Host: So beyond retries and idempotency, there’s a whole other category of failure that’s about permissions and trust. What does that look like in practice?

Guest: It starts with least-privilege tool scoping — the allowed for task types check we already wired into the loop isn’t just a routing detail, it’s the concrete mechanism that keeps a tool scoped to exactly what the task needs, rather than being an afterthought bolted onto a tool that can already do more. And anything irreversible — sending an email, spending money, deleting data — needs a mandatory human approval gate, full stop, not a config flag someone forgot to set. Code execution tools get sandboxed completely, no host filesystem, no network, no ambient credentials beyond exactly what that task needs.

Host: And there’s a sneakier injection risk too, right — not through the prompt itself?

Guest: Right, it comes in through tool outputs — a fetched webpage or a document can contain text that reads like an instruction, and the model’s next loop iteration might just follow it. It’s the same prompt injection problem Module 4 covers, just arriving as an observation instead of the initial prompt, which is exactly why it’s easy to miss. That’s the security side — the other half is a real trade-off: full autonomy is fast and needs no human around, but every action fires without a check, so it only belongs on low-risk reversible stuff, while approval gates buy you a real backstop at the cost of latency and someone having to be there. Same tension shows up in tool-set shape — one agent with every tool is simple to operate but gets slower and dumber as its context fills with tool descriptions, while multiple narrow agents scale better but now you need the ownership and lease coordination from Module 2 to keep them from stepping on each other.

11. From Theory to Enforcement: The Policy-Gated Tool Runtime

Host: So if someone wants to see all of this stop being theory and start being code, that’s the policy-gated tool runtime lab. Walk me through what’s actually enforced there.

Guest: It’s the pipeline version of everything we just talked about, in a specific order for a reason: scope check first because it’s cheapest and shouldn’t leak that a tool even exists to someone with no grant on it, then schema validation so a malformed call can’t burn a well-behaved caller’s rate limit, then the limiter itself keyed per tenant and per tool so a search and a fund transfer never share a budget, and only last the approval gate — because that’s the one stage that blocks on a human, and you don’t pay that latency for a call that was going to fail anyway. And it logs denials just as carefully as successes, because a tool nobody ever approves needs to be visible in the audit trail, not buried.

Host: And the extension exercise ties it straight back to Module 2, right — idempotency keys on the tool call itself?

Guest: Exactly — you add an idempotency key to the ToolCall, wire it through the same IdempotentTaskStore pattern from the distributed systems module, and write a test that fires the same call twice and proves the handler only actually runs once. That’s the whole point of this module in miniature: an agent loop is a control system, and once a step can retry, every side-effecting action needs the same durability guarantee a distributed job would need. Get that right and you’ve genuinely built the thing, not just talked about it — which is as good a place as any to leave this module.

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

An agent is not a smarter model — it’s a control loop that lets a model take actions in the world and observe the results, repeated until a task finishes or a limit is reached. Agent engineering is the discipline of making that loop safe, bounded, and debuggable. Every hard problem that shows up in this module — runaway execution, side effects that need to be retry-safe, permission boundaries, partial failure — is a control-systems problem the rest of this handbook has already built the vocabulary for; agents just apply it to a loop where the “logic” is a model’s output instead of code you wrote.

The agent loop, stripped to its essentials: observe the current context and state, decide the next action (call a tool, ask a clarifying question, or declare the task finished), act by executing that decision, observe the result, and repeat. Nothing about this loop is mysterious — it’s the same observe-decide-act-observe shape as any control system — except that the “decide” step is a model call instead of deterministic code, which means every other step has to compensate for that non-determinism with structure the model itself doesn’t provide.

The model never executes anything

A model producing a tool call doesn’t run code — it emits a structured request (a function name and arguments). Your system validates that request, executes it, and feeds the result back in as the next observation. Every trust and safety boundary in this module exists at that exact seam — between the model’s request to act and your system’s decision to actually act on it.

Two guards in this diagram are what separate a production agent loop from a demo: the step/cost budget check runs before every decision, not just at the start — an agent without an enforced budget is a resource-usage risk, not a reliability feature. And the permission and schema check sits between the model’s decision and actual execution — the model requesting a tool call is not the same as the system agreeing to run it.

Tool calling mechanics. The model is given a schema describing available tools (name, description, argument types) and emits a structured call matching that schema when it decides to use one. Your code validates the call against the schema — malformed or hallucinated tool calls (a tool name that doesn’t exist, or arguments that don’t match) get rejected before execution, with the rejection fed back as an observation the model can react to, not silently ignored or crashed on.

Planning vs. execution separation. Some architectures have the model produce a full plan upfront, then execute it step by step — more predictable and easier for a human to review before anything happens, but less adaptive when reality diverges from the plan mid-execution. Others interleave planning and acting every step (the ReAct pattern: reason about the next step, act, observe, reason again) — more adaptive to new information, harder to review ahead of time since there’s no fixed plan to inspect. Neither is strictly better; the right choice depends on how reviewable the task needs to be before it starts versus how much it needs to adapt mid-flight.

Memory: short-term vs. long-term. Short-term memory is the conversation/context window itself — bounded, and expensive to fill (every token in context costs money and latency on every subsequent call). Long-term memory is retrieval from an external store (Module 8 covers this in depth). An agent that treats its context window as the only memory it has either runs out of room on long tasks or pays to re-send everything every turn; one that never uses the context window loses the immediate continuity a task needs. Production agents need both, used for what each is actually good at.

Termination conditions. A step budget (maximum loop iterations), a cost budget (maximum estimated spend, per Module 4’s cost tracking), a wall-clock timeout, and an explicit “done” signal the model can emit are the four levers that keep a loop from running forever. All four should exist simultaneously — relying on the model to “know when to stop” is not a termination condition, it’s a hope.

Research Note

The paper that named and popularized the reason-act-observe loop this module’s Mental Model section describes — worth reading for how interleaving reasoning traces with actions improves both performance and interpretability over acting without an explicit reasoning step.

Source: Yao et al., "ReAct: Synergizing Reasoning and Acting in Language Models"

A minimal, bounded agent loop with a step budget and a permission-checked tool executor — the two guards from the Architecture diagram, made concrete:

A bounded agent loop with permission-checked tool execution

from __future__ import annotations
from collections.abc import Awaitable, Callable
from dataclasses import dataclass, field
from typing import Any
class ToolCallRejected(RuntimeError):
pass
@dataclass
class ToolCall:
name: str
arguments: dict[str, Any]
@dataclass
class AgentStep:
kind: str # "tool_call" | "final_answer"
tool_call: ToolCall | None = None
final_answer: str | None = None
@dataclass
class Tool:
name: str
handler: Callable[[dict[str, Any]], Awaitable[str]]
allowed_for_task_types: set[str] = field(default_factory=set)
class BoundedAgentLoop:
def __init__(self, tools: dict[str, Tool], max_steps: int = 8) -> None:
self._tools = tools
self._max_steps = max_steps
async def run(
self,
task_type: str,
decide: Callable[[list[str]], Awaitable[AgentStep]],
) -> str:
observations: list[str] = []
for _ in range(self._max_steps):
step = await decide(observations)
if step.kind == "final_answer":
return step.final_answer or ""
assert step.tool_call is not None
observations.append(await self._execute(step.tool_call, task_type))
return "incomplete: step budget exhausted, handing off to a human"
async def _execute(self, call: ToolCall, task_type: str) -> str:
tool = self._tools.get(call.name)
if tool is None:
return f"error: unknown tool '{call.name}'"
if task_type not in tool.allowed_for_task_types:
return f"error: tool '{call.name}' not permitted for task type '{task_type}'"
try:
return await tool.handler(call.arguments)
except Exception as exc:
return f"error: tool '{call.name}' failed: {exc}"

Every failure path — an unknown tool, a permission denial, a tool that raises — returns a string observation rather than raising out of the loop. That’s deliberate: the model gets to see and react to failures the same way it sees successes, instead of the whole task crashing because one tool call didn’t work. The allowed_for_task_types set on each Tool is the permission boundary from the Architecture diagram, enforced in code, not left to the model’s judgment about what it should or shouldn’t do.

An agent is tasked with resolving a support ticket, with tools to search a knowledge base, draft a reply, and escalate to a human. One loop iteration: the model decides to call search_kb with the ticket’s error message; the tool executes and returns three matching articles as the observation; the model, now with that context, decides to call draft_reply referencing one of them.

Now the failure case: draft_reply raises because the templating service is briefly down. Per the Implementation section, that failure becomes an observation — "error: tool 'draft_reply' failed: ..." — not a crashed task. The model, seeing the failure, might retry, might try a different tool, or might reach the step budget without resolving the ticket, at which point the loop’s fallback ("incomplete: step budget exhausted, handing off to a human") fires — a graceful, informative stop instead of a silent hang or an unbounded retry storm.

Runaway loops with no enforced budget

An agent with no step or cost budget can loop indefinitely on a task it can’t actually complete, consuming model calls (and their cost) the whole time. The budget check in this module’s Architecture section exists specifically to bound this — “the model should know when to stop” is not a mitigation.

Tool call hallucination

A model can emit a call to a tool that doesn’t exist, or with arguments that don’t match the expected schema. Validate every tool call against its schema before execution — never getattr() or dynamically dispatch based on unchecked model output.

Side-effecting tool calls retried without idempotency

If a tool call with a real side effect (send an email, charge a payment) fails ambiguously — the exact ambiguous-timeout problem from Module 2 — and the agent retries it, the side effect can execute twice. Any tool with a real-world side effect needs the same idempotency-key discipline Module 2 covers for any other distributed system.

Context window overflow silently drops earlier context

A long-running agent task can exceed its context window, silently truncating early context the model actually needed — with no error, just degraded behavior that looks like the model “forgetting” something it was told several steps ago. Long-running tasks need an explicit summarization or offload-to-long-term-memory strategy before hitting that limit, not a hope that it won’t come up.

Confident progression past a failed step

A model that treats a failed tool call as a minor setback and proceeds anyway — rather than surfacing the uncertainty or stopping — can produce a confident-sounding final answer built on a step that never actually succeeded. Surfacing failed steps explicitly in the final output (or routing to a human when a required step failed) is safer than letting the model paper over it.

Upfront planning vs. interleaved plan-and-act

A full upfront plan is easier to review before execution starts and gives clearer visibility into what the agent intends to do — at the cost of being less adaptive when the plan doesn’t survive contact with real tool results. Interleaved ReAct-style planning adapts step by step, at the cost of having no fixed plan a human can review before anything happens.

Full autonomy vs. human-in-the-loop approval gates

Full autonomy is faster and requires no human availability, but every action executes without a check — appropriate for low-risk, easily reversible actions. A human approval gate before high-risk or irreversible actions (sending an email, spending money, deleting data) adds latency and requires a human to be available, in exchange for a real safety backstop the loop itself can’t provide.

One agent with many tools vs. multiple specialized agents

A single agent with a large tool set is simpler to operate as one system, but its context grows with every tool description and its decision quality can degrade with too many options to choose from. Multiple specialized agents, each with a narrow tool set, coordinated via the ownership and lease patterns from Module 2, scale better in tool-set size at the cost of needing real coordination machinery between them.

  • Least-privilege tool scoping: a tool available to an agent should be scoped to exactly what the task needs — the allowed_for_task_types check in this module’s Implementation section is the concrete mechanism, not an afterthought bolted onto a tool that can already do more.
  • Human approval gates for irreversible or high-risk actions — sending external communication, spending money, deleting data — as a hard requirement, not a configurable nicety.
  • Sandbox any code-execution tool completely, with no access to the host filesystem, network, or credentials beyond what the specific task explicitly requires.
  • Prompt injection via tool outputs: content returned by a tool (a fetched webpage, a document) can contain instructions that steer the next loop iteration — the same failure mode Module 4 covers for direct prompts, but arriving through an observation instead of the initial prompt, which makes it easy to overlook.
  • Context grows every iteration — each loop step adds tokens to the running context, so cost and latency compound over a long task; summarizing or pruning older observations is often necessary well before hitting the hard context-window limit.
  • Parallelize independent tool calls rather than serializing them by default — if a step’s decision genuinely doesn’t depend on another pending call’s result, running them concurrently (bounded, per Module 1) cuts wall-clock time directly.
  • Budget per-step latency separately from total task latency — a single slow tool call shouldn’t be allowed to consume the entire task’s time budget unnoticed.
  • Many concurrent agent sessions need the same bounded-concurrency discipline as any other I/O-heavy workload — see Module 1’s treatment of semaphores and admission control, applied to sessions instead of individual requests.
  • Long-running tasks need durable, checkpointed state so a process restart doesn’t lose in-progress work — the same recovery requirements Module 2 covers for any distributed system, applied to an agent’s own task state.
  • Multi-agent coordination at scale needs the same ownership-lease and durable-workflow-state patterns Module 2 covers for coordinating any concurrent workers, not a bespoke agent-specific mechanism.

Walk me through the agent loop.

Observe the current context, the model decides an action (tool call, clarifying question, or final answer), the system validates and executes that action if it’s a tool call, the result becomes the next observation, and the loop repeats until a termination condition fires.

How do you prevent an agent from looping forever or running up cost?

An enforced step budget, a cost budget tied to estimated spend, and a wall-clock timeout — all three checked before every decision, not just at the start, with a defined graceful-stop behavior (hand off to a human) rather than an abrupt failure.

Why three separate budgets, not one

Each catches a different runaway pattern: a step budget catches infinite tool-call loops even if each call is cheap; a cost budget catches a small number of very expensive calls a step budget wouldn’t flag; a wall-clock timeout catches a loop that’s technically making progress but too slowly to be useful. Relying on only one leaves the other two failure shapes uncovered.

How do you make a tool call with a real side effect safe to retry?

The same idempotency-key pattern from Module 2 — the tool call carries a stable key, and the tool’s handler checks for a prior execution with that key before acting, so a retry (from the agent or from ambiguous failure) doesn’t double-execute the side effect.

When would you use multiple specialized agents instead of one agent with many tools?

When a single agent’s tool set has grown large enough to degrade its decision quality, or when different parts of a task genuinely benefit from different context/instructions — at the cost of needing real coordination (ownership, leases, shared durable state) between the agents, which is not free.

Policy-Gated Tool Runtime implements this module’s safety boundaries as an enforcement pipeline: capability scoping, JSON Schema argument validation, per-tool rate limiting, human approval for high-risk tools with a bounded wait, and an audit log that records denials as carefully as successes. It is labelled production-shaped — the pipeline is complete and tested, the backing stores are in-memory and caller identity is a header rather than a verified token.

Add idempotent tool execution to the bounded agent loop

Extend the BoundedAgentLoop from this module’s Implementation section so that ToolCall carries an idempotency key, and _execute checks a durable store (reusing the IdempotentTaskStore pattern from Module 2) before running a tool’s handler. Write a test that retries the same tool call twice with the same key and asserts the handler executes exactly once.

Before running an agent loop against real tools in production

  • Every tool call is validated against its schema before execution, never dynamically dispatched
  • A step budget, a cost budget, and a wall-clock timeout are all enforced, not just one
  • Every tool is scoped to least privilege for the task type that uses it
  • Irreversible or high-risk actions require a human approval gate
  • Tool calls with real side effects carry an idempotency key
  • Failed steps are surfaced in the final output or routed to a human, not silently papered over
Version Date Change
1.0.0 2026-08-05 Initial publication.