@iamsajaldubey
Module 10The setup planner: an agent you can inspect
@iamsajaldubey
PROJECT LAB / MODULE 10

The setup planner: an agent you can inspect

Watch an agent choose a calculator, compare options and stop at its limit.

The situation

Two fictional setup options cost 3×18 + 2×12 and 4×20 + 1×15. The approved budget is 90. The agent may use a Calculator tool; it has no shopping or weather tool.

Your goal

Build a bounded tool-using agent and explain its actual execution trace.

A first win

Calculate the first option as78 and the remaining budget as12.

Keep this artifact

An agent workflow, a tool-call trace and a three-case boundary report.

Explore the mechanism

This interactive model teaches the mechanism. It does not call a model, search your files or send messages.

Why this works

Tool choice is the distinguishing step

The model can choose a connected tool and use its result. A fixed chain has a predetermined model call instead. Merely writing the word agent in a prompt does not connect capabilities.

Tool output is evidence

The arithmetic result should come from the calculator call and be visible in the trace. A plausible narrated tool call is not proof that an execution happened.

A stopping rule makes failures inspectable

Set a maximum iteration count and an error path. If the task needs an unavailable tool, the agent should explain the missing capability rather than inventing a result.

Build it, step by step

  1. Define tool authority

    Write the task and whitelist: compare supplied costs, use Calculator, return a draft recommendation. It may not purchase, send or browse.

    Check: The allowed actions and missing tools are explicit.

    Need a hint?

    All prices are fictional sample inputs.

  2. Connect the real agent

    Configure n8n AI Agent with a model that supports tool calling and attach the Calculator sub-node. Set Maximum Iterations to4 for this exercise.

    Check: The node has at least one connected tool and a visible limit.

    Need a hint?

    A chat model working in a basic chain may not support tool calling.

  3. Run both options

    Supply the two cost formulas and budget90. Inspect the calculator input/result and the agent final answer.

    Check: A is78, B is95; A fits and B exceeds budget by5.

    Need a hint?

    Verify the trace shows actual tool output, then check the arithmetic independently.

  4. Inspect the trace

    Record request, tool name, input, result, number of iterations and final answer. Mark which claims came from a tool and which from supplied data.

    Check: A partner can replay why the recommendation was made.

    Need a hint?

    Do not treat hidden reasoning text as a reproducible tool log.

  5. Test unsupported requests

    Ask for current weather, an online purchase and a task that exceeds the configured limit.

    Check: The agent reports a capability boundary or stops; no imaginary execution appears.

    Need a hint?

    Use your actual workflow error/stop route to demonstrate the limit.

  6. Evaluate transfer

    Change quantities and budget, rerun, and compare results against a manually computed answer.

    Check: The agent adapts without a new system prompt.

    Need a hint?

    Save your own rubric for math correctness, tool use and unsupported actions.

Build with a clear contract

You may use only a connected Calculator tool. Compare optionA=3*18+2*12 and optionB=4*20+1*15 against a fictional budget90. Return each total, budget difference, tool evidence and a draft recommendation. Do not purchase, browse or send messages. If a tool is unavailable or the task requires another capability, say so. Stop within4 iterations. A human reviews the recommendation.

When it goes sideways

It answers math without a tool trace

The agent narrated a calculation rather than using the connected tool.

Try: Make exact arithmetic tool use part of the rubric and inspect execution data.

No tool call occurs

The model or configuration lacks tool-call support.

Try: Choose a documented supported model and verify the Calculator connection.

It loops or claims unavailable actions

Scope and iteration limits are missing.

Try: Set the limit, narrow authority and test an unavailable capability.

Review your evidence

Tick a criterion only after checking your own artifact. These are self-reported checks, not an automated certification.

Make it your own

Add an approved-note lookup

Add a read-only notes tool before the calculator. Keep its corpus small and require source IDs. Compare lookup failures with arithmetic failures.

  • Only approved notes are reachable.
  • Unknown facts remain unknown.
  • Tool budget and error paths are demonstrated.

Check the mental model

What proves calculator use?

A weather question arrives with only Calculator connected.

Remember the distinction

Tool choice is the distinguishing step

Tool output is evidence

A stopping rule makes failures inspectable

Go to the source

Original community projects. Interactive scenes are teaching simulations. Tool outputs vary. Your evidence stays on this browser unless you export it.