Your goal
Build a bounded tool-using agent and explain its actual execution trace.
Watch an agent choose a calculator, compare options and stop at its limit.
Two fictional setup options cost 3×18 + 2×12 and 4×20 + 1×15. The approved budget is 90. The agent may use a Calculator tool; it has no shopping or weather tool.
Build a bounded tool-using agent and explain its actual execution trace.
Calculate the first option as78 and the remaining budget as12.
An agent workflow, a tool-call trace and a three-case boundary report.
This interactive model teaches the mechanism. It does not call a model, search your files or send messages.
The model can choose a connected tool and use its result. A fixed chain has a predetermined model call instead. Merely writing the word agent in a prompt does not connect capabilities.
The arithmetic result should come from the calculator call and be visible in the trace. A plausible narrated tool call is not proof that an execution happened.
Set a maximum iteration count and an error path. If the task needs an unavailable tool, the agent should explain the missing capability rather than inventing a result.
Write the task and whitelist: compare supplied costs, use Calculator, return a draft recommendation. It may not purchase, send or browse.
Check: The allowed actions and missing tools are explicit.
All prices are fictional sample inputs.
Configure n8n AI Agent with a model that supports tool calling and attach the Calculator sub-node. Set Maximum Iterations to4 for this exercise.
Check: The node has at least one connected tool and a visible limit.
A chat model working in a basic chain may not support tool calling.
Supply the two cost formulas and budget90. Inspect the calculator input/result and the agent final answer.
Check: A is78, B is95; A fits and B exceeds budget by5.
Verify the trace shows actual tool output, then check the arithmetic independently.
Record request, tool name, input, result, number of iterations and final answer. Mark which claims came from a tool and which from supplied data.
Check: A partner can replay why the recommendation was made.
Do not treat hidden reasoning text as a reproducible tool log.
Ask for current weather, an online purchase and a task that exceeds the configured limit.
Check: The agent reports a capability boundary or stops; no imaginary execution appears.
Use your actual workflow error/stop route to demonstrate the limit.
Change quantities and budget, rerun, and compare results against a manually computed answer.
Check: The agent adapts without a new system prompt.
Save your own rubric for math correctness, tool use and unsupported actions.
You may use only a connected Calculator tool. Compare optionA=3*18+2*12 and optionB=4*20+1*15 against a fictional budget90. Return each total, budget difference, tool evidence and a draft recommendation. Do not purchase, browse or send messages. If a tool is unavailable or the task requires another capability, say so. Stop within4 iterations. A human reviews the recommendation.
The agent narrated a calculation rather than using the connected tool.
Try: Make exact arithmetic tool use part of the rubric and inspect execution data.
The model or configuration lacks tool-call support.
Try: Choose a documented supported model and verify the Calculator connection.
Scope and iteration limits are missing.
Try: Set the limit, narrow authority and test an unavailable capability.
Tick a criterion only after checking your own artifact. These are self-reported checks, not an automated certification.
Add a read-only notes tool before the calculator. Keep its corpus small and require source IDs. Compare lookup failures with arithmetic failures.
The model can choose a connected tool and use its result. A fixed chain has a predetermined model call instead. Merely writing the word agent in a prompt does not connect capabilities.
The arithmetic result should come from the calculator call and be visible in the trace. A plausible narrated tool call is not proof that an execution happened.
Set a maximum iteration count and an error path. If the task needs an unavailable tool, the agent should explain the missing capability rather than inventing a result.
Use the primary documentation to verify this part of your build.
Apply it in the project labUse the primary documentation to verify this part of your build.
Apply it in the project labUse the primary documentation to verify this part of your build.
Apply it in the project lab