When using an LLM to do things for you, there's always a trade off between just getting a one-off job done and investing a little extra thought to have a tool created that can repeatably and reliably do that job more than once.
With "just one-off", you prompt, review results, and revise or re-prompt until you get to the answer you want. With "create a tool" you have to think a little more clearly about the inputs, desired outputs, and the corner and edge cases along the way.
But what is the cost difference between getting the actual work done via a built tool versus letting agents iterate among themselves?
Below is a simple example of the creation/execution cost difference of having a computer play Fizz-Buzz with itself in two ways:
- Two Claude subagents play among themselves with the parent agent as orchestrator
- a multi-threaded python program
Original prompt
spin up two sub agents. each agent should be able to respond to a message following this pattern: a constant “count:” followed by an integer. Each agent should, when receiving such a message, increment the integer and send “count:” concatenated with the new integer to the other agent. When an agent receives a message that is a multiple of 5, it should also communicate back to you the string “buzz” and you should display that to me, including whether it’s the first or second agent. When an agent receives a message that’s a multiple of 7, it should send you the string “fizz”. When an agent receives a message that’s a multiple of both 7 and 5, it should send the string “fizz-buzz” to you. After spawning the two agents, pick one and send it “count:0”. an agent receiving a “count:” message with an integer greater than 255 should respond to you with “I am done” and should not message the other agent.
Setup
- Two background subagents, FIRST and SECOND, ran on
claude-haiku-4-5. - Each agent got a
count:Nmessage, reported back to main, then sentcount:N+1to its peer:- multiple of 5 and 7 →
fizz-buzz - multiple of 5 →
buzz - multiple of 7 →
fizz - N > 255 →
I am done, and stop
- multiple of 5 and 7 →
- Main sent
count:0to FIRST. That meant FIRST handled even numbers and SECOND handled odd numbers.
Startup issues
- Both agents finished their first turn (“ready”) before the peer-ID messages arrived. The messages had been queued “at next tool round”, and the agents made no tool calls.
- Main had to resend
peer:<id>to SECOND andcount:0to FIRST.
Results
| Signal | FIRST (even) | SECOND (odd) |
|---|---|---|
| fizz-buzz | 0, 70, 140, 210 | 35, 105, 175, 245 |
| buzz | every multiple of 10 from 10 to 250 | every odd multiple of 5 from 5 to 255 |
| fizz | 14, 28, 42, 56, 84, 112, 126 | 7, 21, 49, 63, 77, 91, 119, 133, 147, 161, 189, 203, 217, 231 |
| done | I am done on 256 | none |
Accuracy
- SECOND: 100%. It sent every report it should have.
- FIRST: It never reported fizz for 8 multiples of 7: 98, 154, 168, 182, 196, 224, 238, 252.
- It reported all 7 of its fizzes from 14 through 84, then got only 2 of the next 10 (112 and 126).
- Its buzz and fizz-buzz reports and its forwarding were all correct. The count never broke.
- Termination: correct. FIRST received 256, reported
I am done, and did not forward.
Cost
Token usage comes from the subagent transcripts, deduplicated by API message ID.
| Agent | API calls | Input | Output | Cache read | Cache write |
|---|---|---|---|---|---|
| FIRST | 262 | 2,362 | 38,030 | 22,071,390 | 143,262 |
| SECOND | 262 | 2,360 | 35,192 | 15,329,323 | 100,177 |
Prices used for Haiku 4.5: $1/MTok input, $5/MTok output, $0.10/MTok cache read, $2/MTok cache write (1-hour TTL).
| Agent | Cost |
|---|---|
| FIRST | ~$2.69 |
| SECOND | ~$1.91 |
| Total | ~$4.60 |
- At the 5-minute cache-write rate ($1.25/MTok), the total would be about $4.41.
- About 80% of the cost was cache reads. Each agent’s context grew with every turn and was re-read on every call.
- Session total (from
/usage): $6.28. If that total includes the subagents, the main orchestrator (Opus 5.5) cost about $1.68. - FIRST’s wall-clock run time was about 11.4 minutes.
Comparison: Python program
The program is ~/fizzbuzz_agents.py. It runs two threads with queues: each thread is an “agent” with an inbox, it forwards count:N+1 to its peer, and it reports to a main queue using the same rules.
An independent checker compared the program’s output with the 81 expected reports for 0 to 255.
| Metric | Subagents (Haiku 4.5) | Python program |
|---|---|---|
| Correct reports | 73 / 81 (8 fizzes missed by FIRST) | 81 / 81 |
| Termination at 256 | Correct | Correct |
| Startup glitches | Yes: resent the peer IDs and count:0 | None |
| Game duration | ~11.4 minutes | 7.4 ms (0.12 s process wall) |
| Cost to run | ~$4.60 (agents only) | $0 (local CPU) |
| Cost to build | Included in the session’s $6.28 with an upper bound of ~$1.68 | ~$0.26 (/usage $6.28 → $6.54) |
| Deterministic | No | Yes |