Friday, October 02, 2026

The Costs of "Prompting Every Time" versus "Build a Tool to Do It"

When using an LLM to do things for you, there's always a trade off between just getting a one-off job done and investing a little extra thought to have a tool created that can repeatably and reliably do that job more than once.

With "just one-off", you prompt, review results, and revise or re-prompt until you get to the answer you want. With "create a tool" you have to think a little more clearly about the inputs, desired outputs, and the corner and edge cases along the way.

But what is the cost difference between getting the actual work done via a built tool versus letting agents iterate among themselves?

Below is a simple example of the creation/execution cost difference of having a computer play Fizz-Buzz with itself in two ways:

  • Two Claude subagents play among themselves with the parent agent as orchestrator
  • a multi-threaded python program
Takeaway: For a fixed, rule-based process like this, a coded tool will give a proven answer and was roughly 90,000× faster and 1/18th the cost in this case.

Writing and running the python cost about $0.26 in total, compared to $4.60 for subagents to play the game. The subagents also dropped the ball about 10% of the time. The subagents' cost grows roughly with the square of the number of turns because each iteration of the game re-reads a growing context.

Subagents only pay off when each iteration needs judgment, not merely arithmetic.

Below is a report Claude wrote about the experiment.

Original prompt

spin up two sub agents. each agent should be able to respond to a message following this pattern: a constant “count:” followed by an integer. Each agent should, when receiving such a message, increment the integer and send “count:” concatenated with the new integer to the other agent. When an agent receives a message that is a multiple of 5, it should also communicate back to you the string “buzz” and you should display that to me, including whether it’s the first or second agent. When an agent receives a message that’s a multiple of 7, it should send you the string “fizz”. When an agent receives a message that’s a multiple of both 7 and 5, it should send the string “fizz-buzz” to you. After spawning the two agents, pick one and send it “count:0”. an agent receiving a “count:” message with an integer greater than 255 should respond to you with “I am done” and should not message the other agent.

Setup

  • Two background subagents, FIRST and SECOND, ran on claude-haiku-4-5.
  • Each agent got a count:N message, reported back to main, then sent count:N+1 to its peer:
    • multiple of 5 and 7 → fizz-buzz
    • multiple of 5 → buzz
    • multiple of 7 → fizz
    • N > 255 → I am done, and stop
  • Main sent count:0 to FIRST. That meant FIRST handled even numbers and SECOND handled odd numbers.

Startup issues

  • Both agents finished their first turn (“ready”) before the peer-ID messages arrived. The messages had been queued “at next tool round”, and the agents made no tool calls.
  • Main had to resend peer:<id> to SECOND and count:0 to FIRST.

Results

SignalFIRST (even)SECOND (odd)
fizz-buzz0, 70, 140, 21035, 105, 175, 245
buzzevery multiple of 10 from 10 to 250every odd multiple of 5 from 5 to 255
fizz14, 28, 42, 56, 84, 112, 1267, 21, 49, 63, 77, 91, 119, 133, 147, 161, 189, 203, 217, 231
doneI am done on 256none

Accuracy

  • SECOND: 100%. It sent every report it should have.
  • FIRST: It never reported fizz for 8 multiples of 7: 98, 154, 168, 182, 196, 224, 238, 252.
    • It reported all 7 of its fizzes from 14 through 84, then got only 2 of the next 10 (112 and 126).
    • Its buzz and fizz-buzz reports and its forwarding were all correct. The count never broke.
  • Termination: correct. FIRST received 256, reported I am done, and did not forward.

Cost

Token usage comes from the subagent transcripts, deduplicated by API message ID.

AgentAPI callsInputOutputCache readCache write
FIRST2622,36238,03022,071,390143,262
SECOND2622,36035,19215,329,323100,177

Prices used for Haiku 4.5: $1/MTok input, $5/MTok output, $0.10/MTok cache read, $2/MTok cache write (1-hour TTL).

AgentCost
FIRST~$2.69
SECOND~$1.91
Total~$4.60
  • At the 5-minute cache-write rate ($1.25/MTok), the total would be about $4.41.
  • About 80% of the cost was cache reads. Each agent’s context grew with every turn and was re-read on every call.
  • Session total (from /usage): $6.28. If that total includes the subagents, the main orchestrator (Opus 5.5) cost about $1.68.
  • FIRST’s wall-clock run time was about 11.4 minutes.

Comparison: Python program

The program is ~/fizzbuzz_agents.py. It runs two threads with queues: each thread is an “agent” with an inbox, it forwards count:N+1 to its peer, and it reports to a main queue using the same rules.

An independent checker compared the program’s output with the 81 expected reports for 0 to 255.

MetricSubagents (Haiku 4.5)Python program
Correct reports73 / 81 (8 fizzes missed by FIRST)81 / 81
Termination at 256CorrectCorrect
Startup glitchesYes: resent the peer IDs and count:0None
Game duration~11.4 minutes7.4 ms (0.12 s process wall)
Cost to run~$4.60 (agents only)$0 (local CPU)
Cost to buildIncluded in the session’s $6.28 with an upper bound of ~$1.68~$0.26 (/usage $6.28 → $6.54)
DeterministicNoYes


Thursday, October 01, 2026

The Third Act Commences

 As of October 16, after just shy of 8 years there, I'll be leaving Zoox to retire.

When I told her, my manager asked me "have you been thinking about this long?". I told her that I've been doing this professionally for 39 years and I've been thinking about it for 38. As Zoox evolves from pure R&D to a live commercial service, the shape of the work needed in my role and the way I work best have slowly diverged. It's best for Zoox to find that right person, and it's best for me to cheer them on from the stands.

I feel very fortunate to be at a place in life where the professional, financial, and familial planets have all aligned in support of this next step. As to what that next step is, it's not terribly concrete beyond:
1. "I've got a bunch of things I've been wanting to do around the house."
2. an 11,000 mile year-long project that I did a 120 mile week-long trial run for in August.

When the project actually happens, my plan is to produce a coffee-table book with a page of words, maps, and pictures for each day's discoveries. There's some video documenting the trial run.

Shameless plug aside, looking back at all my steps and mis-steps along the way, I can't really imagine very much that I would have done differently. I'm grateful to all of those that gave me a chance and supported me (foremost my beautiful and beloved wife). I'm gratified to see where those I gave a chance and supported have gotten to (foremost my beautiful, beloved, and ridiculously talented children).

As a poet once said, "what a long, strange trip it's been" -- time to start another one.