Your agent isn't slow because of the CPU
By Kingsley Torlowei
An agent with twenty tool calls takes three minutes to answer one question. Run it over a thousand inputs and the instinct is to reach for hardware: more cores, a bigger box, a faster runtime.
None of that touches the three minutes. Look at where they go.
Where the time goes
One turn of an agent is a model call, then a tool call, then a model call again. The model call is a network round trip to someone else's GPU. The tool call is a network round trip to someone else's API. Your process sends a request and waits. Then it sends another and waits.
The CPU is idle for almost all of it.
And the turns can't overlap. Turn 18 needs what turn 17 returned. You can't parallelize a conversation, so the latency of one agent on one input is set almost entirely by your model provider and your tools. Infrastructure doesn't get to change that number.
What infrastructure does control is everything around the turns. That's where the hours actually go.
The expensive part is starting over
The agent fails at turn 18 of 20. A rate limit, a tool that timed out, a malformed response. The retry starts at turn 1.
So you pay for turns 1 through 17 twice: twice the wall clock, twice the tokens. And the second run isn't even the same run. The model is non-deterministic, so the new turns 1 through 17 can take a different path to turn 18, or never reach it. You aren't retrying the failure. You're rolling the dice again.
The fix is not a faster turn. It's never running a finished turn again.
Wrap each model call and each tool call as a step, and every step is checkpointed when it returns:
import papayya
from papayya import agent
@agent(name="research")
def research(run, task: dict) -> dict:
messages = [{"role": "user", "content": task["question"]}]
for _ in range(30):
reply = run.step("model", call_model, item_id=task["id"])(messages)
messages.append(reply["message"])
if not reply["tool_calls"]:
return {"answer": reply["message"]["content"]}
for call in reply["tool_calls"]:
result = run.step(f"tool:{call['name']}", run_tool, item_id=task["id"])(call)
messages.append({"role": "tool", "content": result})
papayya.mark_degraded("hit the turn limit without an answer")
When that item fails at turn 18 and you re-drive it, turns 1 through 17 are read back instead of run again. They hand back exactly what they returned the first time, so the conversation that reaches turn 18 is the one that failed there. Not a new sample.
What it saves, measured
We measured it on a document pipeline from our examples: 40 pages, each read by a step that takes 2.5 seconds, about one vision-model call. Page 17 fails during an outage. Once the outage closes, we re-drive it. Then we run the same document from scratch for comparison.
| Wall clock | Steps run | Steps reused | |
|---|---|---|---|
| First pass, fails at page 17 | 68.0s | 16 | 0 |
| Re-drive | 60.3s | 24 | 16 |
| From scratch | 100.6s | 40 | 0 |
The re-drive finishes 40 seconds sooner than starting over. That is exactly the 16 pages that already worked. Skipping them took 0.02 seconds. Across the whole re-drive, a third of a second went to anything other than the pages themselves.
The saving is proportional to where the failure lands. Page 17 of 40 saves 40%. Page 39 of 40 saves 95%.
One caveat, stated plainly: the 2.5 seconds in that measurement is a sleep, not a provider call. So these are seconds, not dollars. The dollars follow the same split, because a reused step makes no model call at all.
Retry what's transient
Some failures do go away on their own: a provider blip, a rate limit, a dropped socket. Those should never reach a person. A step that raises is retried before the item fails: five attempts, waiting 1, 2, 4 and 8 seconds between them.
That has a price when the failure doesn't go away. In the measurement above, the first pass took 68 seconds for 40 seconds of pages. The other 28 were page 17 being retried against an outage that was still open. For work that cannot succeed on a second try, say so, and it fails on the first:
run.step("charge-card", charge, retries=0)(order)
A retry is not a re-drive. Retries happen inside one attempt at an item. When they run out, the item fails and waits for you. Re-driving it is what reuses the turns that worked.
Slow is not dead
A forty-minute agent and a stuck agent look the same from outside: nothing has come back yet.
Every item holds a lease, and a heartbeat from a separate process renews it for as long as the item is working, however slow each turn is. If the worker dies, the heartbeats stop, the lease lapses, and another worker resumes the item from its last checkpoint.
A stuck item gets a ceiling instead. Thirty minutes by default; declare anything up to 24 hours on the agent:
@agent(name="research", max_duration_seconds=3 * 3600)
The ceiling is enforced by a signal, not a thread, so it still fires when a step is holding Python's GIL. That's exactly when a heartbeat thread would have gone quiet.
A thousand of them at once
This is why a slow agent is mostly a throughput problem, not a latency one. Each worker runs one item at a time, and throughput is the number of workers. Six four-second items on three workers finish in about eight seconds, in two waves of three. A thousand three-minute items don't take fifty hours. They take fifty hours divided by the pool.
Running everything at once isn't always what you want, though. One customer's thousand items can starve everyone else's, or push one API key past its limit. So cap it per tenant:
@agent(name="research", concurrency_per_key=2, rate_limit="30/min")
At most two of one tenant's items run at a time, and at most thirty start in any minute. Other tenants' items run around them. A slot is held by the heartbeat, not by a counter, so when the worker running a capped item dies, the slot comes back on its own. We killed one mid-item. The tenant's next item started 59.9 seconds later.
What we haven't built
Said plainly, so you don't build on it:
- Scheduling around your provider's rate limit. The caps above are numbers you choose. We don't yet adjust how many items run from your provider's actual token budget. Today a 429 is retried like any other error, and a run that saturates your key will see more of them.
- Parallel tool calls inside one turn. If a turn makes three independent tool calls, running them at the same time is up to your code or your framework. Each one still goes through its own step.
The line
We don't make a turn faster. Nobody but your model provider can.
We make sure you never pay for one twice.
The full reference, with the retry settings, the ceiling and the caps, is in Long-running agents.