Learn

AI Cost Per Turn: What's a Good Number — and When It's a Leak

If you run an AI agent where your team already works — Slack or Microsoft Teams — the most useful cost metric isn't tokens, licenses, or monthly spend. It's AI cost per turn: what one request-and-response actually costs, including all the reading, tool use, and checking the agent did to answer it. Ask ten vendors what a good cost per turn is and you'll get ten shrugs. Here's what we see running agents in production every day.

A "turn" is the unit of work a human actually asked for

One message in, one completed answer out. That might be a two-line reply, or it might include pulling a report, updating a system, and verifying the result. Per-token pricing tells a manager nothing — tokens are an engineering detail. Cost per turn maps to something real: "we asked for 40 things today and it cost $11." It's the cost-per-click of agent operations.

The benchmarks we actually use

From our own deployments, measured against provider billing — not estimated:

  1. Simple answers: cents. A question the agent can answer from what it already has in front of it should land under a dime. If small talk costs real money, something is wrong.
  2. Working turns: roughly $0.10–$0.60. Pull live numbers, draft a document, file a task, make a small site edit. This is the healthy middle of everyday operations.
  3. Heavy turns: $1–$3, when you chose them. A long analysis, a big build, a deep audit. Expensive turns are fine when they're deliberate — one big ask replacing hours of work is a bargain.
  4. The alarm line: sustained $1.50+ on ordinary turns. When routine conversation starts costing what heavy work should, you don't have an expensive model — you have a leak. This is the threshold our own monitoring alerts on.

What actually drives cost per turn (it's not the model)

The silent killer is context. An agent that keeps its memory in the chat itself re-reads an ever-growing scrollback on every single turn — so the fiftieth message costs several times the fifth, for the same quality of answer. We watched one deployment drift to $2.83 per message on routine turns. Nothing was wrong with the model; the conversation had become the database. Moving working memory into living status documents and resetting the chat dropped ordinary turns back to cents. That pattern is documented in Slack History → Living Status Docs and the broader fix in AI Cost Optimization: Smaller Scopes, Smaller Bills.

How to get to a good number

  1. Measure it. Per-turn cost from real usage data, reconciled against the provider's bill — not vibes. You can't manage a number you're estimating.
  2. Move memory out of the chat. Decisions and status live in documents the agent reads on demand, not in scrollback it re-reads forever.
  3. Make fresh starts free. When memory lives in documents, resetting a conversation costs nothing — so long threads never get expensive in the first place.
  4. Set an alarm, not a vibe. Pick your threshold and have monitoring announce breaches where the team already works, in Slack or Teams. Leaks get caught in hours, not at invoice time.

A good cost per turn isn't a single number — it's cents for small things, real money only when you deliberately spend it, and an alarm that fires when those two drift together.

Find where AI fits in your operation

Tell us how your team works today, and we'll show you the first workflow worth automating.

Tell us about the operation

About you
About the business
About the opportunity

Your details are sent to Big Timber and stored so we can respond. We read every one and reply personally.