Learn

Multiple AI Agents vs One: What Actually Bloats an Agent (Hint: It's Not Knowledge)

Sooner or later, every team running an AI agent in Slack or Microsoft Teams asks the same question: should we be building multiple AI agents vs one agent that just keeps learning more? The instinct behind the question is reasonable. You watch an agent accumulate skills, procedures, tool connections, and project history, and you assume it works like a person — at some point it gets overloaded, spread thin, bloated with knowledge.

The instinct is half right. Something does bloat. But it isn't knowledge — and understanding what actually gets heavy is one of the cheapest lessons in running AI, because the wrong answer to "one agent or many?" either multiplies your maintenance burden for nothing or quietly multiplies your bill.

First, understand what the meter charges for

The meter runs on reading, not thinking. Every time an agent speaks, it first re-reads everything in its head: its standing instructions, the instruction manuals for every tool it's connected to, and the entire conversation so far. AI providers charge per word read (input) and per word written (output), and for a working agent the reading side dominates by a wide margin. An agent that writes a three-sentence reply may have re-read the equivalent of a short novel to produce it.

That single fact reframes the bloat question. The right thing to ask isn't "how much does this agent know?" It's "how much does this agent carry into every single turn?"

Knowledge is not the bloat

An agent's skills, procedures, project notes, and institutional memory live in files — like books on a shelf. A library of five hundred playbooks costs exactly the same as a library of five: nothing, until a task actually opens one. A well-built agent scans a compact index of what it knows, pulls the one relevant procedure when the work calls for it, and leaves the other four hundred ninety-nine on the shelf.

This means an agent can keep learning indefinitely without getting heavier. Every correction it's given, every workflow it masters, every client quirk it records — written to a file, indexed, and dormant until needed. Teams that split agents because one "knows too much" are solving a problem that doesn't exist, and paying for it twice: two agents to maintain, and a knowledge base that now lives in two places and drifts apart.

What actually bloats an agent

Three things get packed into the agent's head on every turn, whether or not they're useful for the message at hand:

  1. Tool manuals ride along. Every tool the agent is connected to ships its instruction manual into the agent's head at the start of every conversation — used or not. Connect forty tools and the agent re-reads forty manuals before answering "what time is the standup?" This is the hidden tax of plug-in convenience: each new connection makes every conversation slightly more expensive, forever.
  2. Payloads pay rent. Ask a plugged-in tool for data and the raw response — often an enormous blob — lands directly in the conversation. The agent then re-reads that blob on every later turn. You don't pay for data once; you pay rent on it for the life of the conversation. Providers do discount re-reads of text they've already seen (a "cache hit," typically around a tenth of fresh price), but discounted baggage is still baggage, and the discount vanishes the moment the opening stack changes — swap one tool in the lineup and the whole stack re-reads at full price.
  3. Scrollback accumulates. A long-running channel conversation is a growing pile of words sitting in front of every new answer. Left unmanaged, an agent in a busy channel ends up re-reading weeks of chatter to answer one question.

Notice what's absent from the list: skill count. Knowledge in files appears in context only when invoked. Tools and conversation history appear always. That asymmetry is the entire answer to the bloat question.

Heavy data belongs outside the conversation

The worst version of payload rent is using the conversation itself as a data pipeline. The fix is an old idea: send a runner on an errand. A small script — running outside the agent's head — fetches the ten thousand rows, does the math, and hands back five lines. Only the five lines ever enter the conversation. The ten thousand rows never touch the meter. Better still, a scheduled script needs no AI at all: a morning report, a threshold alarm, a data pull — zero tokens, every single run.

In one real deployment, a busy Slack channel crept to $2.67 per agent reply — almost all of it re-read scrollback and payload baggage rather than new work. The alarm that caught the problem is itself a scheduled script, which costs nothing. In the same shop, a three-year advertising analysis pushed thousands of rows through scripts and spreadsheets outside the conversation; only the distilled summary ever entered the agent's context. The rule that falls out: data-shaped work goes through scripts, judgment-shaped work goes through the conversation, and everything gets distilled before it enters chat.

So: one agent or many? The three-tier answer

  1. More knowledge → same agent, more files. Skills, procedures, and memory scale essentially for free. If the motivation for splitting is "it knows too much," don't split. Write things down, index them well, and let the shelf grow.
  2. More parallel or heavy work → temporary subagents. A capable agent can spin up throwaway workers on demand: each gets a narrow task, a clean head with only the context that task needs, and a finish line. They do the work, report back a distilled result, and disappear. The orchestrating agent never carries their payloads, and nobody maintains a standing fleet. This is how one agent runs five research threads or builds three design candidates at once without any conversation getting fat — a team, rented by the minute.
  3. A different tool lineup → a separate standing agent. This is the one case where a permanent second agent genuinely pays. When a workload needs its own set of tools all day — a reporting agent wired into ad platforms and analytics, versus an operations agent wired into project management and a website builder — splitting wins twice. Each agent carries fewer manuals per turn, so every conversation in both lanes gets cheaper. And each agent stays scoped to what it's for, which is also the right security posture: the reporting agent physically can't touch the website, and one customer's agent can't see another customer's systems. Deployments scale in agents; knowledge never needs to.

A compact way to hold it: knowledge scales in files, workload scales in temporary subagents, tool lineups scale in separate agents. If you're creating standing agents to solve a knowledge problem, you're reorganizing the library by buying more buildings.

What this means for your deployment

When you evaluate an AI agent setup — or a vendor pitching one — ask three questions. What rides into every turn (how many tools, how big the manuals)? Where does heavy data get crunched (inside the conversation, or outside in scripts)? And what's the splitting rule (are new agents created for tool scope and isolation, or as a band-aid for bloat)? Teams with good answers run lean agents that know enormous amounts. Teams without them run expensive agents that know very little — they just carry a lot.

An agent doesn't get bloated by what it knows — only by what it carries. Keep the shelf full and the backpack light.

Find where AI fits — one agent or many.

Not sure whether your work needs one agent, a swarm of temporary ones, or just a well-placed script? We map that in one working session. Find where AI fits.

Tell us about the operation

About you
About the business
About the opportunity

Your details are sent to Big Timber and stored so we can respond. We read every one and reply personally.