AI Frontiers, part 47: Multi-agent systems — when decomposition helps and when it hurts
Part 47from the AI Frontiers series · 65 parts in all
After AutoGPT demonstrated both the appeal and the dysfunction of a single model in a loop, the natural next move was obvious to everyone building at the time: if one agent is unreliable, use several. Give each one a role, a persona, and a turn-taking protocol. Let them critique each other. The 2023 research wave — CAMEL, MetaGPT, AutoGen, ChatDev, Generative Agents — produced frameworks for exactly this, and the demos were striking. A "software company" of generated personas could produce a plausible application from a one-line brief; a village of simulated characters could sustain days of emergent social behavior (Li et al.; Hong et al.; Wu et al.; Qian et al.; Park et al.).
Eighteen months later, the honest scorecard is mixed in a way that is more interesting than either the enthusiasm or the backlash. Multi-agent decomposition genuinely helps for some problem shapes and reliably hurts for others, and the boundary between them is now reasonably well understood. This entry is about drawing that boundary, because I think most production teams in 2025 will end up with something simpler than the frameworks advertise and more capable than a single prompt.
The case for decomposition, and why it is mostly about context
The strongest argument for multiple agents was never about intelligence; it is about context management. A single agent solving a long task accumulates a transcript containing everything it has tried, every tool result, and every intermediate reasoning step. That transcript grows until it crowds out the information that matters — the context window is finite, and attention over a long transcript is not uniform. Splitting a task into subagents lets each one work in a fresh, small context: the researcher reads twenty documents and returns two pages; the writer never sees the twenty documents. That is a real architectural benefit and it exists independently of any claim about emergent collaboration.
The second argument is specialization of interface. A subagent that only ever calls a search tool can be given a narrow, well-tested tool contract and a tightly scoped prompt; the blast radius of a prompt-injection attempt or a malformed tool call is contained. This is where multi-agent design overlaps with the security discussion in part 35: decomposition is a containment strategy, not just a capability strategy. An agent with read-only access to the wiki, invoked by a coordinator that holds the credentials, is a better security posture than one agent with all the permissions.
The third argument, and the weakest, is the one the frameworks led with: that debate between personas improves reasoning. The evidence for that is much thinner than the demos suggested.
Where debate helps, and where it does not
The cleanest experiment on this came in 2024, when Wang and colleagues separated the two things multi-agent debate changes: the number of samples drawn from the model, and the interaction between them. Their finding, across reasoning benchmarks, was that most of the improvement attributed to debate could be reproduced by simply sampling and majority-voting, with no agents, no roles, and no discussion. Discussion added little beyond the ensemble effect on the reasoning tasks they tested (Wang et al., "Rethinking the Bounds of LLM Reasoning"). This is consistent with the self-correction literature: a model asked to verify its own answer without any external signal tends to agree with itself, and asking it twice in different personas does not supply the missing signal (Huang et al.; Stechly et al.).
Where interaction does help is where the collaborators have asymmetric information or tools. A critic with access to a test suite, a second model with a different training distribution, a verifier that can actually execute the proposed plan, or a human with domain knowledge — these add information rather than opinion, and the multi-agent framing is a reasonable structure for organizing them. Zhang and colleagues' social-psychology study of agent collaboration found the same asymmetry: conformity effects appear in homogeneous groups, and diversity of perspective is what produces gains (Zhang et al.). A debate between two instances of the same model at the same temperature is a coin flip with extra steps.
The replication of that result in the planning domain is equally instructive. Valmeekam and colleagues found that LLM planners produce plans that are syntactically valid and execution-invalid at high rates, and that adding a critic does not help unless the critic can verify plan steps against real preconditions (Valmeekam et al.). The lesson generalizes: multi-agent debate is a proxy for verification, and it is a strictly worse proxy than actual verification. If you can test the answer, test it. If you cannot, ask whether the extra agents are buying you anything beyond sampling.
The costs nobody put on the slide
Four costs scale with agent count, and they are the reason I have watched several multi-agent prototypes get rewritten back into a pipeline.
Superlinear token spend. Each agent independently loads its system prompt, its context, and often the shared task description. A five-agent system with a debate round can easily cost an order of magnitude more tokens than a single well-prompted call. If your evaluation is pass/fail accuracy and not cost-per-success, you will not notice this until the invoice arrives. Cost-aware scoring, which part 42 argued for, is mandatory here.
Error compounding. If each agent in a chain is 95% reliable and the chain has six steps, end-to-end reliability is under 75% before you count the coordination failures. Multi-agent systems multiply steps, and steps multiply failure probability. The framework demos hide this by choosing tasks where the intermediate outputs are never validated and the evaluator is generous.
Non-determinism in the control flow. "The agents decided among themselves who should take the task" is a sentence that sounds clever and produces systems that cannot be reproduced, debugged, or regression-tested. The moment your architecture contains an unobserved negotiation, you have given up on the single most valuable property of software: the ability to rerun it and get the same answer.
Coordination overhead that grows with the square of the group. Every pair of agents is a potential channel, and every channel is a place for information to be lost, duplicated, or corrupted. AutoGen's contribution was largely acknowledging this by making communication explicit and typed (Wu et al.); teams that skip that step rediscover it painfully.
The pattern that survives
What is left standing after the hype — and it is a genuinely good pattern — is a small number of deterministic orchestration shapes with LLM agents as steps. A coordinator that owns the state and makes routing decisions in code. Worker agents invoked as functions with narrow, typed inputs and outputs. An explicit verification step, ideally executable: run the tests, check the schema, hit the API, compare against a reference. Retries with a budget. And a trace, because the next entry is about how you debug any of this.
In that shape, "multi-agent" is a slightly grand name for what a distributed-systems engineer would call a workflow with pluggable steps. The division of labor is authored by a human; the model fills in the parts that require judgment. That design has three virtues the framework version does not: it is reproducible, its cost is bounded and estimable, and each step can be evaluated in isolation — which is the only way any of this gets better over time. Anthropic's own write-up late in 2024 landed on essentially this conclusion, arguing for simple composable patterns and against framework-driven complexity (Anthropic).
So: decompose for context isolation, for permission boundaries, and for parallelism when subtasks are independent. Do not decompose to simulate a committee. The best multi-agent system I have seen this year had three agents and about four hundred lines of orchestration code, and every one of those lines was doing something the model was not allowed to decide.
Two shapes that actually justify multiple agents
After working through a handful of these systems, I have found exactly two decompositions that pay for themselves consistently. Both are boring, and both replace a conversation with an interface.
Shape one: fan-out with a compression boundary. A task requires reading more material than fits comfortably in one context — forty support tickets, a dozen policy documents, a repository's worth of source files. A coordinator splits the work into independent units, dispatches one narrow agent per unit with a schema-constrained output (see part 45), and then a single synthesis step reads only the returned summaries. The benefit is not collaboration; it is that no agent ever sees more context than it can use well, and the returned objects are small enough to validate. This pattern is embarrassingly parallel, cheap to reason about, and easy to evaluate — each worker's output can be scored against a reference extraction independent of the rest.
Shape two: a critic that can check. A generator produces an artifact — a patch, a schema migration, a query, a pricing calculation — and a second component verifies it with a capability the generator lacks: running the tests, executing the query against a test database, comparing against a reference implementation. This is the case where the debate literature underperforms but the verification literature delivers, because the critic is adding information rather than an opinion. The evaluator can be a model, a program, or a model with a program, and the closer it is to executable ground truth the more reliable the loop. When I do use multiple model agents, this is the structure: one that proposes, one that is only permitted to look for reasons the proposal fails, with the second holding a tool the first does not.
Both shapes share three properties: the number of agents is small and fixed, the control flow is written in code rather than negotiated, and every hand-off is a typed object. If a design has none of those properties, the burden of proof is on the design. A useful test to apply before building anything: name the artifact that moves across each boundary, and write down what would have to be true for a single agent with the same tools to fail at the same task. If the answer is "the context would be too long," or "it would need permissions we do not want to grant," the decomposition is doing real work. If the answer is "it would reason better as a team," you are about to build an expensive sampler.
Unwinding a framework prototype
The most common trajectory I have watched goes like this. A team adopts a multi-agent framework because its primitives make the demo easy: agents with names, a group chat, an automatic speaker-selection policy. The prototype works. Then three things happen in order that are worth recognizing by name. Cost per task stops being predictable, because the number of model calls is now a function of the agents' conversation rather than of the input. Debugging stops being possible, because reproducing a failure requires replaying a non-deterministic dialogue. And improvement stops, because there is no step to attribute a quality change to — a score moved, and nobody knows which of four prompts and two routing decisions caused it.
The rewrite, when it comes, is almost mechanical and takes days rather than weeks. Freeze the graph: write down which agent talks to which, and in what order, based on what the prototype actually does when observed. Replace each conversational hand-off with a function call that takes a typed input and returns a typed output, keeping the prompts that were doing real work and deleting the ones that were role-playing. Move the state into one object the coordinator owns, so that a step is a pure function of the state and the step index. Keep an agent boundary only where a subagent genuinely needs a different tool permission or a fresh context, and delete it everywhere else. Finally, re-run the last two weeks of production traffic through the new pipeline and compare pass rate, cost per task, and steps per task.
In every case I have seen, the rewrite is cheaper, faster, more accurate, and — the part that matters most — measurable. That is not an argument that multiple agents are useless. It is an argument that the valuable version of multi-agent design is a small set of loud boundaries inside ordinary software, and that the framework's group chat was always a convenient way to avoid writing the coordinator.
Works Cited
Anthropic. "Building Effective Agents." Anthropic Engineering, 2024, anthropic.com/engineering/building-effective-agents. Accessed 19 Dec. 2024.
Hong, Sirui, et al. "MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework." arXiv, 2023, arxiv.org/abs/2308.00352. Accessed 19 Dec. 2024.
Huang, Jie, et al. "Large Language Models Cannot Self-Correct Reasoning Yet." arXiv, 2023, arxiv.org/abs/2310.01798. Accessed 19 Dec. 2024.
Li, Guohao, et al. "CAMEL: Communicative Agents for 'Mind' Exploration of Large Language Model Society." arXiv, 2023, arxiv.org/abs/2303.17760. Accessed 19 Dec. 2024.
Park, Joon Sung, et al. "Generative Agents: Interactive Simulacra of Human Behavior." arXiv, 2023, arxiv.org/abs/2304.03442. Accessed 19 Dec. 2024.
Qian, Chen, et al. "ChatDev: Communicative Agents for Software Development." arXiv, 2023, arxiv.org/abs/2307.07924. Accessed 19 Dec. 2024.
Shinn, Noah, et al. "Reflexion: Language Agents with Verbal Reinforcement Learning." arXiv, 2023, arxiv.org/abs/2303.11366. Accessed 19 Dec. 2024.
Stechly, Kaya, et al. "On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks." arXiv, 2024, arxiv.org/abs/2402.08115. Accessed 19 Dec. 2024.
Valmeekam, Karthik, et al. "On the Planning Abilities of Large Language Models: A Critical Investigation." arXiv, 2023, arxiv.org/abs/2305.15771. Accessed 19 Dec. 2024.
Wang, Qineng, et al. "Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key?" arXiv, 2024, arxiv.org/abs/2402.18272. Accessed 19 Dec. 2024.
Wu, Qingyun, et al. "AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation." arXiv, 2023, arxiv.org/abs/2308.08155. Accessed 19 Dec. 2024.
Yao, Shunyu, et al. "ReAct: Synergizing Reasoning and Acting in Language Models." arXiv, 2022, arxiv.org/abs/2210.03629. Accessed 19 Dec. 2024.
Zhang, Jintian, et al. "Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View." arXiv, 2023, arxiv.org/abs/2310.02124. Accessed 19 Dec. 2024.