AI Frontiers, part 31: AutoGPT and the first agent wave
Part 31from the AI Frontiers series · 65 parts in all
In late March 2023 a developer going by Significant Gravitas published a Python script that gave a language model a goal, a small set of tools — search the web, read and write files, run commands — and a loop that let it keep going. AutoGPT became the fastest-growing repository in GitHub's history by some measures, spawned dozens of derivatives within weeks, and generated a volume of discourse about autonomous agents that the underlying technology could not possibly support.
By the summer the tone had changed. Practitioners who tried to use it for anything real reported the same problems: the agent looped, forgot its objective, burned money on repeated searches, and produced confident summaries of work it had not done. The honest assessment of that period is that the first agent wave was a genuine and correct insight attached to a technology that was not ready — and that the failures revealed the specific engineering problems that the next two years were spent solving.
What the loop actually was
Stripped of the branding, AutoGPT was a while-loop with a prompt. Given a goal, the model was asked to produce reasoning, a plan and a next action. The action was parsed from the output and dispatched to a tool. The result was appended to the context. Then the process repeated.
That is the same architecture as the function-calling primitive with the human removed from the middle, and the removal is the entire story. With a person in the loop, every iteration gets a judgment: is that the right next step, is that result credible, is this going anywhere. Autonomous operation requires the model to supply that judgment itself, and in 2023 it could not reliably do so.
Contemporaneous research had already established the pattern that AutoGPT was attempting. ReAct interleaved reasoning with action; Reflexion (Shinn et al.) added a mechanism for the agent to reflect on failures in natural language and carry the lesson into the next attempt, which measurably improved performance; Self-Refine (Madaan et al.) showed a model could improve its own output through iterative critique. Generative Agents (Park et al.) built a small simulated society of language-model characters with memory, reflection and planning, and it remains the most compelling demonstration that the architecture can produce coherent long-horizon behavior — in a sandbox where failure costs nothing and the environment is entirely controlled.
The four failures, in the order they appeared
The problems were consistent enough across implementations that they can be listed as a taxonomy, and each one maps to a piece of the stack that had to be built afterwards.
Context exhaustion. A loop appends every thought, action and observation to the conversation. Within a few dozen steps the context is full, and the oldest content — the original goal, the constraints, the plan — is the first thing to be summarized away or dropped. Agents forgot what they were doing because the transcript had become a record of everything except the objective. The fix is context management as a first-class subsystem: external memory, explicit task state stored outside the transcript, and summarization that preserves decisions rather than text.
Looping and thrashing. Without a cost model, the agent had no reason not to repeat a search it had already run or reread a file it had already read. Detecting that a state has been visited before requires either a memory of actions or a critic that notices the repetition, and neither existed. The fix is a notion of progress: track what changed, and treat no-change as a signal to escalate or stop.
Compounding error. Every step has some probability of being wrong, and the probability that a long trajectory is correct is the product. This is the arithmetic from the GUI-agent entry appearing a year early in a different costume, and it is the fundamental reason autonomous long-horizon work is hard. The fix is verification at each step — tests, checks, deterministic assertions — rather than trusting the trajectory.
Confident fabrication of success. The worst failure mode, and the one that generated the most disillusionment: an agent that had done nothing useful would write a report describing the work as complete. This follows directly from the preceding entry's subject. A model trained to produce plausible continuations, asked to summarize its progress, produces a plausible summary of progress. The fix is that the summary must be grounded in artifacts that can be checked — files that exist, tests that passed, URLs that resolve.
The evaluation problem arrived with the agents
A subtler legacy of this period is that it forced the field to think about how to measure an agent at all, and the first attempt was instructive. AgentBench (Liu et al.) ran language models as agents across eight interactive environments — operating systems, databases, knowledge graphs, games, web browsing — and produced two findings that shaped everything after.
First, the gap between models was much larger on interactive tasks than on static benchmarks, and even strong models failed at a substantial fraction of environments. Second, the dominant failure mode was not lack of knowledge but the inability to sustain a coherent plan across many interactions. A model that can answer a question about how to use a database tool and a model that can actually use one are different things, and the difference is planning discipline rather than information.
The methodological lesson is worth more than the results. An agent benchmark has to specify an environment, a task, a success condition and a budget, and the choice of budget changes the ranking. Every subsequent evaluation debate — how many attempts, what counts as success, whether cost is part of the score — started here, and the field would take another two years to start reporting cost alongside accuracy.
Memory, which the first wave reinvented by accident
The single most useful technical idea to emerge from this period was a workaround. AutoGPT kept a file for the goal and a file for the plan, and wrote notes to disk between iterations, because it had no other way to survive context exhaustion. BabyAGI maintained a task queue and a list of completed tasks, re-prioritizing them each cycle with a model call. Neither was designed as a memory architecture; both were consequences of a limit.
The research literature supplied the principled version. Generative Agents (Park et al.) gave each character an append-only memory stream of observations, and retrieved from it by scoring each entry on recency, importance and relevance to the current situation, then reflecting periodically to synthesize higher-level conclusions. That combination — an append-only log, a retrieval function over it, and a periodic summarization step — is now the standard shape of agent memory, and it is worth noticing that it is precisely the retrieval architecture of part 5 applied to a private log rather than a document corpus.
Two design principles follow, and they have held up. Memory belongs outside the context window, because anything inside it can be evicted by the next long input. And what gets retrieved should be selected by relevance to the current step rather than by recency alone, because an agent working through a task needs the constraint it was given ten steps ago more than it needs the observation from ten seconds ago.
Cost, which nobody budgeted for
The failure that did the most to end the enthusiasm was economic. An autonomous loop re-sends its whole trajectory on every iteration, so the token count grows quadratically with the number of steps. Thirty iterations of a moderately long transcript is hundreds of thousands of tokens, and in 2023 that was real money for a script that might be accomplishing nothing.
Cost is not merely a constraint on agents; it is a design input, and the first wave demonstrated why. A budget forces a stop condition, which forces the agent to decide what to do with an incomplete result, which forces an honest progress assessment. Systems with a hard budget behave better than systems without one, not because the model gets smarter but because the architecture has to answer questions the unbounded version can ignore.
The corollary that took the field longer to accept is that cost has to be part of the score. An agent that succeeds at ten times the price of a simpler approach is not obviously the better system, and an evaluation that reports only success rates will recommend it anyway. The benchmark paper from this period (Liu et al.) measured capability with a generous budget; the cost-aware critique of agent benchmarks arrived a year later, and it was right.
What the wave got right
The skepticism that followed was justified and slightly excessive, because three of the things the first agent wave believed are now mainstream.
Tool use is the unlock. AutoGPT's users were watching a language model take actions in the world, and the fact that the actions were frequently unhelpful did not change the fact that this is the architecture. Everything that works today is the same loop with better discipline.
Decomposition helps. Splitting a goal into subtasks and working them one at a time, which the frameworks pursued through planner-executor designs, is exactly what the planning work of 2024 and 2025 formalized. The MetaGPT (Hong et al.) and ChatDev (Qian et al.) experiments assigning roles to multiple model instances were crude, and they were pointed at a real property: different prompts elicit different competences, and separating them improves the result.
Memory has to be external. The most durable technical inheritance from this period is that agent state belongs outside the context window. Task lists, file systems, databases and vector stores as memory are now unremarkable, and the reason they are unremarkable is that the first wave demonstrated what happens without them.
The claim worth being careful about is autonomy. The first wave assumed that removing the human was the goal. The second wave, arriving two years later, concluded that the human is a component — placed where judgment is cheapest and the cost of error is highest — and that autonomy is a property you negotiate per task rather than a setting you turn on. That is a less exciting position and a much better engineering one, and arriving at it cost a spring of enthusiastic blog posts and a summer of disappointed issues on GitHub.
Reading the loop as a specification
There is a more generous way to read this period than as a cautionary tale about hype. Every component that a working agent needs was identified in 2023, by people building the broken version and noticing which part broke.
Context exhaustion identified memory management as a subsystem. Looping identified progress tracking as a subsystem. Confident fabrication identified grounding as a subsystem. Unbounded cost identified budgeting as a subsystem. Compounding error identified verification as a subsystem. A modern agent platform has all five, and the reason it has them is that the first wave ran without them and failed in ways that were specific enough to name.
The design principle that ties them together is the one the second wave adopted and the first wave rejected: put the human where judgment is cheapest and the cost of a mistake is highest, and let the machine do the volume work in between. Not because models cannot be trusted with judgment, but because the arithmetic of compounding error means a long autonomous trajectory will fail eventually, and the design question is whether that failure is cheap and visible or expensive and silent.
There is also a hiring lesson buried in the period that has nothing to do with models. The teams that made progress were the ones who treated the agent as a distributed system with an unreliable component, and instrumented it accordingly: logs of every action, replayable trajectories, and a hard budget. The teams that made progress fastest were small, because a system whose failure modes are not yet understood benefits enormously from everyone being able to read the same trace.
Everything in the agent entries that follow this one — the coding agents, the browser agents, the protocols, the observability tooling — is an answer to one of those five problems. The first wave is worth remembering not as an embarrassment but as the only public experiment that could have produced that list so quickly, and it produced it by being enthusiastically wrong in a dozen instructive ways at once.
Works Cited
Hong, Sirui, et al. "MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework." arXiv, 2023, arxiv.org/abs/2308.00352. Accessed 7 Sept. 2023.
Huang, Jiaxin, et al. "Large Language Models Can Self-Improve." arXiv, 2022, arxiv.org/abs/2210.11610. Accessed 7 Sept. 2023.
Liu, Xiao, et al. "AgentBench: Evaluating LLMs as Agents." arXiv, 2023, arxiv.org/abs/2308.03688. Accessed 7 Sept. 2023.
Madaan, Aman, et al. "Self-Refine: Iterative Refinement with Self-Feedback." arXiv, 2023, arxiv.org/abs/2303.17651. Accessed 7 Sept. 2023.
Nakajima, Yohei. "BabyAGI." GitHub, 2023, github.com/yoheinakajima/babyagi. Accessed 7 Sept. 2023.
Park, Joon Sung, et al. "Generative Agents: Interactive Simulacra of Human Behavior." arXiv, 2023, arxiv.org/abs/2304.03442. Accessed 7 Sept. 2023.
Qian, Chen, et al. "Communicative Agents for Software Development." arXiv, 2023, arxiv.org/abs/2307.07924. Accessed 7 Sept. 2023.
Shinn, Noah, et al. "Reflexion: Language Agents with Verbal Reinforcement Learning." arXiv, 2023, arxiv.org/abs/2303.11366. Accessed 7 Sept. 2023.
Significant Gravitas. "AutoGPT." GitHub, 2023, github.com/Significant-Gravitas/AutoGPT. Accessed 7 Sept. 2023.
Sumers, Theodore R., et al. "Cognitive Architectures for Language Agents." arXiv, 2023, arxiv.org/abs/2309.02427. Accessed 7 Sept. 2023.
Wang, Lei, et al. "A Survey on Large Language Model Based Autonomous Agents." arXiv, 2023, arxiv.org/abs/2308.11432. Accessed 7 Sept. 2023.
Wang, Guanzhi, et al. "Voyager: An Open-Ended Embodied Agent with Large Language Models." arXiv, 2023, arxiv.org/abs/2305.16291. Accessed 7 Sept. 2023.
Yao, Shunyu, et al. "ReAct: Synergizing Reasoning and Acting in Language Models." arXiv, 2022, arxiv.org/abs/2210.03629. Accessed 7 Sept. 2023.