AI Frontiers, part 14: The agentic coding turn β from autocomplete to teammates
Part 14from the AI Frontiers series · 65 parts in all
In February 2025 Anthropic shipped Claude Code as a research preview: a command-line program that reads a repository, edits files, runs commands, and iterates on the results until a task is done (Anthropic, "Claude 3.7 Sonnet and Claude Code"). It is four years distant from GitHub Copilot's original trick β ghost-text completions appearing as you type β and the difference is not the model quality, although that improved a great deal. The difference is that the assistant now has hands. It can read the error message, change the file, run the tests again, and discover that its first idea was wrong.
That change in shape is what this entry is about, because it did more to close the gap between demo and deployment than any single benchmark number. Once the assistant can observe the consequences of its own action, the evaluation question changes from "does the generated code look right?" to "does the test suite pass?" β which is a question software has known how to answer since before any of us were writing it.
Autocomplete was never the point
The first generation of coding models was constrained by its interface. Codex (Chen et al.) established that a language model fine-tuned on public code could complete functions from their signatures, and HumanEval β the benchmark it introduced β became the field's measuring stick for two years. GitHub Copilot went generally available in June 2022 (GitHub), and the widely cited productivity study found developers completing a well-specified HTTP server task about 56% faster than a control group (Peng et al.).
But completion has a hard budget. To be useful, a suggestion must arrive in a few hundred milliseconds, which caps how much context the model can consume and how much it can produce. It must be a single contiguous insertion, because the editor has to render it as you type. And it has no ability to check anything. Copilot's own designers were candid about this: the feature makes competent developers faster and gives less-experienced developers confident code they cannot evaluate, which is the same observation the security literature made independently (Pearce et al.; Perry et al.).
What completion could not do was finish a task. It could not know that the function it wrote does not compile because a header moved three releases ago, and it certainly could not go find out.
The three things that changed
Long context. Reading a repository is a context problem before it is a reasoning problem, and the progression traced in part 7 β from 4K to 200K tokens in roughly a year β is what made "here is the file, here are the three files that import it, here is the failing test" a single prompt rather than a retrieval engineering project.
Tool use. Function calling (June 2023) gave models a structured way to ask for actions, and MCP (2024) standardized how those actions get delivered. A coding agent's toolset is embarrassingly simple β read a file, write a file, search, run a shell command β and that simplicity is why it works. The interface is the same set of primitives a person uses in a terminal.
Verification loops. This is the piece that people underrate. An agent
that can run pytest and read the output has a ground-truth signal about its own
work. It does not need to be right on the first attempt; it needs to be able to tell that it
is wrong. Every agent loop I have watched succeed in practice is built around this β tests,
type checkers, linters, compilers, build failures. The tooling you already had for humans
turned out to be the supervision mechanism for the agent.
SWE-bench was the reframing
The benchmark that made the turn legible is SWE-bench (Jimenez et al.), which takes real GitHub issues and the pull requests that resolved them, and asks whether a system can produce a patch that makes the repository's own tests pass. It is a much harder and much more honest test than HumanEval, because the repository is real, the fix may span several files, and the grading criterion is external to the model.
When it was published in 2023, the best model resolved roughly 4% of the issues. By October 2024, in the same week Anthropic released the updated Claude 3.5 Sonnet with computer use, it reported 49% on the human-validated subset (Anthropic, "Introducing Computer Use"). By early 2025 the frontier was above 60%. A useful part of that progress was not model intelligence at all: SWE-agent (Yang et al.) showed that the agentβcomputer interface β how the agent sees files, what its edit format is, how errors are surfaced β moved scores substantially with the same underlying model. Which is the same lesson the MCP post made about tool design, arriving from a completely different direction.
What the productivity evidence actually supports
Here I want to be careful, because the gap between the enthusiasm and the literature is wide.
The strongest positive result remains Peng et al.'s controlled experiment, and it measures a specific well-specified task with a specific tool. The DORA research program, which has surveyed delivery performance across thousands of organizations for a decade, found in 2024 that a 25% increase in AI adoption was associated with a modest decrease in delivery throughput and a larger decrease in delivery stability (Google Cloud). GitClear's analysis of hundreds of millions of changed lines found rising code churn and duplication in the Copilot era (GitClear). Stack Overflow's 2024 survey found use rising while trust fell: the majority of developers were using AI tools, and a substantial share actively distrusted their accuracy. None of these results contradict each other. Speed at the keystroke level and stability at the delivery level are different measurements, and it would be strange if generating more code faster did not strain review capacity.
Which is the point I would put in front of any team adopting this in 2025: the constraint moved. Writing code used to be the expensive step and review was cheap. Now generation is nearly free and review is the bottleneck. An agent that opens a 900-line pull request has done you no favor if the review takes four hours and the reviewer approves it because it compiles. The tool that made the last eighteen months interesting and the tool that will make the next eighteen months interesting are not the same tool.
The practical architecture of a coding agent
Strip away the branding and every one of these systems is the same loop. Take a goal and a context window. Decide the next action. Execute it against a sandbox with permissions. Observe the result. Decide whether to continue, and manage the context so that the transcript does not drown the task. The differences that matter are all in the surrounding details.
Permissions determine whether anyone can trust it unattended. Reading is cheap to allow; writing files is medium; running arbitrary shell commands is a different category entirely. The systems I have seen deployed in serious settings ask for the first few writes, learn the pattern, then broaden β and always run in a container or a VM, never against a workstation with production credentials in the environment.
Context management determines whether it can work on anything large. Compaction, summarization and file-selection heuristics are the unglamorous engineering that separates an agent that can fix a bug in a small package from one that flails in a monorepo. This is the real substance of the "does it work on real codebases?" question.
Recoverability determines the cost of being wrong. Version control, fast test suites, and small commits are not just good practice for humans; they are the mechanism by which an agent's mistake costs ninety seconds instead of an afternoon. The teams getting value from these tools in 2025 are, without exception, the ones whose repositories were already easy to work in.
What teams had to change
The adoption pattern across the teams I have watched is consistent enough to write down. The repositories that got real value from coding agents had four properties, and all four were things a thoughtful team would have wanted anyway.
A build and test command that works from a clean checkout. This is the single highest-leverage investment. An agent that can run the test suite has a feedback loop; an agent that cannot is guessing. Every hour spent making the test suite faster and more deterministic pays back in agent reliability, because a flaky test teaches the agent that the world is random and it should give up or, worse, edit the test.
A machine-readable description of the project. Which is why every major tool now looks for a conventional instruction file at the repository root β the same file this site's repository keeps, describing build commands, deployment topology and the rules a contributor must not break. Written well, it is the difference between an agent that reads the conventions and one that invents its own.
Small, reviewable units of change. Agents are not naturally concise. They will happily rewrite a file to fix one line, because they optimize for the task in front of them, not for your review queue. Teams that got value pushed back on this explicitly: ask for the minimal diff, ask for the test that proves it, refuse the reformat that came along for the ride.
Cost visibility. A long agent session over a large repository can consume more tokens in an afternoon than a team did in a month before, and the cost is invisible unless you instrument it. The reasoning models of part 11 make this worse by design. Treating model spend as an engineering metric rather than a line item someone reconciles quarterly is the only way to make sane build-versus-buy decisions about it.
The uncomfortable corollary is that an agent is an amplifier of whatever your process already is. In a repository with a fast test suite, a clear build command and a codebase small enough to hold in context, it is genuinely useful. In one where nobody knows how to run the tests locally, it produces plausible diffs that nobody can verify, which is worse than producing nothing at all β because now the diff exists, and someone is going to merge it.
There is one more change coming that nobody has a good answer for yet, and it is not technical. When a teammate writes code, there is a person who can explain why, and who is answerable when it is wrong. When an agent writes code, accountability has to be assigned by convention. The convention that seems to be forming is that the human who directed and merged the change owns it completely β the agent is a tool, not a party. That is the correct rule, and it is also a rule that will be tested the first time an agent introduces a subtle bug that passes review, because the temptation to say "the model did it" is going to be strong. Teams would do well to decide the norm before the incident rather than after it, which is the same reason security policy is written before the breach and not during it.
Four years from ghost text to something that opens a pull request is fast. What has not changed is the part that was always the actual work: deciding what to build, defining what correct means, and being able to tell the difference.
Works Cited
Anthropic. "Claude 3.7 Sonnet and Claude Code." Anthropic, 24 Feb. 2025, www.anthropic.com/news/claude-3-7-sonnet. Accessed 20 Mar. 2025.
---. "Introducing Computer Use, a New Claude 3.5 Sonnet, and Claude 3.5 Haiku." Anthropic, 22 Oct. 2024, www.anthropic.com/news/3-5-models-and-computer-use. Accessed 20 Mar. 2025.
Chen, Mark, et al. "Evaluating Large Language Models Trained on Code." arXiv, 2021, arxiv.org/abs/2107.03374. Accessed 20 Mar. 2025.
GitClear. "AI Copilot Code Quality: 2024 Research." GitClear, 2024, www.gitclear.com/ai_assistant_code_quality_2024. Accessed 20 Mar. 2025.
GitHub. "GitHub Copilot Is Generally Available to All Developers." The GitHub Blog, 21 June 2022, github.blog/news-insights/product-news/github-copilot-is-generally-available-to-all-developers/. Accessed 20 Mar. 2025.
Google Cloud. "Accelerate State of DevOps Report 2024." DORA, 2024, dora.dev/research/2024/dora-report/. Accessed 20 Mar. 2025.
Jimenez, Carlos E., et al. "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" arXiv, 2023, arxiv.org/abs/2310.06770. Accessed 20 Mar. 2025.
Kapoor, Sayash, et al. "AI Agents That Matter." arXiv, 2024, arxiv.org/abs/2407.01502. Accessed 20 Mar. 2025.
Li, Yujia, et al. "Competition-Level Code Generation with AlphaCode." Science, vol. 378, no. 6624, 2022, arxiv.org/abs/2203.07814. Accessed 20 Mar. 2025.
Pearce, Hammond, et al. "Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions." arXiv, 2021, arxiv.org/abs/2108.09293. Accessed 20 Mar. 2025.
Peng, Sida, et al. "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot." arXiv, 2023, arxiv.org/abs/2302.06590. Accessed 20 Mar. 2025.
Perry, Neil, et al. "Do Users Write More Insecure Code with AI Assistants?" arXiv, 2022, arxiv.org/abs/2211.03622. Accessed 20 Mar. 2025.
Stack Overflow. "2024 Developer Survey." Stack Overflow, 2024, survey.stackoverflow.co/2024/. Accessed 20 Mar. 2025.
Yang, John, et al. "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering." arXiv, 2024, arxiv.org/abs/2405.15793. Accessed 20 Mar. 2025.