AI Frontiers, part 65: A field guide to shipping AI systems — what I would do differently
Part 65from the AI Frontiers series · 65 parts in all
Sixty-five entries ago this series started with a transformer paper (Vaswani et al.) and the question of what its architecture implied. The arc since then has followed the technology out of the research lab and into the ordinary business of building software that people depend on: weights, retrieval, context windows, sparsity, quantization, caching, protocols, agents, evaluation, security, and money. The last entry is the one I would want a reader to keep, so it is not a summary. It is the field guide — the practices I would keep, the things I got wrong, and the habits that produced the difference between systems that got better and systems that got replaced.
The organizing idea, stated once more and then dropped: the frontier is where the ideas are born and the industry is where they get cheap. Everything useful in this series is a consequence of taking that sentence seriously, because it tells you where to invest your attention. Not in chasing capability, which arrives on someone else's schedule, but in the machinery that converts capability into reliable, affordable, accountable work.
Ten practices I would keep
1. Build the evaluation before the feature. It is the same argument as writing the test first, and it is true for the same reason: a measurement you design in advance constrains the solution space usefully, while a metric you invent afterward describes whatever you built. Every system in this series that improved had a private test set and a cost column; every system that churned had a demo.
2. Start with the smallest model that passes, and route later. It is far easier to justify escalating a request to a bigger model than to explain why every request was paying frontier prices. The two-tier pattern — small models as the workforce, frontier models as arbiters, and a router deciding which — has held through every model generation since 2023.
3. Make the trace the primary artifact. Full prompts, retrieval candidates before and after filtering, tool calls with arguments, token counts, cost, and a replay mode. If a system cannot be re-run from a recording, it cannot be debugged, evaluated, or improved, and part 49 argued it is also the raw material for everything else.
4. Treat prompts as configuration and schemas as interfaces. Version both, review changes to both, and keep them out of application code. The schema is the contract between the model and the rest of your system, and constrained decoding makes it enforceable rather than aspirational.
5. Contain rather than detect. Least privilege, egress allowlists, sandboxes, read/write separation, and human confirmation on irreversible actions. The consensus of part 57 was that a model's judgment cannot be a security boundary, and that the defenses that work are the ones from 1975.
6. Measure the human as carefully as the model. Acceptance rates, override accuracy, time per item, and the rate at which reviewers catch something the system got wrong. A review queue without those numbers is either unnecessary or not working, and part 62 showed how quickly it becomes a signature.
7. Own the loop from production back to training. Corrected outputs become evaluation cases, evaluation cases gate releases, releases produce traces, traces produce training data. This is the only compounding asset in an applied AI business, because it is the only one that does not reset when the base model changes.
8. Instrument cost per completed task, including retries and reviews. A unit economics number you can see is a design constraint; one you discover from a finance review is a crisis. Part 50 and part 64 are the same argument at different altitudes.
9. Keep correctness in the software and judgment in the model. Deterministic verification — a test suite, a schema, a database constraint, a type system — beats asking a model to be careful. Almost every reliability win of the last three years came from moving a responsibility out of the model and into ordinary engineering.
10. Design the rollback before the rollout. One config change, reversible in minutes, is what allows a team to be aggressive about trying things. Systems that cannot roll back spend their caution budget on avoidance instead of on progress.
Five things I got wrong
The first was overrating prompt engineering. The 2022 and 2023 advice that a sufficiently clever prompt obviates other work aged badly. Prompts matter enormously; they are also the least durable asset in the stack, and the number of systems whose entire quality story rested on a prompt that a base model update broke is a large number.
The second was underrating retrieval. I spent part of 2023 assuming that larger context windows would make search infrastructure redundant. Part 52 is the correction: the window is a resource and retrieval is the discipline of spending it, and selection became more valuable as the budget grew, not less.
The third was believing multi-agent architectures would be the agent architecture. Part 47 records the correction — decomposition pays for context isolation and permission boundaries, and almost nothing else. Most of what I called coordination was really context management wearing a costume.
The fourth was treating evaluation as a launch gate. It is an operating function, with a budget, an owner, and a weekly cadence, and teams that treated it as a milestone discovered that quality drifts without anyone touching the code.
The fifth was underestimating integration and change management. The technical work in every project I have seen succeed was a minority of the effort. The rest was process design, training, exception handling, and the human work of changing how people do their jobs. The measured evidence agrees: the gains are real but jagged (Dell'Acqua et al.), concentrated where tasks are bounded and expertise thin (Noy and Zhang; Brynjolfsson et al.), and they arrive after an investment phase that shows up as disappointing measured productivity for a while (Brynjolfsson et al., "The Productivity J-Curve").
The parts of software engineering that did not change
The most useful thing this field has clarified is how much of the old discipline still applies. Fred Brooks's argument that there is no single technique that yields an order of magnitude in software productivity remains the best description of why applied AI is hard (Brooks). The observational literature on machine learning systems made the same point about hidden debt (Sculley et al.). The interaction guidelines, the experimentation discipline, the release engineering, and the observability practices all carry over with small modifications (Amershi et al.; Kohavi et al.; Beyer et al.). Conway's law still describes why your architecture looks like your org chart (Conway). Test-first still works, now applied to prompts and judgments as much as to functions (Beck). And the opposite of a single silver bullet is Sutton's argument that general methods which absorb computation keep winning: the model layer will improve on its own schedule, whether or not your architecture is ready for it (Sutton).
What is genuinely new is narrower than the discourse suggests, and worth stating precisely because it is the whole difficulty. We are building systems whose central component cannot be specified, only measured; whose failure mode is confident plausibility rather than an exception; whose cost and latency are dominated by generated text; and whose behavior changes when someone else updates a dependency you do not control. Everything in this series is an attempt to build reliable software out of that material, and the techniques that worked are consistent: measure relentlessly, constrain structurally, contain the blast radius, keep the loop running, and put the model where the verification is cheap.
The arc, in one paragraph
Five years ago a paper's architecture became everybody's foundation. Then the weights leaked and an ecosystem formed; the frontier got multimodal and then reasoning-shaped; retrieval and context engineering turned into infrastructure; sparsity, quantization, and caching made scale affordable; agents arrived, failed in instructive ways, and came back as workflows with narrow tools and real permissions; protocols standardized the plumbing; security forced the field to take containment seriously; biology showed what rigorous validation looks like; and the economics settled into the pattern that this whole series has been describing — capability is commoditized on a quarterly rhythm, and durable advantage lives in evaluation, data, integration, and trust, which is to say in the unglamorous parts that nobody demos.
The technology will keep improving, faster than any of us expects and in directions most of us will guess wrong. The practices above will not need much revision, because they are the habits of people who build things that other people depend on, and that job has not changed since people started writing software. Measure what you ship. Write down why. Keep the humans able to disagree. Make the reversible choice when the evidence is thin. And when the next model makes the impossible cheap, be positioned to spend the difference on something better than you had.
That is the field guide. Thanks for reading sixty-five of these.
The pre-ship checklist
All of this compresses into a list short enough to run before a launch, which is the form it has taken on the teams I have worked with. Twenty questions, each of which has a yes or no answer.
Measurement. Do we have a private evaluation set built from real traffic? Is it versioned, and is the split protected from accidental inspection? Do we measure cost per completed task, including retries and reviews? Do we have a p95 latency number and a budget? Can we tell whether a change improved things, and by how much given the noise?
Observability. Is there a full trace with prompts, retrieval candidates, tool calls, and cost? Can we replay a recorded interaction? Do we know which model version served each request, and which policy version made each decision? Are traces redacted, retained for a defined period, and stored in the required jurisdiction?
Robustness. Do we know the system's effective operating envelope, and does the product route outside it to a human rather than guessing? Is there a mechanism that forces abstention when confidence is low? Does the system fail safely when its dependencies are unavailable or slow?
Security and containment. What untrusted text can it read, and what credentials does it hold? Can data leave through a channel we have not enumerated? Which actions are irreversible, and which of those require human confirmation? If it were compromised right now, would we find out from our own telemetry or from a customer?
Operations. Can we roll back in one change, in minutes, and have we practiced? Does someone own the quality dashboard and the review queue? Is there a documented response for each alert, or is the response "look at it"? Do we know who to call when a model provider changes behavior at 2 a.m. in a region we do not monitor closely?
Commercials. Do we know what the customer pays today for this work? Does the price rest on a measurement the customer can inspect? Do we know which of the four margin sources — proprietary data, workflow embedding, verification, distribution — we actually have? And if the base model were replaced tomorrow with a slightly worse and much cheaper one, would we know how to trade between them?
A team that can answer those questions has done the work that the last sixty-four entries described. A team that cannot has a prototype with a launch date, which is a different thing and should be labeled as such.
What I would tell a new engineer
Three pieces of advice, in the order I would give them.
Learn the measurement side first. The engineers who became most valuable in this field over the last three years were not the ones with the best prompt intuition; they were the ones who could design an evaluation, read a trace, and tell whether an observed difference was real. That skill transfers across every model generation, every vendor, and every architecture, which is more than can be said for anything else in the stack.
Be suspicious of fluency, including your own. The characteristic failure of this era is a system that sounds right, and the characteristic professional habit is to demand evidence: where did this come from, what would falsify it, what does the trace show, what does the test say. Fluency is a signal about the generator, not about the claim, and internalizing that is the single most useful defense against both the technology's weaknesses and your own susceptibility to them.
Build the boring parts well. Configuration, versioning, logging, schema validation, rollbacks, tests. They feel like a distraction from the interesting work and they are the reason the interesting work compounds. Every impressive AI system I have seen hold up over years was, underneath, a well-built piece of ordinary software that happened to have a model in it. The model is the easy part, and it always was — which is exactly why the durable advantage lives where the work is hard.
Works Cited
Amershi, Saleema, et al. "Guidelines for Human-AI Interaction." Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 2019. Accessed 2 July 2026.
Beck, Kent. Test-Driven Development: By Example. Addison-Wesley, 2002. Accessed 2 July 2026.
Beyer, Betsy, et al., editors. Site Reliability Engineering: How Google Runs Production Systems. O'Reilly Media, 2016. Accessed 2 July 2026.
Brooks, Frederick P. "No Silver Bullet: Essence and Accidents of Software Engineering." IEEE Computer, vol. 20, no. 4, 1987, pp. 10–19. Accessed 2 July 2026.
Brynjolfsson, Erik, Danielle Li, and Lindsey R. Raymond. "Generative AI at Work." NBER Working Paper 31161, 2024, www.nber.org/papers/w31161. Accessed 2 July 2026.
Brynjolfsson, Erik, Daniel Rock, and Chad Syverson. "The Productivity J-Curve: How Intangibles Complement General Purpose Technologies." American Economic Journal: Macroeconomics, vol. 13, no. 1, 2021, pp. 333–372. Accessed 2 July 2026.
Conway, Melvin E. "How Do Committees Invent?" Datamation, vol. 14, no. 4, 1968, pp. 28–31. Accessed 2 July 2026.
Dell'Acqua, Fabrizio, et al. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality." Harvard Business School Working Paper 24-013, 2024. Accessed 2 July 2026.
Kohavi, Ron, Diane Tang, and Ya Xu. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press, 2020. Accessed 2 July 2026.
Noy, Shakked, and Whitney Zhang. "Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence." Science, vol. 381, 2023, pp. 187–192. Accessed 2 July 2026.
Sculley, D., et al. "Hidden Technical Debt in Machine Learning Systems." Advances in Neural Information Processing Systems, 2015. Accessed 2 July 2026.
Sutton, Rich. "The Bitter Lesson." Incomplete Ideas, 2019, incompleteideas.net/IncIdeas/BitterLesson.html. Accessed 2 July 2026.
Vaswani, Ashish, et al. "Attention Is All You Need." Advances in Neural Information Processing Systems, 2017. Accessed 2 July 2026.