Trust is earned, not given

A different perspective

2026-06-04 · Projects

AI Frontiers, part 64: The economics of applied AI — where the margin actually is

Part 64from the AI Frontiers series · 65 parts in all

The price of a million tokens fell by more than two orders of magnitude between 2023 and 2026, and the median AI product still does not have software margins. That combination is worth explaining, because the explanation is not that the technology failed. It is that inference cost was never the dominant cost in an applied system, and the parts that dominate — evaluation, integration, human review, and the unglamorous work of making a probabilistic component behave deterministically enough to sell — do not fall on anyone else's schedule.

This entry is about the money. It follows part 50, which was about making a single system cheap, and widens the frame to the business: the full cost stack, where measured value creation actually shows up in the research literature, why some AI businesses have durable margins while most do not, and what the falling cost curve means for how you should be building right now.

The full cost stack

Inference is one line. Here is the whole bill for a production AI product, roughly in order of how often it is underestimated.

Integration and change management. The largest and most ignored cost. Getting a model to produce good output is a fraction of getting an organization to change what it does. The work is process design, training, exception handling, and the political labor of persuading the people whose job the system touches. This is measured in person-months and almost never appears in an AI project's business case, which is why so many pilots succeed technically and commercially evaporate.

Evaluation and quality infrastructure. The private test sets, judges, review queues, and monitoring from part 60 are a permanent operating cost, not a launch cost. In systems I have costed, this runs at roughly the same order of magnitude as inference, and it is the line item most likely to be cut in a downturn — with a quality collapse following two quarters later.

Human review. Where the workflow gates output, the review capacity described in part 62 is a labor cost with a real hourly rate. A system that requires two minutes of specialist review per item cannot be cheaper than the two minutes it consumes, no matter how cheap the tokens are.

Data acquisition and maintenance. Curation, labeling, refresh, and the licensing questions of part 37. For retrieval-heavy products this includes the boring cost of keeping a corpus current, which is headcount that never stops.

Inference and orchestration. Real, and now small enough that it is rarely the reason a product fails. Where it still matters is at high volume, which is exactly where the routing and caching levers of parts 50 and 59 pay off.

Platform and on-call. Someone has to run it at three in the morning. AI systems have more of these failure modes than ordinary services, because the failure can be "the answers got worse" rather than "the process died."

Two accounting consequences follow. Gross margin should be computed per completed task, not per call, since retries, escalations, and reviews are part of the unit. And the business case should be built on the change in a full-time-equivalent hour, not on the token bill, because the token bill is the smallest defensible part of the value story.

What the measured evidence says about value

The productivity literature for generative AI is young, unusually well-measured for a new general-purpose technology, and narrower than the marketing suggests. Four results frame it.

Assistive tools help most where expertise is thin and the task is bounded. A randomized study of professional writing tasks found large reductions in time and increases in quality, with the largest gains for lower-ability participants, which compresses performance differences rather than amplifying them (Noy and Zhang). Randomized work on customer support found substantial productivity gains concentrated among less-experienced workers, with improved issue resolution and better customer sentiment (Brynjolfsson et al., "Generative AI at Work"). The randomized trial of a coding assistant found large speedups on a well-specified implementation task (Peng et al.).

Difficulty is jagged, and the sharp edge is where projects fail. The field study of knowledge workers on realistic consulting tasks found that the same technology that dramatically improved performance inside its competence produced worse results than unaided humans outside it — the now-standard finding that capability has a ragged boundary and users cannot see where it lies (Dell'Acqua et al.). This is the empirical foundation for everything in part 62 about trust calibration, and it is the single most commercially important finding in the literature, because it means deployment quality depends on routing around the boundary.

Exposure is large; automation is smaller. Estimates of the share of tasks exposed to language-model capabilities are large — the frequently cited figure is that around half of work tasks could be substantially assisted, with a much smaller fraction fully automatable (Eloundou et al.). The gap between exposure and automation is where the integration and verification costs live, and it is why economy-wide projections of near-term productivity gains are more modest than the task-level results imply (Acemoglu).

The payoff lags the investment. The productivity J-curve analysis of prior general-purpose technologies argues that measured productivity often falls during the build-out phase, because firms are investing in intangible capital — process redesign, data, training — that national accounts do not measure well (Brynjolfsson et al., "The Productivity J-Curve"). If that pattern holds, the organizations reporting disappointing returns in 2025 and 2026 are not evidence that the technology does not work; they are evidence that they are still in the investment phase, and the ones that will do well are those that finish the intangible work rather than abandoning it halfway.

Where durable margin comes from

Four sources, and only four, show up repeatedly in applied AI businesses that hold their pricing.

Proprietary data that improves the system. Not data you happen to have, but data whose accumulation is a byproduct of operating the business: interaction traces, corrected outputs, outcome labels, and the evaluation history that goes with them. This is the compounding loop from parts 49, 60, and 62, and it is the only source of margin that gets stronger with use.

Workflow embedding. Systems that own the process — the queue, the approvals, the audit trail, the integration with five other systems — are expensive to replace for reasons that have nothing to do with model quality. Distribution and switching cost are as decisive here as in any enterprise software market, and considerably more decisive than the model.

Verification infrastructure. The ability to be right in a domain where being wrong is expensive, built from domain-specific checks, review workflows, and regulatory evidence. This is the hardest to copy and the most valuable, because it is where a customer's risk is transferred and where the price can be justified per unit of risk rather than per token.

Distribution. Access to the customers, the channel, or the interface through which the work already flows. Unglamorous and decisive.

What is missing from that list is the model, and that is the honest summary of the last three years: a business whose only asset is a good prompt against someone else's model is competing on a component that is being commoditized on a quarterly basis. That does not make such businesses worthless — many are doing real integration work and will earn real margins — but it does mean the moat has to be somewhere else, and the earlier a team identifies where, the less of its roadmap it spends re-implementing the same wrapper.

The pricing question

Three pricing models are in use, with different risk profiles and different measurement burdens.

Per seat. Familiar, forecastable, and misaligned with value when usage varies by an order of magnitude across customers. It works when the product replaces an existing per-seat line item, which is the most common and most defensible situation.

Per token or per call. Straightforward pass-through, which means your margin is fixed by the difference between your cost and your price, and the customer captures the benefit of falling inference prices instead of you. Fine for infrastructure, poor for applications.

Per outcome. The most attractive and the hardest. It requires defining the outcome precisely, measuring it automatically, and absorbing the variance of a probabilistic system in your own revenue. It also requires the accountability machinery of part 62 to be real, because the customer will inspect the definition of "resolved" the moment the invoice arrives. Where it works — high-volume, verifiable, previously labor-costed tasks — it commands the best economics of the three, because it prices the risk transfer rather than the tokens.

The practical advice I would give an operator is to price against whatever the customer currently pays for the work, then defend that price with measurement. If the system saves a six-dollar-per-item process and costs eighty cents to run, the price is not eighty cents plus a markup; it is a fraction of six dollars, justified by an audit trail the customer can inspect. Getting the measurement right is what makes that conversation possible, which is why building your own evaluation keeps turning out to be a commercial investment rather than an engineering chore.

What the cost curve means for strategy

Three strategic implications of a technology whose unit cost falls faster than almost any in history.

Do not optimize for the price of today's model. Capability per dollar has improved faster than price per token has fallen, so the systems that win are the ones that can spend more inference where it buys quality — a theme from part 50 — rather than the ones engineered around a 2024 cost structure that no longer exists.

Invest in the loops, not the prompt. Data collection, evaluation, tracing, and review infrastructure are the assets that survive every model upgrade. Prompts do not. Any engineering hour spent making the loop better has a longer half-life than an hour spent tuning a prompt against a model that will be deprecated next quarter, and this is the closest thing to a reliable rule in applied AI.

Expect the J-curve to be real and budget for it. The organizations that have gotten returns so far did the unglamorous work: redesigning the process, building the evaluation, training the reviewers, and shipping the second version after the first one disappointed. The technology's economics are genuinely favorable. They are just favorable on a schedule set by integration and verification work, and no amount of cheaper inference substitutes for finishing it.

Three shapes, three margin structures

The cost model above produces very different business outcomes depending on which shape a product takes, and the shapes are recognizable enough to be worth naming.

Infrastructure and tooling. Sells to builders, usually per token or per seat, with gross margins that look like cloud infrastructure — high in percentage terms once at scale, punished by supporting costs at low volumes, and vulnerable to the platform owner absorbing the functionality. The defensible position is usually the one that is hardest for a platform to copy: deep integration with a specific ecosystem, or performance that comes from accumulated operational knowledge rather than from a better algorithm.

Vertical applications. Sell to an industry, price against the labor being replaced, and carry the integration and compliance burden that generic vendors avoid. These have the best gross margins available in applied AI, because the price is anchored to human work rather than to compute, and the worst sales cycles, because the buyer is a regulated organization comparing you to a process it already understands. The margin is real and the moat is the domain detail.

Internal systems. No revenue, only savings, which means the business case has to be built from a before-and-after measurement the finance function will accept. These are the projects where the discipline of this series pays off most directly, because the value is exactly the delta in cost per task, quality-adjusted, and any team that cannot measure that delta cannot defend the investment. They also tend to be the most durable: an internal system that works produces a compounding operational advantage that competitors cannot buy, because it depends on process change rather than on a capability that anyone can license.

For each shape the same discipline applies. Know the cost per completed task, know the customer's alternative, and know which of the four margin sources you actually have. A product that can answer those three questions is investable even if the model it depends on gets replaced next quarter, because the value was never in the model.

Works Cited

Acemoglu, Daron. "The Simple Macroeconomics of AI." Economic Policy, 2025, www.nber.org/papers/w32487. Accessed 4 June 2026.

Brynjolfsson, Erik, Danielle Li, and Lindsey R. Raymond. "Generative AI at Work." NBER Working Paper 31161, 2024, www.nber.org/papers/w31161. Accessed 4 June 2026.

Brynjolfsson, Erik, Daniel Rock, and Chad Syverson. "The Productivity J-Curve: How Intangibles Complement General Purpose Technologies." American Economic Journal: Macroeconomics, vol. 13, no. 1, 2021, pp. 333–372. Accessed 4 June 2026.

Dell'Acqua, Fabrizio, et al. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality." Harvard Business School Working Paper 24-013, 2024. Accessed 4 June 2026.

Eloundou, Tyna, et al. "GPTs Are GPTs: Labor Market Impact Potential of LLMs." Science, vol. 384, 2024. Accessed 4 June 2026.

Noy, Shakked, and Whitney Zhang. "Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence." Science, vol. 381, 2023, pp. 187–192. Accessed 4 June 2026.

Peng, Sida, et al. "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot." arXiv, 2023, arxiv.org/abs/2302.06590. Accessed 4 June 2026.