Trust is earned, not given

A different perspective

2023-08-10 · Projects

AI Frontiers, part 30: Tool use arrives β€” function calling and the plugin experiment

Part 30from the AI Frontiers series · 65 parts in all

Two things happened to language models in 2023 that looked similar and turned out to be completely different in their consequences. In March, OpenAI opened a plugin ecosystem: third-party services could declare an API to ChatGPT, and the model could call them on a user's behalf (OpenAI, "ChatGPT Plugins"). In June, the API gained function calling, a much smaller feature in which a model could return a structured request for a named function with typed arguments, and the calling application would decide what to do about it (OpenAI, "Function calling and other API updates").

The plugin store was the ambitious one, and it failed. The function calling primitive was a request/response format with no ecosystem attached, and it became the foundation of everything that followed β€” agent frameworks, retrieval pipelines, and eventually a formal protocol for connecting models to tools. The reasons for that divergence are instructive, and they are mostly about where the boundary between the model and the application is drawn.

The research case, which was already overwhelming

Tool use was not a product idea looking for a justification. The literature had established the value of giving a language model an external action well before anyone shipped it.

WebGPT (Nakano et al.) put a browsing interface in front of a model, let it issue search queries and quote passages, and got better factual answers than the underlying model could produce alone. Toolformer (Schick et al.) showed that a model could learn when to call an API by self-supervision: sample candidate calls, keep the ones that reduce the loss on the following tokens, and fine-tune on the result. TALM (Parisi et al.) reached the same place with a different recipe, and LaMDA's use of a calculator and a search tool was documented in its own paper (Thoppilan et al.). The MRKL paper (Karpas et al.) named the architecture β€” a router directing queries to language, knowledge and reasoning modules β€” and the augmented-language-model survey (Mialon et al.) catalogued the whole space.

Two results from 2023 sharpened the case. ReAct (Yao et al.) interleaved reasoning traces with actions, showing that a model that thinks, acts and observes outperforms one that only thinks β€” the pattern that every agent since has followed. And Gorilla (Patil et al.) attacked the problem that makes tool use hard in practice: given thousands of possible APIs, the model has to select the right one and call it with valid arguments. Their finding that retrieval over API documentation beats memorization is the same result the retrieval literature reached for documents, and it is the reason tool selection is a search problem rather than an act of recall.

Why the plugin model did not work

Plugins made a specific bet: that exposing services to a chat interface would produce an app ecosystem, with the same dynamics as a mobile app store. Four properties of the setup worked against it.

The interface was the product owner's, not the developer's. A plugin ran inside someone else's conversation surface, with the interaction pattern chosen by the platform and the context of its use invisible to the service. There was no way to design an experience for a particular task, and no way to observe the traffic that arrived.

Discovery was unsolved. A plugin could only be invoked when a user thought to enable it, and the model's decision to call one was unreliable. A user who has to remember to turn on the right plugin has already done the hard part of the task; the automation that remains is marginal.

The economics were hostile. Developers bore the cost of an integration, received no revenue of their own, and competed for attention in a store where the platform's own capabilities were improving every quarter. Anything a plugin did that was valuable enough to matter was a candidate for being absorbed into the base product.

The reliability was not there. A plugin that fails half the time teaches users not to use it, and in mid-2023 the components β€” model-driven selection, argument construction, error recovery, authentication to a third-party service β€” were all individually marginal. Multiplying four marginal things produces a product nobody trusts.

The lesson is not that plugins were a mistake. It is that the value of an integration lives with whoever owns the interface where the user's task actually happens, and a general-purpose chat window is usually not that place.

Why function calling worked

The June feature appears much less ambitious and is much more useful, for structural reasons.

It does not specify where the model runs, who calls the function, or what the user interface looks like. The model emits a structured request; the application executes it, decides whether to ask for confirmation, formats the result, and controls the loop. That leaves every consequential decision β€” permissions, retries, logging, when to give up β€” with the developer, which is where they have to be, because those decisions depend on the domain and not on the model.

It is also a contract about shape rather than about behavior. A function is described by a name, a purpose and a set of typed parameters, and the model is asked to fill that shape rather than to write prose. The output is parseable, which means it can be validated, logged, replayed and tested. This is the same argument as the schema argument in the vision pipelines: forcing the model to emit a structure converts an unverifiable claim into an ordinary error-handling problem.

The third property is that function calling composes with everything else. A tool call is a message in a conversation, so it fits inside retrieval-augmented systems, inside multi-step reasoning, and inside the rest of the stack. Plugins were a closed surface; function calling was a format.

The ecosystem that formed in the gap

By the summer of 2023 the practical tool-use layer was being built by frameworks rather than platforms. LangChain and LlamaIndex supplied abstractions for tools, agents and memory, and their rapid adoption was a symptom: developers needed to move the prompt-assembly and tool-execution logic into their own code, next to their own data and permissions, and the frameworks gave them a starting point.

The research kept moving in parallel. ToolLLM and ToolBench (Qin et al.) built a dataset of over 16,000 real APIs and showed that a model fine-tuned on tool-use trajectories could handle unfamiliar tools; RestGPT (Song et al.) applied the pattern to REST APIs specifically; HuggingGPT (Shen et al.) chained models together as tools, with a language model as the planner. Each of these was an attempt to answer the same question β€” how does a model choose among thousands of possible actions β€” and they converged on the answer the retrieval literature already had: do not put the whole catalogue in the prompt, retrieve the plausible candidates and let the model choose among a handful.

That convergence matters because it is the design principle the eventual standard would formalize two years later. A tool registry is a search index. Fewer, better-named tools beat many overlapping ones. And the model's ability to choose correctly degrades much faster with the number of options than with the difficulty of the task, which is a property of the selection problem rather than of the model.

What a tool description has to contain

Since the model's choice is driven entirely by text, the description of a tool is a specification written for a reader with no context, no curiosity and no willingness to ask a clarifying question. A surprising amount of tool-use reliability is determined before any model is involved, by how well that specification was written.

The elements that matter, in roughly descending order: a name that distinguishes the tool from its neighbours, parameter descriptions that state expected format as well as meaning, at least one example of a valid call, an explicit statement of what the tool does not handle, and a documented error surface so the model can distinguish a retryable failure from a permanent one. The example is doing disproportionate work, because it teaches the shape of a valid argument in a way that a prose description does not.

Two failure patterns follow from getting this wrong. The first is ambiguity between overlapping tools: if a search tool and a lookup tool have similar descriptions, the model will pick between them arbitrarily and the choice will look random in your logs. The second is catalogue bloat. Selection accuracy falls as the number of available tools rises, and it falls faster when names and descriptions overlap. The remedy is the same as in retrieval: keep the library large and the prompt small, and choose candidates per request rather than presenting everything at once.

A failure taxonomy worth instrumenting

Tools introduce specific failure modes, and each one needs a different fix. Being able to count them separately is the difference between improving a system and guessing at it.

Wrong tool. The model selected a plausible but incorrect function. Fix: better descriptions, fewer candidates, and examples that show the intended boundary.

Right tool, invalid arguments. Right function, schema violation or nonsensical value. Fix: stricter schemas with enumerations and formats, plus a validation layer that returns a structured error the model can act on rather than a stack trace.

Fabricated arguments. A field populated with a plausible invented value β€” an order number that does not exist, a date outside the range. This is hallucination wearing a schema, and it is caught the same way: validate before executing, and never let a plausible identifier turn into a database write.

Ignored results. The tool returned useful data and the model answered as if it had not. Fix: keep results small, place them immediately before the question, and make the instruction to use them explicit.

Success on error. The tool failed and the model reported success. The most damaging mode and the reason a tool result must carry an unambiguous status field rather than embedding the error in prose the model is free to interpret.

Finally, the control-flow question that plugins required and function calling leaves to you: how many iterations are allowed, and what happens when the limit is reached. An unbounded loop of tool calls is an unbounded bill and an unbounded latency, and the correct answer is a budget expressed in iterations and tokens rather than a hope that the model will stop.

What tool use changed about the shape of applications

Three shifts followed, and they are visible in retrospect more clearly than they were at the time.

The model stopped being the whole application. A system built around completion was a prompt and a response. A system built around tools is a model plus a permission model plus a set of integrations plus an audit log, and the model is one component among several. The teams that understood this built products; the teams that treated the model as the product built demos.

Evaluation got harder and more concrete at the same time. Harder, because a failure could now be a wrong tool choice, a malformed argument, a failed API call, or a bad interpretation of a valid response, and untangling those requires tracing. More concrete, because each of those failures is mechanically observable in a way that a bad paragraph is not. Tool use traded an ambiguous failure for a legible one, which is a good trade even when it feels like more work.

Security became an application concern rather than a model concern. Once a model can cause a side effect, the interesting question stops being what the model says and becomes what it is allowed to do. That question has no model-side answer, and the entire apparatus of allowlists, confirmation prompts and scoped credentials that followed is the industry discovering that prompting is not a permission system.

The most durable thing to come out of 2023 in this area, then, is unglamorous: a JSON schema for describing a function, and the convention that the calling application maintains control. Everything in the agent era β€” the frameworks, the protocols, the safety architecture β€” is built on that small decision, and the ambitious version of the idea is the one that died.

Works Cited

Karpas, Ehud, et al. "MRKL Systems: A Modular, Neuro-Symbolic Architecture That Combines Large Language Models, External Knowledge Sources and Discrete Reasoning." arXiv, 2022, arxiv.org/abs/2205.00445. Accessed 10 Aug. 2023.

Mialon, GrΓ©goire, et al. "Augmented Language Models: A Survey." arXiv, 2023, arxiv.org/abs/2302.07842. Accessed 10 Aug. 2023.

Nakano, Reiichiro, et al. "WebGPT: Browser-Assisted Question-Answering with Human Feedback." arXiv, 2021, arxiv.org/abs/2112.09332. Accessed 10 Aug. 2023.

OpenAI. "ChatGPT Plugins." OpenAI, 23 Mar. 2023, openai.com/index/chatgpt-plugins/. Accessed 10 Aug. 2023.

---. "Function Calling and Other API Updates." OpenAI, 13 June 2023, openai.com/index/function-calling-and-other-api-updates/. Accessed 10 Aug. 2023.

Parisi, Aaron, et al. "TALM: Tool Augmented Language Models." arXiv, 2022, arxiv.org/abs/2205.12255. Accessed 10 Aug. 2023.

Patil, Shishir G., et al. "Gorilla: Large Language Model Connected with Massive APIs." arXiv, 2023, arxiv.org/abs/2305.15334. Accessed 10 Aug. 2023.

Qin, Yujia, et al. "Tool Learning with Foundation Models." arXiv, 2023, arxiv.org/abs/2304.08354. Accessed 10 Aug. 2023.

Schick, Timo, et al. "Toolformer: Language Models Can Teach Themselves to Use Tools." arXiv, 2023, arxiv.org/abs/2302.04761. Accessed 10 Aug. 2023.

Shen, Yongliang, et al. "HuggingGPT: Solving AI Tasks with ChatGPT and Its Friends in Hugging Face." arXiv, 2023, arxiv.org/abs/2303.17580. Accessed 10 Aug. 2023.

Song, Yifan, et al. "RestGPT: Connecting Large Language Models with Real-World RESTful APIs." arXiv, 2023, arxiv.org/abs/2306.06624. Accessed 10 Aug. 2023.

Thoppilan, Romal, et al. "LaMDA: Language Models for Dialog Applications." arXiv, 2022, arxiv.org/abs/2201.08239. Accessed 10 Aug. 2023.

Yao, Shunyu, et al. "ReAct: Synergizing Reasoning and Acting in Language Models." arXiv, 2022, arxiv.org/abs/2210.03629. Accessed 10 Aug. 2023.