Trust is earned, not given

A different perspective

2025-09-10 · Projects

AI fundamentals, part 8: From lab to product — context, tools, and failure modes

Part 8from the AI fundamentals series · 8 parts in all

The series closer. Everything so far — tokenization, attention, training, scaling, sampling, evaluation, alignment — describes the model. Shipping it in a product adds one more layer: the context you assemble around the model. In practice, that layer decides whether an LLM feature feels magical or broken. This part is the field guide.

The context window is your real product surface

The model knows nothing except what's in front of it. Every good LLM feature is a function that builds context:

def build_context(user_message, history, docs, budget_tokens=6000):
    # Assemble the prompt under a token budget, priority-ordered.
    # The model can only use what survives this function.
    parts = [
        ("system",  SYSTEM_RULES),                      # fixed rules, never cut
        ("memory",  summarize(history) if long(history) else history),
        ("knowledge", retrieve_top_k(user_message, docs, k=4)),  # most relevant
        ("history", tail(history, keep_last=6)),        # recent turns matter most
        ("user",    user_message),                      # the actual question
    ]
    ctx, used = [], count_tokens(SYSTEM_RULES)
    for name, text in parts[1:]:
        t = count_tokens(text)
        if used + t > budget_tokens:                    # budget exhausted:
            text = compress_or_drop(name, text)         # trim oldest/least relevant
            t = count_tokens(text)
        ctx.append((name, text)); used += t
    return ctx

Three context disciplines that prevent most production bugs: summarize, don't truncate (a compressed summary of old conversation beats a hard cut mid-sentence); retrieve, don't memorize (RAG: search your corpus at question time and paste the top matches in, so knowledge stays current and auditable); and order by importance — models attend most to the beginning and end of long contexts; the middle is where facts get dropped ("lost in the middle").

Tool use: letting the model act

LLMs hallucinate facts; they're much better at deciding which known function to call. That asymmetry is the entire tool-use pattern — the model emits structured intent, your code does the work:

TOOLS = {
    "get_order_status": {
        "args": ["order_id"],
        "handler": lambda order_id: orders_db.lookup(order_id),  # typed, allow-listed
    },
    # Note what is NOT here: any tool that writes, pays, emails, or deletes.
}

def agent_turn(user_message, history):
    # One agent loop. The model proposes a tool call; WE validate and run it.
    # The model never executes anything by itself.
    ctx = build_context(user_message, history, docs=[])
    reply = model_call(ctx, tools=list(TOOLS))
    if reply.get("tool_call"):
        name, args = reply["tool_call"]["name"], reply["tool_call"]["args"]
        if name not in TOOLS:                       # allow-list check first
            return "That capability isn't available."
        result = TOOLS[name]["handler"](**args)     # our code, our validation
        return model_call(ctx + [tool_result(result)])   # answer with real data
    return reply["text"]

The security rule from our RMA article is universal: the model proposes, your code disposes. Validate every argument against server-side state, allow-list the tool set, and keep state-changing operations in deterministic code the model can request but never directly trigger.

The failure-mode catalog (learn these by heart)

The checklist for shipping

Before any LLM feature goes live: a frozen golden eval (part 6) with the failure list reviewed; an owner for the prompt (it's code — version it, review it); guardrails on inputs and outputs; a cost ceiling per request; a fallback path when the provider is down; and a human review queue for the first weeks. None of that is AI. All of it is why one AI feature becomes trusted and another becomes a liability.

That completes the series. You now know what happens between "user types a question" and "answer appears" — tokenization into bytes, attention over context, a stack of transformer blocks trained under scaling laws, sampled token by token with a KV cache, aligned with human feedback, evaluated against benchmarks that can lie — and you've written the essential ideas yourself in Python. The best next step is the one you take with your own corpus: build a tiny model, break it, and fix it. Everything above will make sense retroactively.