AI fundamentals, part 8: From lab to product — context, tools, and failure modes
Part 8from the AI fundamentals series · 8 parts in all
The series closer. Everything so far — tokenization, attention, training, scaling, sampling, evaluation, alignment — describes the model. Shipping it in a product adds one more layer: the context you assemble around the model. In practice, that layer decides whether an LLM feature feels magical or broken. This part is the field guide.
The context window is your real product surface
The model knows nothing except what's in front of it. Every good LLM feature is a function that builds context:
def build_context(user_message, history, docs, budget_tokens=6000):
# Assemble the prompt under a token budget, priority-ordered.
# The model can only use what survives this function.
parts = [
("system", SYSTEM_RULES), # fixed rules, never cut
("memory", summarize(history) if long(history) else history),
("knowledge", retrieve_top_k(user_message, docs, k=4)), # most relevant
("history", tail(history, keep_last=6)), # recent turns matter most
("user", user_message), # the actual question
]
ctx, used = [], count_tokens(SYSTEM_RULES)
for name, text in parts[1:]:
t = count_tokens(text)
if used + t > budget_tokens: # budget exhausted:
text = compress_or_drop(name, text) # trim oldest/least relevant
t = count_tokens(text)
ctx.append((name, text)); used += t
return ctx
Three context disciplines that prevent most production bugs: summarize, don't truncate (a compressed summary of old conversation beats a hard cut mid-sentence); retrieve, don't memorize (RAG: search your corpus at question time and paste the top matches in, so knowledge stays current and auditable); and order by importance — models attend most to the beginning and end of long contexts; the middle is where facts get dropped ("lost in the middle").
Tool use: letting the model act
LLMs hallucinate facts; they're much better at deciding which known function to call. That asymmetry is the entire tool-use pattern — the model emits structured intent, your code does the work:
TOOLS = {
"get_order_status": {
"args": ["order_id"],
"handler": lambda order_id: orders_db.lookup(order_id), # typed, allow-listed
},
# Note what is NOT here: any tool that writes, pays, emails, or deletes.
}
def agent_turn(user_message, history):
# One agent loop. The model proposes a tool call; WE validate and run it.
# The model never executes anything by itself.
ctx = build_context(user_message, history, docs=[])
reply = model_call(ctx, tools=list(TOOLS))
if reply.get("tool_call"):
name, args = reply["tool_call"]["name"], reply["tool_call"]["args"]
if name not in TOOLS: # allow-list check first
return "That capability isn't available."
result = TOOLS[name]["handler"](**args) # our code, our validation
return model_call(ctx + [tool_result(result)]) # answer with real data
return reply["text"]
The security rule from our RMA article is universal: the model proposes, your code disposes. Validate every argument against server-side state, allow-list the tool set, and keep state-changing operations in deterministic code the model can request but never directly trigger.
The failure-mode catalog (learn these by heart)
- Hallucination — fluent, confident, false. Mitigations: RAG with citations, lower temperature for factual tasks, "answer only from the provided context" instructions, and evals that check factual claims.
- Prompt injection — hostile text in retrieved documents or user input instructing the model ("ignore previous rules"). Mitigations: treat all retrieved text as data, never as instructions; scope the agent (part 7's layers); keep tools read-only where possible.
- Sycophancy & over-refusal — the alignment trade-offs showing up in production: models that agree with wrong users, or refuse benign ones. Mitigate with eval cases for both, and prompt examples of the right stance.
- Latency & cost — the KV-cache economics from part 5 scale with context. Route easy requests to small models, cache aggressively, and keep contexts as short as the task allows.
- Non-determinism — temperature > 0 means the same input can produce different outputs. For tests, pin temperature 0; for UX, design flows that tolerate variation; for audits, log everything.
The checklist for shipping
Before any LLM feature goes live: a frozen golden eval (part 6) with the failure list reviewed; an owner for the prompt (it's code — version it, review it); guardrails on inputs and outputs; a cost ceiling per request; a fallback path when the provider is down; and a human review queue for the first weeks. None of that is AI. All of it is why one AI feature becomes trusted and another becomes a liability.
That completes the series. You now know what happens between "user types a question" and "answer appears" — tokenization into bytes, attention over context, a stack of transformer blocks trained under scaling laws, sampled token by token with a KV cache, aligned with human feedback, evaluated against benchmarks that can lie — and you've written the essential ideas yourself in Python. The best next step is the one you take with your own corpus: build a tiny model, break it, and fix it. Everything above will make sense retroactively.