<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel><title>AI Frontiers — Trust is earned, not given</title><link>https://www.bobhuang.com/series/ai-frontiers/</link><description>Reading the research and the engineering that turned language models into products: one paper, technique, or production problem per part, from the transformer paper to what I would do differently shipping AI systems.</description><language>en</language><lastBuildDate>17 Sep 2026 12:00:00 GMT</lastBuildDate><atom:link href="https://www.bobhuang.com/series/ai-frontiers/feed.xml" rel="self" type="application/rss+xml"/><item><title>AI Frontiers, part 65: A field guide to shipping AI systems — what I would do differently</title><link>https://www.bobhuang.com/blog/979/ai-frontiers-part-65-field-guide-to-shipping-ai-systems/</link><guid>https://www.bobhuang.com/blog/979/ai-frontiers-part-65-field-guide-to-shipping-ai-systems/</guid><pubDate>02 Jul 2026 12:00:00 GMT</pubDate><description>The last entry in the series: the ten practices I would keep, the five things I got wrong, and what five years of watching this stack get built actually taught me about shipping software that thinks.</description></item><item><title>AI Frontiers, part 64: The economics of applied AI — where the margin actually is</title><link>https://www.bobhuang.com/blog/978/ai-frontiers-part-64-economics-of-applied-ai-where-the-margin-is/</link><guid>https://www.bobhuang.com/blog/978/ai-frontiers-part-64-economics-of-applied-ai-where-the-margin-is/</guid><pubDate>04 Jun 2026 12:00:00 GMT</pubDate><description>Token prices fell by orders of magnitude and most AI products still do not have software margins. The full cost stack, the measured productivity evidence, and the four things that actually protect an applied AI business.</description></item><item><title>AI Frontiers, part 63: Interoperability — standards, protocols, and portable context</title><link>https://www.bobhuang.com/blog/977/ai-frontiers-part-63-interoperability-standards-and-portable-context/</link><guid>https://www.bobhuang.com/blog/977/ai-frontiers-part-63-interoperability-standards-and-portable-context/</guid><pubDate>07 May 2026 12:00:00 GMT</pubDate><description>MCP, A2A, OpenAPI, OpenTelemetry: the agent stack is acquiring protocols faster than it is acquiring practice. What to standardize, what to keep proprietary, and why your evaluation set is the most portable asset you own.</description></item><item><title>AI Frontiers, part 62: Human-in-the-loop design — review queues and trust calibration</title><link>https://www.bobhuang.com/blog/976/ai-frontiers-part-62-human-in-the-loop-review-queues-trust-calibration/</link><guid>https://www.bobhuang.com/blog/976/ai-frontiers-part-62-human-in-the-loop-review-queues-trust-calibration/</guid><pubDate>09 Apr 2026 12:00:00 GMT</pubDate><description>A human reviewer is a system component with a latency budget, an error distribution, and a fatigue curve. Designing the queue, the interface, and the escalation rules is where human-in-the-loop systems succeed or quietly become rubber stamps.</description></item><item><title>AI Frontiers, part 61: Retrieval at scale — sharding, freshness, and cache invalidation</title><link>https://www.bobhuang.com/blog/975/ai-frontiers-part-61-retrieval-at-scale-sharding-and-freshness/</link><guid>https://www.bobhuang.com/blog/975/ai-frontiers-part-61-retrieval-at-scale-sharding-and-freshness/</guid><pubDate>05 Mar 2026 12:00:00 GMT</pubDate><description>Retrieval systems look elegant at a million vectors and get difficult at a billion. Approximate indexes, shard routing, the freshness problem that vector stores handle badly, and the invalidation rules that decide what your users actually see.</description></item><item><title>AI Frontiers, part 60: Continuous evaluation in production — drift, regressions, rollbacks</title><link>https://www.bobhuang.com/blog/974/ai-frontiers-part-60-continuous-evaluation-in-production-drift/</link><guid>https://www.bobhuang.com/blog/974/ai-frontiers-part-60-continuous-evaluation-in-production-drift/</guid><pubDate>05 Feb 2026 12:00:00 GMT</pubDate><description>A pre-launch evaluation set is stale the day it ships. What to monitor, how to get labels cheaply from production, how to detect drift without drowning in false alarms, and how to make rollback a one-line change.</description></item><item><title>AI Frontiers, part 59: Model routing — choosing a model per request, and proving it works</title><link>https://www.bobhuang.com/blog/973/ai-frontiers-part-59-model-routing-choosing-a-model-per-request/</link><guid>https://www.bobhuang.com/blog/973/ai-frontiers-part-59-model-routing-choosing-a-model-per-request/</guid><pubDate>08 Jan 2026 12:00:00 GMT</pubDate><description>Once you have more than one model, every request is a decision. How routers are built, why offline accuracy numbers overstate them, and how to evaluate a policy you can only observe one arm of.</description></item><item><title>AI Frontiers, part 58: Data residency and sovereignty — where inference is allowed to run</title><link>https://www.bobhuang.com/blog/972/ai-frontiers-part-58-data-residency-and-sovereignty-inference/</link><guid>https://www.bobhuang.com/blog/972/ai-frontiers-part-58-data-residency-and-sovereignty-inference/</guid><pubDate>04 Dec 2025 12:00:00 GMT</pubDate><description>Model inference creates data flows that no cloud-region dropdown covers. A practical map of the legal stack, the deployment options, and the surprising places your prompts actually travel.</description></item><item><title>AI Frontiers, part 57: Security of agentic systems — permissions, sandboxes, and blast radius</title><link>https://www.bobhuang.com/blog/971/ai-frontiers-part-57-security-of-agentic-systems-permissions-and-sandboxes/</link><guid>https://www.bobhuang.com/blog/971/ai-frontiers-part-57-security-of-agentic-systems-permissions-and-sandboxes/</guid><pubDate>06 Nov 2025 12:00:00 GMT</pubDate><description>An agent is a program that takes instructions from its input, holds credentials, and can act. That combination is a new security model, and the defenses that work come from 1975, not from prompt engineering.</description></item><item><title>AI Frontiers, part 56: Science with models — from conjecture to verification</title><link>https://www.bobhuang.com/blog/970/ai-frontiers-part-56-science-with-models-from-conjecture-to-verification/</link><guid>https://www.bobhuang.com/blog/970/ai-frontiers-part-56-science-with-models-from-conjecture-to-verification/</guid><pubDate>09 Oct 2025 12:00:00 GMT</pubDate><description>Where models genuinely accelerate science — search with a verifier, fast surrogates, closed experimental loops — and why every real result still ends at an experiment that a model cannot run.</description></item><item><title>AI Frontiers, part 55: Biology and protein models — what AlphaFold changed</title><link>https://www.bobhuang.com/blog/969/ai-frontiers-part-55-biology-and-protein-models-what-alphafold-changed/</link><guid>https://www.bobhuang.com/blog/969/ai-frontiers-part-55-biology-and-protein-models-what-alphafold-changed/</guid><pubDate>04 Sep 2025 12:00:00 GMT</pubDate><description>Protein structure prediction was the first scientific field where a model genuinely changed the daily work. A tour of AlphaFold, protein language models, generative design — and the limits that wet-lab validation keeps rediscovering.</description></item><item><title>AI Frontiers, part 54: Robotics and embodied models — the sim-to-real gap</title><link>https://www.bobhuang.com/blog/968/ai-frontiers-part-54-robotics-and-embodied-models-sim-to-real-gap/</link><guid>https://www.bobhuang.com/blog/968/ai-frontiers-part-54-robotics-and-embodied-models-sim-to-real-gap/</guid><pubDate>07 Aug 2025 12:00:00 GMT</pubDate><description>Vision-language-action models turned robot demos into something that looks like general manipulation. The distance between a demo and a deployment is physics, data, and reliability — and it is measured in nines, not benchmarks.</description></item><item><title>AI Frontiers, part 53: Reasoning distillation — putting thinking into small models</title><link>https://www.bobhuang.com/blog/967/ai-frontiers-part-53-reasoning-distillation-thinking-into-small-models/</link><guid>https://www.bobhuang.com/blog/967/ai-frontiers-part-53-reasoning-distillation-thinking-into-small-models/</guid><pubDate>03 Jul 2025 12:00:00 GMT</pubDate><description>DeepSeek-R1's release made chain-of-thought distillation a production technique: a small model can learn to reason from a large model's traces. What transfers, what only looks like it transferred, and how to run the loop without inflating your latency budget.</description></item><item><title>AI Frontiers, part 52: Long-context retrieval — position, recency, and what gets lost</title><link>https://www.bobhuang.com/blog/966/ai-frontiers-part-52-long-context-retrieval-position-and-recency/</link><guid>https://www.bobhuang.com/blog/966/ai-frontiers-part-52-long-context-retrieval-position-and-recency/</guid><pubDate>05 Jun 2025 12:00:00 GMT</pubDate><description>A million-token window did not replace retrieval. Why models lose information in the middle of long inputs, how context windows were extended, and the engineering pattern that actually works at scale.</description></item><item><title>AI Frontiers, part 51: The evaluation of agents — tasks, traces, and cost-aware scoring</title><link>https://www.bobhuang.com/blog/965/ai-frontiers-part-51-evaluation-of-agents-traces-and-cost-aware-scoring/</link><guid>https://www.bobhuang.com/blog/965/ai-frontiers-part-51-evaluation-of-agents-traces-and-cost-aware-scoring/</guid><pubDate>24 Apr 2025 12:00:00 GMT</pubDate><description>Agent benchmarks told a story of rapid progress that mostly reflected measurement choices. What to hold out, what to count, why pass^k beats pass@k, and how to build an agent evaluation you can trust.</description></item><item><title>AI Frontiers, part 50: Inference cost engineering — caching, routing, and unit economics</title><link>https://www.bobhuang.com/blog/964/ai-frontiers-part-50-inference-cost-engineering-unit-economics/</link><guid>https://www.bobhuang.com/blog/964/ai-frontiers-part-50-inference-cost-engineering-unit-economics/</guid><pubDate>03 Apr 2025 12:00:00 GMT</pubDate><description>Inference cost is not a line item, it is an architectural property. A cost model for LLM systems, the four levers that actually move the number, and the unit economics that decide whether an AI feature is a business.</description></item><item><title>AI Frontiers, part 49: Agent observability — tracing, replay, and debugging loops</title><link>https://www.bobhuang.com/blog/963/ai-frontiers-part-49-agent-observability-tracing-and-replay/</link><guid>https://www.bobhuang.com/blog/963/ai-frontiers-part-49-agent-observability-tracing-and-replay/</guid><pubDate>06 Mar 2025 12:00:00 GMT</pubDate><description>An agent that fails in production tells you nothing unless you captured the trajectory. What to record, how to make runs replayable, and why the trace is the most valuable artifact an AI system produces.</description></item><item><title>AI Frontiers, part 48: Fine-tuning versus prompting — the decision nobody makes explicitly</title><link>https://www.bobhuang.com/blog/962/ai-frontiers-part-48-fine-tuning-versus-prompting-the-decision/</link><guid>https://www.bobhuang.com/blog/962/ai-frontiers-part-48-fine-tuning-versus-prompting-the-decision/</guid><pubDate>06 Feb 2025 12:00:00 GMT</pubDate><description>Fine-tuning and prompting are not competitors; they change different things. A decision procedure for choosing between them, and the evidence on when each one actually wins.</description></item><item><title>AI Frontiers, part 47: Multi-agent systems — when decomposition helps and when it hurts</title><link>https://www.bobhuang.com/blog/961/ai-frontiers-part-47-multi-agent-systems-when-decomposition-helps/</link><guid>https://www.bobhuang.com/blog/961/ai-frontiers-part-47-multi-agent-systems-when-decomposition-helps/</guid><pubDate>19 Dec 2024 12:00:00 GMT</pubDate><description>Between 2023 and 2024 the field tried to turn one model into a team. The results split cleanly: decomposition wins when the subproblems are genuinely separable and loses when the bottleneck is verification.</description></item><item><title>AI Frontiers, part 46: Small models on device — NPUs, memory budgets, and privacy</title><link>https://www.bobhuang.com/blog/960/ai-frontiers-part-46-small-models-on-device-npus-and-privacy/</link><guid>https://www.bobhuang.com/blog/960/ai-frontiers-part-46-small-models-on-device-npus-and-privacy/</guid><pubDate>27 Nov 2024 12:00:00 GMT</pubDate><description>On-device inference is a memory-bandwidth problem wearing a TOPS costume. What fits on a phone, why NPUs are less magic than the spec sheet suggests, and why privacy regulation may be the strongest argument for local inference.</description></item><item><title>AI Frontiers, part 45: Structured output — schemas, grammars, and constrained decoding</title><link>https://www.bobhuang.com/blog/959/ai-frontiers-part-45-structured-output-schemas-constrained-decoding/</link><guid>https://www.bobhuang.com/blog/959/ai-frontiers-part-45-structured-output-schemas-constrained-decoding/</guid><pubDate>07 Nov 2024 12:00:00 GMT</pubDate><description>Asking a model to return JSON politely is not a contract. A tour of constrained decoding — regex, grammars, logit masking — and the harder question of whether valid output is actually correct output.</description></item><item><title>AI Frontiers, part 44: Synthetic preference data and the RLAIF flywheel</title><link>https://www.bobhuang.com/blog/958/ai-frontiers-part-44-synthetic-preference-data-rlaif-flywheel/</link><guid>https://www.bobhuang.com/blog/958/ai-frontiers-part-44-synthetic-preference-data-rlaif-flywheel/</guid><pubDate>10 Oct 2024 12:00:00 GMT</pubDate><description>Human preference labels cost dollars each and hours of latency; model-generated ones cost fractions of a cent. How RLAIF, UltraFeedback, and self-rewarding pipelines work — and the four ways the flywheel spins itself into a wall.</description></item><item><title>AI Frontiers, part 43: Embeddings as infrastructure — search, clustering, and drift</title><link>https://www.bobhuang.com/blog/957/ai-frontiers-part-43-embeddings-as-infrastructure-search-clustering-drift/</link><guid>https://www.bobhuang.com/blog/957/ai-frontiers-part-43-embeddings-as-infrastructure-search-clustering-drift/</guid><pubDate>05 Sep 2024 12:00:00 GMT</pubDate><description>The unglamorous layer under every retrieval system: how embedding models are trained, why dense similarity is not relevance, what breaks at index scale, and how embedding drift quietly rots a search product.</description></item><item><title>AI Frontiers, part 42: Evaluation-driven development — building your own test set</title><link>https://www.bobhuang.com/blog/956/ai-frontiers-part-42-evaluation-driven-development-build-your-own-test-set/</link><guid>https://www.bobhuang.com/blog/956/ai-frontiers-part-42-evaluation-driven-development-build-your-own-test-set/</guid><pubDate>01 Aug 2024 12:00:00 GMT</pubDate><description>The one practice that separates teams that improve from teams that churn: a small, private, versioned evaluation set built from your own traffic, with a written rubric and a cost column.</description></item><item><title>AI Frontiers, part 41: Voice interfaces — ASR, TTS, and the latency budget</title><link>https://www.bobhuang.com/blog/955/ai-frontiers-part-41-voice-interfaces-asr-tts-latency-budget/</link><guid>https://www.bobhuang.com/blog/955/ai-frontiers-part-41-voice-interfaces-asr-tts-latency-budget/</guid><pubDate>11 Jul 2024 12:00:00 GMT</pubDate><description>Speech recognition is largely solved and speech synthesis is largely solved. The hard part is the hundred milliseconds of turn-taking between them, which is why good voice interfaces are rarer than good speech models.</description></item><item><title>AI Frontiers, part 40: Data curation — deduplication, filtering, and why clean beats big</title><link>https://www.bobhuang.com/blog/954/ai-frontiers-part-40-data-curation-deduplication-clean-beats-big/</link><guid>https://www.bobhuang.com/blog/954/ai-frontiers-part-40-data-curation-deduplication-clean-beats-big/</guid><pubDate>06 Jun 2024 12:00:00 GMT</pubDate><description>The part of pretraining that decides the model's ceiling and gets the least attention. Deduplication, quality classifiers and the biases they encode, multilingual coverage, and the provenance problem underneath it all.</description></item><item><title>AI Frontiers, part 39: The GPU supply chain and the cost of a training cluster</title><link>https://www.bobhuang.com/blog/953/ai-frontiers-part-39-gpu-supply-chain-cost-of-a-training-cluster/</link><guid>https://www.bobhuang.com/blog/953/ai-frontiers-part-39-gpu-supply-chain-cost-of-a-training-cluster/</guid><pubDate>02 May 2024 12:00:00 GMT</pubDate><description>Accelerators, high-bandwidth memory, advanced packaging, interconnect, power and permits. Why the constraints on frontier training are industrial rather than algorithmic, and what that does to the market.</description></item><item><title>AI Frontiers, part 38: Training infrastructure — FlashAttention, FSDP, and the plumbing of scale</title><link>https://www.bobhuang.com/blog/952/ai-frontiers-part-38-training-infrastructure-flash-attention-fsdp/</link><guid>https://www.bobhuang.com/blog/952/ai-frontiers-part-38-training-infrastructure-flash-attention-fsdp/</guid><pubDate>04 Apr 2024 12:00:00 GMT</pubDate><description>Training a frontier model is a distributed systems problem wearing a machine learning costume. Data, tensor, pipeline and expert parallelism, activation recomputation, and the failure rates that decide whether a run finishes.</description></item><item><title>AI Frontiers, part 37: The long tail of licensing — what open weights actually permit</title><link>https://www.bobhuang.com/blog/951/ai-frontiers-part-37-long-tail-of-licensing-what-open-weights-permit/</link><guid>https://www.bobhuang.com/blog/951/ai-frontiers-part-37-long-tail-of-licensing-what-open-weights-permit/</guid><pubDate>07 Mar 2024 12:00:00 GMT</pubDate><description>Two models both described as open can carry obligations that differ in every dimension a business cares about. Reading the licenses, the training data provenance problem, and the rights questions courts had not yet answered.</description></item><item><title>AI Frontiers, part 36: Benchmarks and contamination — measuring models honestly</title><link>https://www.bobhuang.com/blog/950/ai-frontiers-part-36-benchmarks-and-contamination-measuring-honestly/</link><guid>https://www.bobhuang.com/blog/950/ai-frontiers-part-36-benchmarks-and-contamination-measuring-honestly/</guid><pubDate>08 Feb 2024 12:00:00 GMT</pubDate><description>A benchmark that becomes a training target stops measuring capability. How contamination happens, how to detect it, and what a private evaluation set buys you that no leaderboard can.</description></item><item><title>AI Frontiers, part 35: Mistral, Mixtral, and the efficiency-first open ecosystem</title><link>https://www.bobhuang.com/blog/949/ai-frontiers-part-35-mistral-mixtral-efficiency-first-open-ecosystem/</link><guid>https://www.bobhuang.com/blog/949/ai-frontiers-part-35-mistral-mixtral-efficiency-first-open-ecosystem/</guid><pubDate>04 Jan 2024 12:00:00 GMT</pubDate><description>A 7B model with sliding-window attention, then a sparse mixture-of-experts that beat much larger dense models at a fraction of the active compute. Why efficiency became the open ecosystem's organizing principle.</description></item><item><title>AI Frontiers, part 34: Prompt injection — the vulnerability class nobody can patch</title><link>https://www.bobhuang.com/blog/948/ai-frontiers-part-34-prompt-injection-vulnerability-nobody-can-patch/</link><guid>https://www.bobhuang.com/blog/948/ai-frontiers-part-34-prompt-injection-vulnerability-nobody-can-patch/</guid><pubDate>07 Dec 2023 12:00:00 GMT</pubDate><description>Instructions and data share one channel, so text a model reads can steer it. Why remote code execution is the right frame, what the research established in 2023, and the defenses that reduce rather than eliminate the risk.</description></item><item><title>AI Frontiers, part 33: Mechanistic interpretability — looking inside the black box</title><link>https://www.bobhuang.com/blog/947/ai-frontiers-part-33-mechanistic-interpretability-inside-the-black-box/</link><guid>https://www.bobhuang.com/blog/947/ai-frontiers-part-33-mechanistic-interpretability-inside-the-black-box/</guid><pubDate>02 Nov 2023 12:00:00 GMT</pubDate><description>Induction heads, superposition, sparse autoencoders and circuits that can be edited: what a decade of trying to read a neural network actually produced, and why it matters more than a visualization.</description></item><item><title>AI Frontiers, part 32: Serving is the product — PagedAttention and continuous batching</title><link>https://www.bobhuang.com/blog/946/ai-frontiers-part-32-serving-is-the-product-paged-attention-batching/</link><guid>https://www.bobhuang.com/blog/946/ai-frontiers-part-32-serving-is-the-product-paged-attention-batching/</guid><pubDate>05 Oct 2023 12:00:00 GMT</pubDate><description>vLLM's memory manager took an operating-system idea and made inference several times cheaper. Why the serving layer, not the model, decides a product's unit economics.</description></item><item><title>AI Frontiers, part 31: AutoGPT and the first agent wave</title><link>https://www.bobhuang.com/blog/945/ai-frontiers-part-31-autogpt-and-the-first-agent-wave/</link><guid>https://www.bobhuang.com/blog/945/ai-frontiers-part-31-autogpt-and-the-first-agent-wave/</guid><pubDate>07 Sep 2023 12:00:00 GMT</pubDate><description>In spring 2023 a script that looped a language model against itself became the fastest-growing repository in GitHub history. Why the autonomous agent dream arrived early, failed publicly, and still changed how we build.</description></item><item><title>AI Frontiers, part 30: Tool use arrives — function calling and the plugin experiment</title><link>https://www.bobhuang.com/blog/944/ai-frontiers-part-30-tool-use-arrives-function-calling-plugins/</link><guid>https://www.bobhuang.com/blog/944/ai-frontiers-part-30-tool-use-arrives-function-calling-plugins/</guid><pubDate>10 Aug 2023 12:00:00 GMT</pubDate><description>In June 2023 models learned to ask for a function by name with typed arguments. Plugins tried to build an app store on top, and failed. What survived was the primitive: a model that can act rather than only answer.</description></item><item><title>AI Frontiers, part 29: The open-weight leap — LLaMA, Alpaca, and the leak that built an ecosystem</title><link>https://www.bobhuang.com/blog/943/ai-frontiers-part-29-open-weight-leap-llama-alpaca-ecosystem/</link><guid>https://www.bobhuang.com/blog/943/ai-frontiers-part-29-open-weight-leap-llama-alpaca-ecosystem/</guid><pubDate>06 Jul 2023 12:00:00 GMT</pubDate><description>Meta released LLaMA to researchers in February 2023 under a gated license, and the weights circulated within a week. What followed — Alpaca, Vicuna, RedPajama, Dolly, StarCoder — set the template for open models for years.</description></item><item><title>AI Frontiers, part 28: Diffusion models — how image generation actually works</title><link>https://www.bobhuang.com/blog/942/ai-frontiers-part-28-diffusion-models-how-image-generation-works/</link><guid>https://www.bobhuang.com/blog/942/ai-frontiers-part-28-diffusion-models-how-image-generation-works/</guid><pubDate>15 Jun 2023 12:00:00 GMT</pubDate><description>From nonequilibrium thermodynamics to latent diffusion and ControlNet: the mechanics of denoising, classifier-free guidance, why latent space made image generation affordable, and the licensing mess underneath.</description></item><item><title>AI Frontiers, part 27: Alignment before the assistant era — RLHF and its discontents</title><link>https://www.bobhuang.com/blog/941/ai-frontiers-part-27-alignment-before-the-assistant-era-rlhf/</link><guid>https://www.bobhuang.com/blog/941/ai-frontiers-part-27-alignment-before-the-assistant-era-rlhf/</guid><pubDate>18 May 2023 12:00:00 GMT</pubDate><description>Reinforcement learning from human feedback is why language models became assistants. It also optimizes against a proxy, which is why reward hacking, sycophancy and overoptimization were predicted before they were observed.</description></item><item><title>AI Frontiers, part 26: Hallucination — the failure mode that will not go away</title><link>https://www.bobhuang.com/blog/940/ai-frontiers-part-26-hallucination-failure-mode-that-will-not-go-away/</link><guid>https://www.bobhuang.com/blog/940/ai-frontiers-part-26-hallucination-failure-mode-that-will-not-go-away/</guid><pubDate>13 Apr 2023 12:00:00 GMT</pubDate><description>Why a model that has read the internet invents citations: the sources of hallucination, the calibration evidence, and the mitigations that actually work — retrieval, verification, and knowing when to abstain.</description></item><item><title>AI Frontiers, part 25: Instruction tuning — from FLAN to Alpaca</title><link>https://www.bobhuang.com/blog/939/ai-frontiers-part-25-instruction-tuning-flan-to-alpaca/</link><guid>https://www.bobhuang.com/blog/939/ai-frontiers-part-25-instruction-tuning-flan-to-alpaca/</guid><pubDate>30 Mar 2023 12:00:00 GMT</pubDate><description>A 137B model that had been instruction-tuned beat a 175B model that had not. How FLAN, T0, InstructGPT and the Flan Collection turned a base model into an assistant — and why Alpaca cost $600.</description></item><item><title>AI Frontiers, part 24: The prompt is not the product — in-context learning and its limits</title><link>https://www.bobhuang.com/blog/938/ai-frontiers-part-24-prompt-is-not-the-product-in-context-learning-limits/</link><guid>https://www.bobhuang.com/blog/938/ai-frontiers-part-24-prompt-is-not-the-product-in-context-learning-limits/</guid><pubDate>16 Feb 2023 12:00:00 GMT</pubDate><description>In early 2023 prompt engineering looked like a discipline and a moat. The research said otherwise: random labels work, order matters more than wording, and in-context learning looks like task location rather than instruction following.</description></item><item><title>AI Frontiers, part 23: Where the frontier goes next</title><link>https://www.bobhuang.com/blog/937/ai-frontiers-part-23-where-the-frontier-goes-next/</link><guid>https://www.bobhuang.com/blog/937/ai-frontiers-part-23-where-the-frontier-goes-next/</guid><pubDate>17 Sep 2026 12:00:00 GMT</pubDate><description>The last entry in the series: the four resources that bound progress, why evaluation is now the binding constraint, what the open-weights diffusion means for moats, and what I would build next.</description></item><item><title>AI Frontiers, part 22: Constitutional AI and scalable oversight</title><link>https://www.bobhuang.com/blog/936/ai-frontiers-part-22-constitutional-ai-scalable-oversight/</link><guid>https://www.bobhuang.com/blog/936/ai-frontiers-part-22-constitutional-ai-scalable-oversight/</guid><pubDate>16 Jul 2026 12:00:00 GMT</pubDate><description>RLHF does not scale, and the field knows it. From preference learning to AI feedback, debate, self-critique and deliberation training — the methods for supervising systems that humans can no longer fully check.</description></item><item><title>AI Frontiers, part 21: Synthetic data and the model-collapse debate</title><link>https://www.bobhuang.com/blog/935/ai-frontiers-part-21-synthetic-data-model-collapse-debate/</link><guid>https://www.bobhuang.com/blog/935/ai-frontiers-part-21-synthetic-data-model-collapse-debate/</guid><pubDate>21 May 2026 12:00:00 GMT</pubDate><description>If models train on model output, do they degrade? The recursion argument, the accumulation counterargument, verification as the real fix, and the distillation economy that made synthetic data unavoidable.</description></item><item><title>AI Frontiers, part 20: The RAG-to-agentic-RAG evolution</title><link>https://www.bobhuang.com/blog/934/ai-frontiers-part-20-rag-to-agentic-rag-evolution/</link><guid>https://www.bobhuang.com/blog/934/ai-frontiers-part-20-rag-to-agentic-rag-evolution/</guid><pubDate>19 Mar 2026 12:00:00 GMT</pubDate><description>Retrieval-augmented generation grew from one embedding lookup into a planning loop with query rewriting, self-critique and graph structure. What changed, what it costs, and when long context beats retrieval outright.</description></item><item><title>AI Frontiers, part 19: KV-cache engineering — the memory wall of long context</title><link>https://www.bobhuang.com/blog/933/ai-frontiers-part-19-kv-cache-engineering-memory-wall/</link><guid>https://www.bobhuang.com/blog/933/ai-frontiers-part-19-kv-cache-engineering-memory-wall/</guid><pubDate>15 Jan 2026 12:00:00 GMT</pubDate><description>Long context is not a model feature, it is a memory budget. The arithmetic of the KV cache, grouped-query attention, latent attention, paging, prefix caching and eviction — and why serving is where context length is actually decided.</description></item><item><title>AI Frontiers, part 18: Speculative decoding — faster inference for free</title><link>https://www.bobhuang.com/blog/932/ai-frontiers-part-18-speculative-decoding-faster-inference-for-free/</link><guid>https://www.bobhuang.com/blog/932/ai-frontiers-part-18-speculative-decoding-faster-inference-for-free/</guid><pubDate>20 Nov 2025 12:00:00 GMT</pubDate><description>Draft, then verify: the only widely deployed trick that makes inference two to three times faster without changing the output distribution at all. How it works, where it stops working, and what it reveals about serving.</description></item><item><title>AI Frontiers, part 17: Quantization — running 70B on a laptop</title><link>https://www.bobhuang.com/blog/931/ai-frontiers-part-17-quantization-running-70b-on-a-laptop/</link><guid>https://www.bobhuang.com/blog/931/ai-frontiers-part-17-quantization-running-70b-on-a-laptop/</guid><pubDate>18 Sep 2025 12:00:00 GMT</pubDate><description>From GPTQ and AWQ to GGUF k-quants and 1.58-bit BitNet: how weight quantization actually works, what it costs in quality, and the memory arithmetic that decides what fits on your desk.</description></item><item><title>AI Frontiers, part 16: Test-time scaling and the reasoning models</title><link>https://www.bobhuang.com/blog/930/ai-frontiers-part-16-test-time-scaling-reasoning-models/</link><guid>https://www.bobhuang.com/blog/930/ai-frontiers-part-16-test-time-scaling-reasoning-models/</guid><pubDate>17 Jul 2025 12:00:00 GMT</pubDate><description>DeepSeek-R1 published its reinforcement-learning recipe in January 2025 and the reasoning-model playbook stopped being a secret. What actually works, where the compute goes, and when thinking longer is a waste.</description></item><item><title>AI Frontiers, part 15: Computer use and GUI agents</title><link>https://www.bobhuang.com/blog/929/ai-frontiers-part-15-computer-use-gui-agents/</link><guid>https://www.bobhuang.com/blog/929/ai-frontiers-part-15-computer-use-gui-agents/</guid><pubDate>15 May 2025 12:00:00 GMT</pubDate><description>When there is no API, the screen becomes the interface. How screenshot-driven agents perceive, ground and act — and why per-step reliability, not intelligence, is the wall they hit.</description></item><item><title>AI Frontiers, part 14: The agentic coding turn — from autocomplete to teammates</title><link>https://www.bobhuang.com/blog/928/ai-frontiers-part-14-agentic-coding-turn/</link><guid>https://www.bobhuang.com/blog/928/ai-frontiers-part-14-agentic-coding-turn/</guid><pubDate>20 Mar 2025 12:00:00 GMT</pubDate><description>SWE-bench went from 4% to over 60% in eighteen months and coding assistants grew hands. What actually changed in the agent loop, what the productivity evidence supports, and why review is now the bottleneck.</description></item><item><title>AI Frontiers, part 13: DeepSeek-V3 and the cost-efficiency shock</title><link>https://www.bobhuang.com/blog/927/ai-frontiers-part-13-deepseek-v3-cost-efficiency-shock/</link><guid>https://www.bobhuang.com/blog/927/ai-frontiers-part-13-deepseek-v3-cost-efficiency-shock/</guid><pubDate>16 Jan 2025 12:00:00 GMT</pubDate><description>DeepSeek-V3 reported a frontier-class training run for about $5.6M of GPU time. Reading the technical report: MLA, fine-grained experts, auxiliary-loss-free balancing, FP8, and what the headline number really measures.</description></item><item><title>AI Frontiers, part 12: Model Context Protocol — plumbing for the agent era</title><link>https://www.bobhuang.com/blog/926/ai-frontiers-part-12-model-context-protocol-plumbing-agent-era/</link><guid>https://www.bobhuang.com/blog/926/ai-frontiers-part-12-model-context-protocol-plumbing-agent-era/</guid><pubDate>05 Dec 2024 12:00:00 GMT</pubDate><description>Anthropic published the Model Context Protocol on 25 November 2024. Why a JSON-RPC standard for tools was overdue, what it specifies, and why its security model is the interesting part.</description></item><item><title>AI Frontiers, part 11: OpenAI o1 and inference-time compute</title><link>https://www.bobhuang.com/blog/925/ai-frontiers-part-11-openai-o1-inference-time-compute/</link><guid>https://www.bobhuang.com/blog/925/ai-frontiers-part-11-openai-o1-inference-time-compute/</guid><pubDate>19 Sep 2024 12:00:00 GMT</pubDate><description>On 12 September 2024 OpenAI shipped o1 — a model that spends more compute thinking before it answers. The research behind it: chain of thought, process reward models, and the economics of buying capability at inference time.</description></item><item><title>AI Frontiers, part 10: Multimodal turns one — vision, audio, and the GPT-4o moment</title><link>https://www.bobhuang.com/blog/924/ai-frontiers-part-10-multimodal-turns-one-vision-audio-gpt-4o/</link><guid>https://www.bobhuang.com/blog/924/ai-frontiers-part-10-multimodal-turns-one-vision-audio-gpt-4o/</guid><pubDate>18 Jul 2024 12:00:00 GMT</pubDate><description>GPT-4o landed a year after GPT-4 first accepted an image. How CLIP, Flamingo and LLaVA taught language models to see, why audio is the harder modality, and what multimodality changes in production.</description></item><item><title>AI Frontiers, part 9: Small language models — Phi, Gemma, and the desktop class</title><link>https://www.bobhuang.com/blog/923/ai-frontiers-part-9-small-language-models-phi-gemma-desktop-class/</link><guid>https://www.bobhuang.com/blog/923/ai-frontiers-part-9-small-language-models-phi-gemma-desktop-class/</guid><pubDate>16 May 2024 12:00:00 GMT</pubDate><description>By mid-2024 a 3.8B model could match GPT-3.5-class benchmarks. Microsoft's Phi line, Google's Gemma, and the data-quality thesis behind them — plus what small models change for cost, privacy, and architecture.</description></item><item><title>AI Frontiers, part 8: Mixture of experts and the memory-feasible frontier</title><link>https://www.bobhuang.com/blog/922/ai-frontiers-part-8-mixture-of-experts-memory-feasible-frontier/</link><guid>https://www.bobhuang.com/blog/922/ai-frontiers-part-8-mixture-of-experts-memory-feasible-frontier/</guid><pubDate>21 Mar 2024 12:00:00 GMT</pubDate><description>Mixtral made mixture-of-experts famous in December 2023: more total knowledge, same per-token compute. The 1991 idea behind it, what Mixtral actually is, and why MoE trades your memory bill for your GPU-hours bill.</description></item><item><title>AI Frontiers, part 7: The context window wars — from 4K to 200K (and the road to a million)</title><link>https://www.bobhuang.com/blog/921/ai-frontiers-part-7-context-window-wars/</link><guid>https://www.bobhuang.com/blog/921/ai-frontiers-part-7-context-window-wars/</guid><pubDate>18 Jan 2024 12:00:00 GMT</pubDate><description>Context windows grew 100x in three years. Why position encodings were the bottleneck, how long-context models actually extend, and why bigger windows changed RAG economics without ending it.</description></item><item><title>AI Frontiers, part 6: Vector databases — the missing infrastructure layer</title><link>https://www.bobhuang.com/blog/920/ai-frontiers-part-6-vector-databases-infrastructure-layer/</link><guid>https://www.bobhuang.com/blog/920/ai-frontiers-part-6-vector-databases-infrastructure-layer/</guid><pubDate>16 Nov 2023 12:00:00 GMT</pubDate><description>In 2023 an infrastructure category appeared, raised hundreds of millions, and became the default neighbor of every LLM deployment. What a vector database actually does, why approximate nearest-neighbor search is the interesting part, and whether the category survives Postgres.</description></item><item><title>AI Frontiers, part 5: Retrieval-augmented generation before it was everywhere</title><link>https://www.bobhuang.com/blog/919/ai-frontiers-part-5-rag-before-it-was-everywhere/</link><guid>https://www.bobhuang.com/blog/919/ai-frontiers-part-5-rag-before-it-was-everywhere/</guid><pubDate>21 Sep 2023 12:00:00 GMT</pubDate><description>RAG spent 2023 going from an Oxford paper to the default enterprise pattern. Where it came from, why it beats fine-tuning for knowledge, and what the first year of production deployments got wrong.</description></item><item><title>AI Frontiers, part 4: Llama 2 and the open-weights turning point</title><link>https://www.bobhuang.com/blog/918/ai-frontiers-part-4-llama-2-open-weights-turning-point/</link><guid>https://www.bobhuang.com/blog/918/ai-frontiers-part-4-llama-2-open-weights-turning-point/</guid><pubDate>20 Jul 2023 12:00:00 GMT</pubDate><description>July 2023: Meta ships Llama 2 with a license businesses can actually sign. What the release contained, why the license terms were the real story, and how a chaotic month of benchmarks reshaped the open-model ecosystem.</description></item><item><title>AI Frontiers, part 3: LoRA and QLoRA — fine-tuning on a single GPU</title><link>https://www.bobhuang.com/blog/917/ai-frontiers-part-3-lora-and-qlora-fine-tuning-single-gpu/</link><guid>https://www.bobhuang.com/blog/917/ai-frontiers-part-3-lora-and-qlora-fine-tuning-single-gpu/</guid><pubDate>01 Jun 2023 12:00:00 GMT</pubDate><description>Two papers turned fine-tuning from a datacenter-only activity into something a person with one GPU can do: LoRA (2021) and QLoRA (May 2023). What they changed, how they work, and what the open-model ecosystem looked like the week Guanaco shipped.</description></item><item><title>AI Frontiers, part 2: GPT-4 and the emergent abilities debate</title><link>https://www.bobhuang.com/blog/916/ai-frontiers-part-2-gpt-4-emergent-abilities-debate/</link><guid>https://www.bobhuang.com/blog/916/ai-frontiers-part-2-gpt-4-emergent-abilities-debate/</guid><pubDate>04 May 2023 12:00:00 GMT</pubDate><description>GPT-4 arrived in March 2023, and two months later a paper argued its boldest abilities were a mirage. This article works through both sides of the emergence debate and what it means for people building on these models.</description></item><item><title>AI Frontiers, part 1: Attention is all you need — reading the transformer paper five years later</title><link>https://www.bobhuang.com/blog/915/ai-frontiers-part-1-attention-is-all-you-need-five-years-later/</link><guid>https://www.bobhuang.com/blog/915/ai-frontiers-part-1-attention-is-all-you-need-five-years-later/</guid><pubDate>19 Jan 2023 12:00:00 GMT</pubDate><description>A working engineer re-reads 'Attention Is All You Need' in 2023: what the transformer actually changed, what it merely enabled, and why the paper's 2017 vocabulary makes it harder to read today than it should be.</description></item></channel></rss>