AI Frontiers
Reading the research and the engineering that turned language models into products: one paper, technique, or production problem per part, from the transformer paper to what I would do differently shipping AI systems.
- Part 1Attention is all you need — reading the transformer paper five years later2023-01-19
- Part 2GPT-4 and the emergent abilities debate2023-05-04
- Part 3LoRA and QLoRA — fine-tuning on a single GPU2023-06-01
- Part 4Llama 2 and the open-weights turning point2023-07-20
- Part 5Retrieval-augmented generation before it was everywhere2023-09-21
- Part 6Vector databases — the missing infrastructure layer2023-11-16
- Part 7The context window wars — from 4K to 200K (and the road to a million)2024-01-18
- Part 8Mixture of experts and the memory-feasible frontier2024-03-21
- Part 9Small language models — Phi, Gemma, and the desktop class2024-05-16
- Part 10Multimodal turns one — vision, audio, and the GPT-4o moment2024-07-18
- Part 11OpenAI o1 and inference-time compute2024-09-19
- Part 12Model Context Protocol — plumbing for the agent era2024-12-05
- Part 13DeepSeek-V3 and the cost-efficiency shock2025-01-16
- Part 14The agentic coding turn — from autocomplete to teammates2025-03-20
- Part 15Computer use and GUI agents2025-05-15
- Part 16Test-time scaling and the reasoning models2025-07-17
- Part 17Quantization — running 70B on a laptop2025-09-18
- Part 18Speculative decoding — faster inference for free2025-11-20
- Part 19KV-cache engineering — the memory wall of long context2026-01-15
- Part 20The RAG-to-agentic-RAG evolution2026-03-19
- Part 21Synthetic data and the model-collapse debate2026-05-21
- Part 22Constitutional AI and scalable oversight2026-07-16
- Part 23Where the frontier goes next2026-09-17
- Part 24The prompt is not the product — in-context learning and its limits2023-02-16
- Part 25Instruction tuning — from FLAN to Alpaca2023-03-30
- Part 26Hallucination — the failure mode that will not go away2023-04-13
- Part 27Alignment before the assistant era — RLHF and its discontents2023-05-18
- Part 28Diffusion models — how image generation actually works2023-06-15
- Part 29The open-weight leap — LLaMA, Alpaca, and the leak that built an ecosystem2023-07-06
- Part 30Tool use arrives — function calling and the plugin experiment2023-08-10
- Part 31AutoGPT and the first agent wave2023-09-07
- Part 32Serving is the product — PagedAttention and continuous batching2023-10-05
- Part 33Mechanistic interpretability — looking inside the black box2023-11-02
- Part 34Prompt injection — the vulnerability class nobody can patch2023-12-07
- Part 35Mistral, Mixtral, and the efficiency-first open ecosystem2024-01-04
- Part 36Benchmarks and contamination — measuring models honestly2024-02-08
- Part 37The long tail of licensing — what open weights actually permit2024-03-07
- Part 38Training infrastructure — FlashAttention, FSDP, and the plumbing of scale2024-04-04
- Part 39The GPU supply chain and the cost of a training cluster2024-05-02
- Part 40Data curation — deduplication, filtering, and why clean beats big2024-06-06
- Part 41Voice interfaces — ASR, TTS, and the latency budget2024-07-11
- Part 42Evaluation-driven development — building your own test set2024-08-01
- Part 43Embeddings as infrastructure — search, clustering, and drift2024-09-05
- Part 44Synthetic preference data and the RLAIF flywheel2024-10-10
- Part 45Structured output — schemas, grammars, and constrained decoding2024-11-07
- Part 46Small models on device — NPUs, memory budgets, and privacy2024-11-27
- Part 47Multi-agent systems — when decomposition helps and when it hurts2024-12-19
- Part 48Fine-tuning versus prompting — the decision nobody makes explicitly2025-02-06
- Part 49Agent observability — tracing, replay, and debugging loops2025-03-06
- Part 50Inference cost engineering — caching, routing, and unit economics2025-04-03
- Part 51The evaluation of agents — tasks, traces, and cost-aware scoring2025-04-24
- Part 52Long-context retrieval — position, recency, and what gets lost2025-06-05
- Part 53Reasoning distillation — putting thinking into small models2025-07-03
- Part 54Robotics and embodied models — the sim-to-real gap2025-08-07
- Part 55Biology and protein models — what AlphaFold changed2025-09-04
- Part 56Science with models — from conjecture to verification2025-10-09
- Part 57Security of agentic systems — permissions, sandboxes, and blast radius2025-11-06
- Part 58Data residency and sovereignty — where inference is allowed to run2025-12-04
- Part 59Model routing — choosing a model per request, and proving it works2026-01-08
- Part 60Continuous evaluation in production — drift, regressions, rollbacks2026-02-05
- Part 61Retrieval at scale — sharding, freshness, and cache invalidation2026-03-05
- Part 62Human-in-the-loop design — review queues and trust calibration2026-04-09
- Part 63Interoperability — standards, protocols, and portable context2026-05-07
- Part 64The economics of applied AI — where the margin actually is2026-06-04
- Part 65A field guide to shipping AI systems — what I would do differently2026-07-02
Subscribe to this series: RSS feed · everything
Other series: React.js · Node.js · AI-Enabled RMA · DotNetCode · DevOps on AWS · DevOps on Azure · DevOps on Google Cloud · AI fundamentals · Stripe in C# · Shopify in C# · Windows domains · A math practice toy · all series