Trust is earned, not given

A different perspective

AI Frontiers

Reading the research and the engineering that turned language models into products: one paper, technique, or production problem per part, from the transformer paper to what I would do differently shipping AI systems.

65 parts · 2023-01-19 → 2026-09-17

  1. Part 1Attention is all you need — reading the transformer paper five years later2023-01-19
  2. Part 2GPT-4 and the emergent abilities debate2023-05-04
  3. Part 3LoRA and QLoRA — fine-tuning on a single GPU2023-06-01
  4. Part 4Llama 2 and the open-weights turning point2023-07-20
  5. Part 5Retrieval-augmented generation before it was everywhere2023-09-21
  6. Part 6Vector databases — the missing infrastructure layer2023-11-16
  7. Part 7The context window wars — from 4K to 200K (and the road to a million)2024-01-18
  8. Part 8Mixture of experts and the memory-feasible frontier2024-03-21
  9. Part 9Small language models — Phi, Gemma, and the desktop class2024-05-16
  10. Part 10Multimodal turns one — vision, audio, and the GPT-4o moment2024-07-18
  11. Part 11OpenAI o1 and inference-time compute2024-09-19
  12. Part 12Model Context Protocol — plumbing for the agent era2024-12-05
  13. Part 13DeepSeek-V3 and the cost-efficiency shock2025-01-16
  14. Part 14The agentic coding turn — from autocomplete to teammates2025-03-20
  15. Part 15Computer use and GUI agents2025-05-15
  16. Part 16Test-time scaling and the reasoning models2025-07-17
  17. Part 17Quantization — running 70B on a laptop2025-09-18
  18. Part 18Speculative decoding — faster inference for free2025-11-20
  19. Part 19KV-cache engineering — the memory wall of long context2026-01-15
  20. Part 20The RAG-to-agentic-RAG evolution2026-03-19
  21. Part 21Synthetic data and the model-collapse debate2026-05-21
  22. Part 22Constitutional AI and scalable oversight2026-07-16
  23. Part 23Where the frontier goes next2026-09-17
  24. Part 24The prompt is not the product — in-context learning and its limits2023-02-16
  25. Part 25Instruction tuning — from FLAN to Alpaca2023-03-30
  26. Part 26Hallucination — the failure mode that will not go away2023-04-13
  27. Part 27Alignment before the assistant era — RLHF and its discontents2023-05-18
  28. Part 28Diffusion models — how image generation actually works2023-06-15
  29. Part 29The open-weight leap — LLaMA, Alpaca, and the leak that built an ecosystem2023-07-06
  30. Part 30Tool use arrives — function calling and the plugin experiment2023-08-10
  31. Part 31AutoGPT and the first agent wave2023-09-07
  32. Part 32Serving is the product — PagedAttention and continuous batching2023-10-05
  33. Part 33Mechanistic interpretability — looking inside the black box2023-11-02
  34. Part 34Prompt injection — the vulnerability class nobody can patch2023-12-07
  35. Part 35Mistral, Mixtral, and the efficiency-first open ecosystem2024-01-04
  36. Part 36Benchmarks and contamination — measuring models honestly2024-02-08
  37. Part 37The long tail of licensing — what open weights actually permit2024-03-07
  38. Part 38Training infrastructure — FlashAttention, FSDP, and the plumbing of scale2024-04-04
  39. Part 39The GPU supply chain and the cost of a training cluster2024-05-02
  40. Part 40Data curation — deduplication, filtering, and why clean beats big2024-06-06
  41. Part 41Voice interfaces — ASR, TTS, and the latency budget2024-07-11
  42. Part 42Evaluation-driven development — building your own test set2024-08-01
  43. Part 43Embeddings as infrastructure — search, clustering, and drift2024-09-05
  44. Part 44Synthetic preference data and the RLAIF flywheel2024-10-10
  45. Part 45Structured output — schemas, grammars, and constrained decoding2024-11-07
  46. Part 46Small models on device — NPUs, memory budgets, and privacy2024-11-27
  47. Part 47Multi-agent systems — when decomposition helps and when it hurts2024-12-19
  48. Part 48Fine-tuning versus prompting — the decision nobody makes explicitly2025-02-06
  49. Part 49Agent observability — tracing, replay, and debugging loops2025-03-06
  50. Part 50Inference cost engineering — caching, routing, and unit economics2025-04-03
  51. Part 51The evaluation of agents — tasks, traces, and cost-aware scoring2025-04-24
  52. Part 52Long-context retrieval — position, recency, and what gets lost2025-06-05
  53. Part 53Reasoning distillation — putting thinking into small models2025-07-03
  54. Part 54Robotics and embodied models — the sim-to-real gap2025-08-07
  55. Part 55Biology and protein models — what AlphaFold changed2025-09-04
  56. Part 56Science with models — from conjecture to verification2025-10-09
  57. Part 57Security of agentic systems — permissions, sandboxes, and blast radius2025-11-06
  58. Part 58Data residency and sovereignty — where inference is allowed to run2025-12-04
  59. Part 59Model routing — choosing a model per request, and proving it works2026-01-08
  60. Part 60Continuous evaluation in production — drift, regressions, rollbacks2026-02-05
  61. Part 61Retrieval at scale — sharding, freshness, and cache invalidation2026-03-05
  62. Part 62Human-in-the-loop design — review queues and trust calibration2026-04-09
  63. Part 63Interoperability — standards, protocols, and portable context2026-05-07
  64. Part 64The economics of applied AI — where the margin actually is2026-06-04
  65. Part 65A field guide to shipping AI systems — what I would do differently2026-07-02

Subscribe to this series: RSS feed · everything

Other series: React.js · Node.js · AI-Enabled RMA · DotNetCode · DevOps on AWS · DevOps on Azure · DevOps on Google Cloud · AI fundamentals · Stripe in C# · Shopify in C# · Windows domains · A math practice toy · all series