AI Frontiers, part 10: Multimodal turns one β vision, audio, and the GPT-4o moment
Part 10from the AI Frontiers series · 65 parts in all
On 13 May 2024 OpenAI announced GPT-4o, and the demo everyone remembers is the one where a model looks through a phone camera, reads a handwritten equation on a sheet of paper, and coaches the person holding it β all in one voice conversation with roughly the latency of a human reply. What made that demo land was not that a model could describe an image. GPT-4 could do that in March 2023, and GPT-4V shipped to developers on DevDay that November. What was new was the shape of the pipeline: one model taking text, vision and audio in, and producing text, vision and audio out, without the three-stage relay of speech-to-text β language model β text-to-speech that every voice assistant had used since 2011 (OpenAI, "Hello GPT-4o").
It is worth being precise about what a year of multimodality actually bought, because the hype and the engineering tell different stories. The engineering story is about how vision got bolted onto language models in the first place β with contrastive encoders and learned adapters β and why that architecture has a natural ceiling. The hype story is about assistants that see. This entry sits between the two, and it connects back to part 1: everything here is the attention mechanism applied to a longer, messier token stream.
How a language model learned to see: encoders and adapters
The foundation was CLIP. Radford and colleagues at OpenAI trained an image encoder and a text encoder with a contrastive objective over 400 million imageβcaption pairs: push the embedding of an image and its caption together, push everything else apart. The result was a model that could classify ImageNet with zero task-specific training, and, more importantly for us, a shared embedding space where "a photo of a dog" is a point you can compute against an image (Radford et al., "Learning Transferable Visual Models"). CLIP was not a captioning model; it was a matcher. But a matcher is exactly the piece a language model needs, because the language model already knows how to continue text given a prefix.
The trick that turned a matcher into a conversational partner was the adapter. Flamingo, from DeepMind, froze both a vision encoder and a large language model and trained only the cross-attention layers that let the language model attend to visual features. It handled interleaved sequences of images and text, so a prompt could contain four images and a question about them (Alayrac et al.). BLIP-2 compressed visual features into a small set of query tokens with a Q-Former before handing them to a frozen LLM (Li et al.). And LLaVA took the most economical version of the idea: a single learned linear projection from CLIP features into the language model's embedding space, trained on 595,000 visual instruction examples that were themselves generated by GPT-4 from captions (Liu et al., "Visual Instruction Tuning").
That last detail is the one I keep coming back to. The instruction data for vision was synthesized by a language model working from image captions. The vision capability came from CLIP; the conversational capability came from distilling GPT-4's behavior over descriptions of images. Fine-tuning only the projector and the LLM β LLaVA-1.5 (Liu et al., "Improved Baselines") showed you could get surprisingly far with a plain MLP projector and more data β meant the recipe was accessible to anyone with a few thousand dollars of GPU time. By 2024 there was a small zoo of open vision-language models: Idefics2, MM1 from Apple (McKinzie et al.), Qwen-VL, InternVL. The architecture had converged on "pretrained vision encoder + adapter + the language model you already had."
Why audio is the harder modality
Vision is a static, spatial problem with a heavily pretrained encoder available off the shelf. Audio is a temporal problem with expectations β and those expectations are ruthless. Humans tolerate a caption that is slightly wrong; we do not tolerate a half-second of silence in a conversation before the other party answers, and we absolutely do not tolerate a response that arrives before we have stopped speaking.
The pre-2024 industry consensus was a cascade: Whisper transcribes, the LLM reasons, a TTS engine speaks. Whisper (Radford et al., "Robust Speech Recognition") was excellent at its job β 680,000 hours of weakly supervised multilingual audio, robust to accents and noise β and the cascade's virtue was that each stage was debuggable and independently swappable. Its vice was latency and information loss. Prosody, tone, laughter and hesitation were stripped by the transcript and could never come back. When you asked a cascaded assistant a question and it heard a sigh, it heard the word "okay".
End-to-end speech models had been explored β Google's AudioPaLM merged a speech encoder with a language model and could do speech-to-speech translation while preserving the speaker's voice (Rubenstein et al.); Meta's SeamlessM4T (Communication et al.) pursued similar ground for translation β but they were research systems. GPT-4o's contribution was to make the end-to-end version a product with conversational latency, and to train one model on text, audio and images together so that the audio branch inherited knowledge learned from text. That inheritance matters: a speech model trained only on audio data has seen a tiny fraction of the world's concepts compared with a model that read the web.
Google had made the same architectural bet earlier and more explicitly. Gemini 1.0 was presented from the outset as natively multimodal β trained on interleaved text, images, audio and video rather than assembled from a language model plus adapters (Gemini Team, "Gemini: A Family of Highly Capable Multimodal Models"). The interesting thing about the Gemini 1.5 report from this spring is not the benchmark table; it is the demonstration that a sufficiently long context window turns video from a model architecture problem into a token-count problem. Feed the frames in as a sequence. That is the context window argument again, wearing a camera.
Benchmarks, and the gap between seeing and reasoning
Multimodal evaluations in 2023β2024 were, charitably, uneven. MM-Vet (Yu et al.) tested composite abilities β recognition plus spatial reasoning plus math β and GPT-4V scored in the high sixties. MMMU, a college-level exam drawn from thirty subjects, put GPT-4V around 56% against roughly 89% for human experts (Yue et al.). MathVista showed a similar pattern: strong at the perception step, weaker once the question required chaining. The pattern is consistent enough to state as a rule of thumb β vision models of this generation read images well and reason about them less well, and the failure mode is a confident, fluent, wrong answer rather than a refusal.
That is why I take the "Dawn of LMMs" analysis (Zhang et al.) seriously as a design document rather than a leaderboard. It catalogued hundreds of image-related capabilities and found that model behavior was extremely sensitive to prompt phrasing and to the position of the image in the prompt β small changes produced large output changes with no change in model weights. Anyone who shipped a vision feature in 2023 learned this the hard way: the differences between a working document-extraction pipeline and a broken one were often literally the order of the image and the instruction.
What actually shipped, and what it cost
Away from demos, the production wins were unglamorous. Accessibility was the first one: Be My Eyes began routing requests from blind and low-vision users to GPT-4V in 2023, and "describe this" is a task that vision models do genuinely well. Document understanding was the second: receipts, invoices, forms and packing slips, where the model combines OCR-grade reading with schema-aware extraction and replaces a fragile template-based pipeline. The third was quality control on images that previously needed a human β a photographed label, a damaged part, a screenshot of a broken UI. And the fourth was video, arriving slowly: not understanding a film, but answering "did anything cross the line in this thirty seconds?"
The cost structure is the part people underestimate. Images are not free at the token level: GPT-4V tokenizes an image into a base cost plus a per-tile cost, so a high-resolution page can consume thousands of tokens before the question is even asked. Multiply by a document workflow and the arithmetic gets uncomfortable fast. This is the second reason small specialized models matter so much in multimodal pipelines: the frontier model handles the hard twenty percent of documents, and a compact vision model β or a classical CV pipeline β handles the easy eighty at a fraction of the cost.
Building a vision pipeline that survives contact with users
If you were wiring this up in 2024 rather than reading about it, the decisions you actually made were boring ones, and they were the ones that determined whether the feature shipped. Four of them recurred in every project I touched.
Resolution before model. A vision model's accuracy on a photographed receipt collapses if you downscale the image to save tokens, and the fix is almost never a better model β it is cropping, deskewing and sending three sharp tiles instead of one mushy page. The tile accounting is public: a base cost per image plus a per-tile cost, so the difference between a 512-pixel and a 1500-pixel image is measurable in dollars per thousand documents. Budget the tokens the way you would budget pixels.
Decide what needs a language model at all. A large fraction of "vision" tasks are retrieval problems. If the question is "which of our 40,000 catalog images looks like this one?", a CLIP-style embedding plus a vector index answers it in milliseconds at a thousandth of the cost of a generative call β and the machinery is identical to what part 6 described for text. Reserve the VLM for the cases where you need a sentence, a decision or a structured record.
Make the model emit schemas, not prose. The single highest-leverage change in every document pipeline was switching from "extract the invoice details" to "return JSON matching this schema" and then validating. Fluency is the enemy of extraction: a model that writes a beautiful paragraph is harder to audit than one that either fills six fields correctly or fails validation. Failures then become ordinary error-handling problems instead of mysteries.
Keep a human in the loop where the cost of being wrong is asymmetric. Confidence in multimodal output is not calibrated, and the observed failure mode is a confident wrong answer about a detail nobody thought to check. The pipelines that held up in production put a low-cost review step β a diff, a highlighted region, a thumbs-down button that logs the image β on the fraction of outputs where a wrong answer would matter. That review queue is not a temporary concession to immature models; it is where your evaluation data comes from.
None of this is specific to GPT-4o, and that is the point. The interfaces stabilized faster than the capabilities did. By mid-2024 you could write against one schema for images and swap the model underneath β frontier, open-weight or a specialized OCR system β and the engineering decision that mattered most was how much of your pipeline assumed the model was going to be perfect.
The interface is the message
Here is my honest read, a year in. The interesting part of multimodality is not that models can see. It is that seeing changes what a model is for. A text-only model is something you go to, deliberately, with a question you have already formulated. A model that can see, hear and speak is something that can be present while you do something else β and that is a different product category, with different latency budgets, different failure tolerances and different privacy consequences.
Which is why the most consequential line in the GPT-4o announcement was probably the one about latency, not the one about modalities. Every previous assistant failed at the turn-taking level before it failed at the intelligence level. If the turn-taking problem is solved, the next set of problems becomes visible: what a model does with a persistent camera, how a conversation that is never written down is governed, and who is accountable when a model that is fluent and confident is simply wrong about what it saw. We spent the last year teaching models to perceive. The next year is about deciding what they should be allowed to look at.
Works Cited
Alayrac, Jean-Baptiste, et al. "Flamingo: a Visual Language Model for Few-Shot Learning." arXiv, 2022, arxiv.org/abs/2204.14198. Accessed 18 July 2024.
Communication, Seamless, et al. "SeamlessM4T: Massively Multilingual & Multimodal Machine Translation." arXiv, 2023, arxiv.org/abs/2308.11596. Accessed 18 July 2024.
Gemini Team, Google. "Gemini: A Family of Highly Capable Multimodal Models." arXiv, 2023, arxiv.org/abs/2312.11805. Accessed 18 July 2024.
Li, Junnan, et al. "BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models." arXiv, 2023, arxiv.org/abs/2301.12597. Accessed 18 July 2024.
Liu, Haotian, et al. "Improved Baselines with Visual Instruction Tuning." arXiv, 2023, arxiv.org/abs/2310.03744. Accessed 18 July 2024.
---. "Visual Instruction Tuning." arXiv, 2023, arxiv.org/abs/2304.08485. Accessed 18 July 2024.
McKinzie, Brandon, et al. "MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training." arXiv, 2024, arxiv.org/abs/2403.09611. Accessed 18 July 2024.
OpenAI. "Hello GPT-4o." OpenAI, 13 May 2024, openai.com/index/hello-gpt-4o/. Accessed 18 July 2024.
Radford, Alec, et al. "Learning Transferable Visual Models From Natural Language Supervision." Proceedings of the 38th International Conference on Machine Learning, 2021, arxiv.org/abs/2103.00020. Accessed 18 July 2024.
---. "Robust Speech Recognition via Large-Scale Weak Supervision." arXiv, 2022, arxiv.org/abs/2212.04356. Accessed 18 July 2024.
Rubenstein, Paul K., et al. "AudioPaLM: A Large Language Model That Can Speak and Listen." arXiv, 2023, arxiv.org/abs/2306.12925. Accessed 18 July 2024.
Yu, Weihao, et al. "MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities." arXiv, 2023, arxiv.org/abs/2308.02490. Accessed 18 July 2024.
Yue, Xiang, et al. "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI." arXiv, 2023, arxiv.org/abs/2311.16502. Accessed 18 July 2024.
Zhang, Xinlu, et al. "The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)." arXiv, 2023, arxiv.org/abs/2309.17421. Accessed 18 July 2024.