AI Frontiers, part 54: Robotics and embodied models — the sim-to-real gap
Part 54from the AI Frontiers series · 65 parts in all
Everything in this series so far has been about systems whose mistakes cost money or credibility. Robotics is where mistakes cost fingers and hardware. That is why the field is worth a full entry even though almost nobody reading this will deploy a manipulator this year: the constraints are so much tighter that the techniques which survive there are informative about the ones that will survive everywhere else.
Between 2023 and 2025 the field went through the same transformation language did, one cycle later. Vision-language-action models — one network that takes camera images, a task instruction, and proprioceptive state, and emits motor commands — replaced the hand-tuned pipeline of detection, pose estimation, motion planning, and control that had dominated manipulation for two decades. RT-2 demonstrated that a language model's world knowledge could be transferred to low-level control, which was the surprise that launched the subfield (Brohan et al., "RT-2"). OpenVLA and its successors made the approach reproducible on open data and open weights (Kim et al.), and generalist policies trained across many robot embodiments began to appear. By 2025, general-purpose robot foundation models were the announced strategy of every large lab with a hardware program.
And yet the deployed fleet of useful general-purpose robots remains small. The reason is not that the models are bad at the demos. It is that a demo is a demonstration of capability on a distribution the team selected, and deployment is a question about the tails.
Why embodiment makes everything harder
Five properties distinguish embodied systems from the software systems the rest of this series describes.
Errors are physical and often irreversible. A wrong token is a retry. A wrong trajectory is a broken gripper, a dropped object, or an injured person. That converts reliability requirements from "good enough to ship with monitoring" into "demonstrably safe under fault conditions," a standard that no amount of benchmark performance addresses.
Data collection is expensive and slow. Language data is scraped from the internet, where humans deposited decades of text for free. Robot data has to be created by teleoperating a physical machine in real time, one trajectory at a time, in an environment someone built. The Open X-Embodiment collaboration pooled more than a million trajectories from many labs across many robots (O'Neill et al.) — an enormous community effort that is, in token terms, a small corpus.
The distribution shift is continuous and physical. Lighting changes across an afternoon. Friction changes with humidity. A new box is a millimeter thinner. Vision-language models tolerate aesthetic drift; a policy that tolerates a two-millimeter drift in the position of a grasp point drops the object.
Long-horizon tasks compound. Every argument in part 51 about compounding errors applies with a physical multiplier: a task with thirty steps, each ninety-nine percent reliable, succeeds about seventy-four percent of the time, and the failures are the expensive kind.
The feedback loop is slower than the improvement cycle. In software, you can ship a fix in an hour. In hardware, a fleet learns at the speed at which you can collect and label new data and validate it against safety requirements, which is measured in quarters.
The lineage, compressed
End-to-end visuomotor learning — mapping pixels directly to joint commands — was shown to be trainable with a real robot in the loop more than a decade ago, at the cost of enormous data collection effort (Levine et al.). The same paper's key insight survived every later reinvention: the hard part is not the policy architecture but the ability to collect enough data in the right distribution, and the practical trick is to make collection cheap and parallel.
Then came the transformer era. RT-1 established a scalable, high-capacity policy trained on a large multi-task robot dataset (Brohan et al., "RT-1"). RT-2 co-fine-tuned a vision-language model on robot trajectories and internet data, and produced the result that mattered for everything after: emergent semantic generalization, where the policy correctly handled objects and instructions it had never seen in a robot context, because it inherited the language model's representation of them (Brohan et al., "RT-2"). OpenVLA made the recipe open and reproducible on Open X-Embodiment data with a 7B backbone (Kim et al.), and continuous-action formulations such as π0 replaced token-by-token action decoding with flow matching for smoother, faster control (Black et al.).
Underneath all of this, an older line of work supplied the data-efficiency trick that makes fine-tuning on a new task tractable at all. Action Chunking with Transformers showed that predicting a short horizon of actions at once, rather than one action at each step, dramatically improved both stability and success rate on fine manipulation with a low-cost bimanual platform (Zhao et al.), and diffusion policies made multimodal action distributions — crucial when there are several ways to grasp an object — tractable (Chi et al.). Most production manipulation stacks in 2025 combine a large pretrained backbone with one of these action heads, and the action head is usually where the engineering win is.
What sim-to-real actually requires
Simulation is the only way to get unlimited practice, and the reason it does not simply solve the problem is that simulated physics is wrong in ways that matter. The literature on closing that gap is older than the current wave and still the right place to start.
Randomize the differences instead of modeling them. Domain randomization trains a policy across a distribution of visual and dynamic parameters wide enough that reality falls inside it (Tobin et al.; Peng et al.). It works, it is cheap, and its cost is that the policy becomes somewhat conservative, because it is optimizing for robustness across a range rather than peak performance at a point.
Make simulation fast enough to be a data source. GPU-parallel simulation — Isaac Gym and its successors (Makoviychuk et al.; Mittal et al.) — turned simulator wall-clock into something closer to a training-loop resource than an experiment budget. Combined with randomization, this produced the familiar recipe: train a robust policy in simulation, then adapt with a small amount of real data.
Use real data to correct, not to learn from scratch. Residual learning and sim-to-real fine-tuning with a few dozen real demonstrations are usually more efficient than either sim-only or real-only. The hardware trend helps here: low-cost arms with good proprioception made it practical for a lab to gather hundreds of real trajectories in a day (Chi et al.).
The honest summary is that sim-to-real is not a transfer problem with a solution; it is a distribution-matching problem with a budget, and every deployment reopens it. The teams that ship are the ones that treat the physical world as the only evaluation environment that counts while using simulation purely to reduce the number of real-world failures needed to get there.
Where foundation models help, and where they do not
The most useful way to divide the problem is by frequency. Language and vision models are excellent at the slow, semantic layers: what is in the scene, which object was referred to, what the task decomposes into, which of five skills applies. They are poor at the fast layers: closed-loop control at hundreds of hertz, contact dynamics, force regulation. Every architecture that works reflects this split. A high-level policy selects skills or produces waypoints; a low-level controller, often classical, executes them. Even the end-to-end models are, in practice, producing action chunks at ten to fifty hertz that a conventional controller tracks.
This division has a direct consequence for how to spend engineering effort. Semantic robustness — grounding a novel instruction, recognizing a new object, recovering from an ambiguous request — is where a bigger model buys measurable progress, because it is a representation problem. Motor reliability is where simulation, calibration, mechanical design, and control buy progress, and no amount of model scale substitutes for a gripper that does not slip.
What a deployment actually looks like
Three things characterize the robots that are doing useful work in 2025: they are constrained, measured, and instrumented.
Constrained: fixed cells, known parts, defined task families, human supervision with an e-stop. The advertised generality of generalist policies is most valuable as fast task onboarding — teaching a new bin or a new part in an afternoon rather than a quarter — not as open-world autonomy.
Measured: success rate per task per shift, intervention rate, cycle time, and the distribution of failure categories. The reliability arithmetic of part 41 applies: once a task needs many sequential steps, only near-perfect per-step reliability is tolerable, and the way to get there is to remove steps and add checks, not to improve the model's average score.
Instrumented: every intervention recorded with camera, state, and action, forming the dataset that makes the next version better. That is the same trace-to-training-data loop as part 49, except the traces are physical and the labeling is a human explaining what the robot did wrong. Fleets that capture this systematically compound an advantage that competitors cannot buy.
And the compute constraints look a lot like the on-device constraints of part 46: power budgets, thermal limits, real-time deadlines, and a hard requirement that the system fail safely rather than slowly.
Why this entry belongs in the series
I include robotics not because I expect readers to build manipulators, but because the field is an early warning system for everything else. Latency budgets, distribution shift, data scarcity, compounding error, safety envelopes, and the impossibility of measuring the tail without deploying — these are the constraints that will bind software agents next. When a robot team tells you they spend more time on evaluation environments and data collection than on models, they are describing the future of applied AI in general. The frontier models democratized the intelligence; the durable advantage is in the environment, the data, and the instrumentation, and nowhere is that clearer than where the feedback is physics.
How to evaluate a robot, and why the answer is not a benchmark
Language models get benchmarks because text is cheap to copy and answers are cheap to check. Robots get something harder, and the difficulty is worth stating because it is the same problem that applied AI in general is heading toward.
Success is task-defined, not answer-defined. "Did the robot place the part in the fixture" has a physical ground truth, but the interesting measurements are around it: placement accuracy, cycle time, how many attempts, how much force, what the operator did when it went wrong. A success rate without those numbers cannot distinguish a policy that works from one that succeeds expensively and dangerously.
Distribution matters more than average. A policy at ninety percent success on the training distribution and forty percent on a new bin of parts is not a ninety percent policy. The evaluation has to be stratified by the axes the physical world varies along: lighting, object pose, material, clutter, and human presence. Building that variation deliberately — rather than waiting for production to supply it — is the robotics equivalent of constructing a held-out evaluation set, and it is the same discipline described in part 42.
Simulation is a screening tool, not evidence. Simulation benchmarks are cheap, reproducible, and useful for comparing architectures and for catching regressions during development. They are not evidence of deployment readiness, because the gap between simulated and real performance is exactly the quantity under investigation. The honest use of a simulator is to run a thousand variations and then test the survivors physically, which is the same relationship a fast statistical screen has to a randomized trial.
Failures need categories, not counts. Dropped object, missed grasp, collision, timeout, operator intervention, incorrect object. The distribution across those categories is the actionable output, because each maps to a different fix — a gripper change, a calibration routine, a perception upgrade, a policy change — and because the categories that matter for safety are a subset of the categories that matter for throughput.
Safety is a separate argument. A reliability measurement is not a safety case. Demonstrating that a system is safe under fault conditions requires a hazard analysis, an argument about what happens when each component fails, and probably a mechanical or supervisory layer that bounds the consequences regardless of what the policy decides. Nothing in this entry about model quality substitutes for that, and the teams that treat the two as the same problem are the ones that end up with both.
Works Cited
Black, Kevin, et al. "π0: A Vision-Language-Action Flow Model for General Robot Control." arXiv, 2024, arxiv.org/abs/2410.24164. Accessed 7 Aug. 2025.
Brohan, Anthony, et al. "RT-1: Robotics Transformer for Real-World Control at Scale." arXiv, 2022, arxiv.org/abs/2212.06817. Accessed 7 Aug. 2025.
Brohan, Anthony, et al. "RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control." arXiv, 2023, arxiv.org/abs/2307.15818. Accessed 7 Aug. 2025.
Chi, Cheng, et al. "Diffusion Policy: Visuomotor Policy Learning via Action Diffusion." arXiv, 2023, arxiv.org/abs/2303.04137. Accessed 7 Aug. 2025.
Kim, Moo Jin, et al. "OpenVLA: An Open-Source Vision-Language-Action Model." arXiv, 2024, arxiv.org/abs/2406.09246. Accessed 7 Aug. 2025.
Levine, Sergey, et al. "End-to-End Training of Deep Visuomotor Policies." arXiv, 2015, arxiv.org/abs/1504.00702. Accessed 7 Aug. 2025.
Makoviychuk, Viktor, et al. "Isaac Gym: High Performance GPU-Based Physics Simulation for Robot Learning." arXiv, 2021, arxiv.org/abs/2108.10470. Accessed 7 Aug. 2025.
Mittal, Mayank, et al. "Orbit: A Unified Simulation Framework for Interactive Robot Learning Environments." arXiv, 2023, arxiv.org/abs/2301.04195. Accessed 7 Aug. 2025.
O'Neill, Abby, et al. "Open X-Embodiment: Robotic Learning Datasets and RT-X Models." arXiv, 2023, arxiv.org/abs/2310.08864. Accessed 7 Aug. 2025.
Peng, Xue Bin, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel. "Sim-to-Real Transfer of Robotic Control with Dynamics Randomization." arXiv, 2017, arxiv.org/abs/1710.06537. Accessed 7 Aug. 2025.
Tobin, Josh, et al. "Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World." arXiv, 2017, arxiv.org/abs/1703.06907. Accessed 7 Aug. 2025.
Zhao, Tony Z., et al. "Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware." arXiv, 2023, arxiv.org/abs/2304.13705. Accessed 7 Aug. 2025.