I'm writing this to leave a timestamp I can be graded against. Every probability here is a subjective estimate. At the end I list the specific observations that would make me change my bet.
My bet on the endgame
Let me say up front that this is a question I find genuinely fascinating. But what I'm betting on isn't any single technical approach. It's the architecture that emerges once the division of labor stabilizes — plus a judgment about where the value ends up:
The endgame architecture has three layers. A centrally trained learned prior (a VLA plus a world model) covers the task distribution. Deferred computation at deployment — search, verification and constraint solving inside the learned model — corrects for distribution shift. A hard-constraint layer issues certificates. All three are necessary. The real design problem is the interfaces between them.
The prior becomes a commodity. Three to five vendors, cross-embodiment, priced per token. Value accrues in the deferred-computation layer: the world-model fine-tunes, verifiers, constraint definitions and failure data held by whoever owns the task.
Deployment order is set by who pays for failure. Settings that are constraint-dense, forgiving on reaction time and expensive to fail close the loop first — industry, logistics. Homes come last.
Subjective probabilities:
Hybrid: learned prior + test-time search / verification — ~75%
Pure end-to-end policy, no deferred computation — ~15%
Explicit-model planning at the core, learning only for perception — <5%
Something else — ~5%
One more: an explicit constraint / certificate layer becomes a standard component of industrial robots before 2030 — about 55%. Lower for consumer robots.
The rest of this piece explains why I'm betting this way, where the frontier actually is, and what would make me change the numbers.
A different axis: where the computation happens
The usual axis in the policy-versus-planning debate is whether the problem can be modeled as an MDP. The strongest result on that axis is an existence claim: when the objective depends on history, and there's no way to summarize that history in a fixed handful of numbers (no finite-dimensional sufficient statistic), no state-feedback law exists, and the policy has to be defined over histories. That's correct. But it tells you nothing about efficiency, and nothing about architecture.
A more useful axis: is the computation for a decision amortized into parameters, or deferred and solved at deployment time? Baking a plan into a controller is amortization; planning online is deferral. RL amortizes, MPC defers. A VLA amortizes; test-time search defers.
The payoff of this axis is that both sides' costs can be written down symmetrically:
Amortization's cost doesn't grow with state dimension — it grows with how many task and constraint distributions you have to cover. The query is O(1). Out-of-distribution failures are silent: the policy doesn't know it has left the distribution.
Deferral's cost grows polynomially with horizon and dimension, and exponentially with the branching of uncertainty. Its failures are explicit — infeasible, didn't converge — but that "explicit" is only trustworthy with a convex or complete solver.
Notice that each side has its own exponential: amortization is exponential in coverage, deferral in uncertainty. Any argument that counts only one axis reaches a one-sided conclusion. Count dimensions alone and planning wins outright; count real-time closed-loop behavior alone and policy does.
Why neither layer can be dropped
Write that symmetry as two necessity claims.
What you can't skip: solving online at deployment
Deferred computation is necessary. Any system with hard constraints needs a layer that solves online at deployment and can return a feasibility certificate. A policy's constraint-satisfaction rate is an in-distribution statistic, not a certificate.
More concretely: for constraints that only arrive at deployment — the obstacle set produced by perception, where the fixture happens to sit today, a window length the user specifies — a policy has to cover the entire constraint family in its training distribution, while a planner takes the constraint as an input and pays nothing to cover it. The exponential really shows up in how finely the constraint has to be described, not in the time horizon.[1]
What you also can't skip: closed-loop reaction
Amortization is necessary. Any behavior that has to react in closed loop under uncertainty, faster than a solver can solve, must be amortized into a policy. Open-loop planning — receding horizon included — is structurally suboptimal under partial observability and randomness: it never probes the situation to see it more clearly before committing (control theory calls this dual control), and expanding every possible branch into a scenario tree puts you straight back on an exponential. The core value of a feedback policy is precisely that it amortizes closed-loop behavior into a function.
If both claims hold, the architecture isn't either/or. What's left is the interface: what the policy hands the planner (a warm start, a mode sequence, a terminal cost), and what the planner hands the actuators (a certificate, a refusal, or a corrected action).
Why the third layer stands on its own
Why pull hard constraints out as a separate third layer? Feasibility certificates exist only for convex (or convexifiable) formulations. Non-convex cases need a complete solver, and sampling-based planners are only probabilistically complete — sample long enough and they will "almost surely" find a solution, but when they don't, they can't tell you that none exists. So being correct-by-construction is a property of the formulation, not of "planning."[2] The design rule follows: push hard constraints into a convex safety layer — control barrier functions, reachable sets, corridor constraints — between the policy and the actuators, and leave everything else to soft verification.
Four drivers and a template
Inference compute prices (high confidence). Still falling, and pushing the marginal cost of deferred computation toward zero. This is the underlying force behind every test-time-compute approach, and it isn't going to reverse.
But keep the two kinds of compute separate. What's getting cheaper is inference compute, and it feeds deferred computation. The big money going out the door right now buys training compute, and that feeds amortization. In September 2026, Figure signed with Nscale for up to 100,000 NVIDIA Vera Rubin GPUs on an initial $3.5 billion commitment, and the announcement says plainly that the compute is for training its next generation of models. The commitment is larger than all the funding Figure has publicly disclosed to date — roughly $1.9 to $2.3 billion — and the first of that compute doesn't come online until the second half of 2027. It's a clean statement of intent: the industry is answering the coverage problem with training compute.
If this essay's framework holds, that money can't buy what it's trying to buy. The exponential that binds on the amortization side is in coverage — not in compute, and not in state dimension. Compute can wring more out of the same data. It cannot conjure contact that was never recorded.
The non-homogeneity of robot data (medium confidence). Coverage across embodiments and scenes grows more than an order of magnitude more slowly than it does for language. A pure policy has to flatten the "constraint novelty vs. failure rate" curve with data, at enormous cost. But counter-evidence is appearing — zero-shot cross-embodiment transfer is starting to work — which weakens "there will never be enough data" while strengthening "the prior becomes a commodity": a prior that genuinely works across embodiments only needs a few vendors.
Contact physics (high confidence). Contact dynamics are non-smooth and their parameters aren't identifiable; you can't write the ODE down. That rules out explicit-model planning as the core: the model has to be learned. Planning doesn't disappear — it comes back as search inside a learned model, not as numerical optimization over hand-written equations of motion. The real opposition was never planning versus learning. It's learning a policy versus learning a model and planning in it. The same fact shapes the certificate layer for contact tasks: you can't get a pre-contact planning guarantee, so verification moves to post-contact inspection.
LLMs as the template (medium-high confidence). Pretraining → RL post-training → test-time compute plus verifiers: language models have walked that road, and robotics is replaying it two to three years behind. Robotics has one advantage over language here: physical outcomes are verifiable, so verifiers have ground truth. In text, the verifier itself is the bottleneck. In robotics it isn't.
The first three drivers imply a hybrid architecture. The fourth sets the timetable.
Where the frontier is
Deferred computation is already mainstream. Verifiability is almost untouched.
Deferred computation: already mainstream
τ₀-VLA (August 2026) is the most direct evidence. It's trained on 40,115 hours of real robot data. Across four long-horizon tasks, 10 trials each, the hierarchical version averages 45% success against 27.5% without the hierarchy. World-model-guided search at planning time lifts closed-loop success on a book-organizing task — one that includes initial arrangements never seen in training — from 60% to 90% (6 of 10 trials to 9 of 10); the other two tasks go from 50% to 70%. The largest gain lands on the task with out-of-distribution cases, which is the shape the theory of deferred computation predicts. But with 10 trials per task, a single trial is 10 points: this is consistent in direction, not confirmation. The absolute numbers also say this is still a research artifact, not a deployable stack.
π0.7 (April 2026) shows the progress on the prior side: multi-stage tasks in unseen environments; zero-shot cross-embodiment transfer — a robot that had never seen a laundry-folding task folds shirts with 80% success; and, out of the box, performance on tasks like operating an espresso machine that matches specialized RL-fine-tuned models. At runtime, subtask instructions come from a learned high-level language policy, or from a human coaching it; subgoal images come from a lightweight world model.
World models are already inside the main loop — but as conditional generators, not as verifiers. π0.7 also distills autonomous data from RL-post-trained agents, failures included, into the generalist model. That is the mechanism of prior commoditization: what specialist RL learns gets absorbed by the generalist, and the specialist no longer needs to exist on its own. The same team's multi-scale memory work pushes tasks to as long as fifteen minutes, handling path dependence with architecture rather than with state augmentation — an empirical confirmation of footnote 1.
Cosmos Policy turns amortize-versus-defer into a runtime switch: in direct policy mode it outputs only actions; in planning mode it uses predicted future states and values to rank candidate trajectories. That's one concrete implementation of "the interface is the core design problem."
Verifiability: almost nobody is working on it
The part of the frontier that isn't on this road — and what I think is the real gap:
In every system above, "verification" means a learned value function doing soft ranking. There is no certificate. World-model error is unbounded; search lifts success from 60% to 90%, but it does not push the probability of violating a constraint down to a number anyone can state.
Search happens at the subtask or semantic level. The low-level policy still executes action chunks open-loop, with no online layer enforcing continuous constraints.
Model uncertainty quantification is missing. Without it there can be no certificate layer — a certificate is conditional on the model, and if model error has no bound, the certificate means nothing.
The gap isn't caused by technical difficulty. It's caused by incentives. Frontier labs optimize benchmark success on household tasks, where failure is cheap; the demand for certificates comes from industrial deployment, which isn't in their objective function. So this layer probably won't be filled in by frontier labs. It will be built one deployment task at a time.
I can say that with some confidence because one industry has already walked this road.
A precedent that already played out: defect inspection
What happened in two years
Around 2019, industrial defect inspection in China was still in the stage of educating customers. By around 2021 the paradigm had largely settled — and in the same window, the market turned into a price war. From customers starting to make serious demands to everyone doing it the same way took about two years.
In those two years, customers wanted two things: proof of certainty, and an explanation of how it worked. The supply side's answer was to explain the principles — convolutions, why this patch of pixels was called a defect. I gave that explanation many times, and the more I gave it, the clearer one thing became:
They couldn't follow it — and they shouldn't have had to.
You can't get a quality engineer to sign off on a line's verdicts by teaching them neural networks. What they wanted wasn't to understand the model. It was for something else to catch it when the model was wrong.
What finally settled wasn't a better explanation. It was a structure: the model makes the first call, hard rules backstop it, conventional metrology intercepts, and the parameters stay adjustable on site.
Hard rules are not metrology
One distinction matters here, or this precedent turns into "just add some rules": hard rules and conventional metrology are not the same thing. Hard rules are an engineering compromise — a pile of human-written if-statements, and the bigger the pile, the less anyone can say whether the model or the rules made the call. Metrology is something else. It produces a quantity with units and a stated uncertainty, backed by a calibration chain traceable to a reference standard. It doesn't explain the model. It goes around the model.
That is what a certificate is: its correctness doesn't depend on whether the model is right.
This maps onto the design rule above: feasibility certificates exist only in formulations you can actually model. Metrology is that formulation. Rules are not.
What this precedent can't be used for
I take three things from this history — but first, what it can't be used for. Inspection and contact work are not the same problem. Inspection decisions are low-dimensional and the takt time is fixed. More importantly, the measuring instruments already exist: calipers, vision measuring machines and profilometers are sitting right there, and wiring them into the process is an engineering problem. Contact tasks have no such ruler, so a certificate layer for them first has to answer "measure it with what?" Of the three points below, the first two are about what happened in this industry, and I'm confident in them. The third is extrapolation. Read it as a hypothesis.
First, the certificate layer was built after explanation failed. Customers wanted proof of certainty. The industry tried explaining the principles and got nowhere. What actually solved it wasn't getting people to understand the model — it was putting an independent measurement outside it.
Second, this layer was forced into existence by customers, not invented by researchers. No paper proposed this structure. It's the shape that got negotiated across the table from customers who refused to accept statistical guarantees. That's exactly why I don't expect frontier labs to build the certificate layer — not because they can't, but because nothing in their objective function pays for a missed defect.
Third, ownership of the constraints is migrating. Early on, field engineers tuned the parameters; over the last two years that has gradually passed to customers, who now tune them themselves. It looks like product maturity. What's really happening is that the right to define pass/fail is moving from the party that builds the model to the party that bears the cost of failure. If robotics reaches the same point — where whoever pays starts demanding proof — my guess is that the asset that accrues will be the same thing: not the model, but how the constraints are defined, and by whom. That's bet #2 above. Inspection gives it one case that has already happened. It is only one case.
The same goes for "two years." It was achieved with measuring instruments already in hand, and it doesn't transfer directly. All it supports is this: once the paying side genuinely starts demanding proof, this layer may not take long to build.
It also explains something from earlier: the inspection industry skipped the deferred-computation layer entirely and converged straight to "amortize plus certificate." The reason is prediction P2 below. Takt time in inspection is fixed, and the required reaction time is far shorter than any online solve — which puts inspection on the left side of the crossover, where online search can't deliver.
Falsifiable predictions
These four can be tested directly on systems that exist today:
P1 — Constraint-novelty curve. Hold in-distribution success fixed, and put on the x-axis how far the obstacle or fixture geometry sits from the nearest example in the training set (Hausdorff distance). A pure policy's failure rate rises monotonically with that distance; a hybrid's stays roughly flat, with a ceiling set by warm-start quality × solver completeness.
P2 — Latency crossover. Put "required reaction time / solve time" on the x-axis and there is a crossover: below 1, hybrid ≈ policy; above 1, hybrid ≫ policy. Where the crossover sits is measurable.
P3 — Silent failure rate. A pure policy's constraint violations go unflagged; the fraction of a hybrid's violations that get flagged ≈ solver completeness. That is the entire safety argument, and it can be counted directly.
P4 — Parameter extrapolation. On path-dependent tasks with a dynamically specified window length, a policy with sequence memory handles windows inside its training range and fails outside it; a planner holds for any parameter. Push the window parameter outside the training range and you have the test.
What would change my bet
Search gains that shrink with data. If τ₀-style planning-time search gains go to zero after a tenfold increase in data, coverage can ultimately be solved by amortization, and the 15% on pure policy has to go up. This is the most important indicator; the ratio worth tracking is search gain over data scale.
Zero-shot success above 95% without search. If the next generation of generalist models crosses that line with no test-time search, same conclusion.
A market that tolerates accidents. If the certificate layer keeps failing to show up in industrial deployments, and failures simply get absorbed as an operating cost, then I've overestimated the demand for verifiability, and the 55% comes down.
All three can be tracked today. I'll be back in a year to grade this.
Where I concede
The prior still has to be learned, and the data hunger is real. A hybrid architecture reduces the coverage required; it doesn't eliminate it. I can't put a bound on by how much.
On contact and hybrid dynamics, planner completeness is weak, and the combinatorial explosion of mode sequences has no good solution. That part of the load still has to be carried mostly by learned mode proposals.
Certificates are conditional on the model. Model uncertainty quantification isn't optional — it's the premise the whole argument rests on. And it happens to be exactly what the frontier lacks most.
If you know someone about to sign off on a robot cell because it passed 99.7% of its trials — send them this before they sign.
Notes
[1] This corrects claims of the form "representation complexity ~ d^N." d^N is the cardinality of the history space — the complexity of a lookup table, not the parameter count of a structured function class. A time delay is a shift register, a sliding window is a bounded buffer, a spectrum is a filter bank; RNNs and attention implement these exactly with very few parameters. The counterexample is language modeling: unbounded history, heavily path-dependent, solved by transformers with a polynomial number of parameters. What grows exponentially is the family of constraints you have to cover, not the length of the history.
[2] A local NLP solver reporting "infeasible" only means it didn't find a solution — not that none exists. Obstacle avoidance is exactly a non-convex feasibility problem, and contact is a hybrid, combinatorial one. "Return no solution if it's physically infeasible" comes with a discount in precisely these settings.
References
τ₀-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation, arXiv:2608.16885 (Aug 17, 2026) — https://arxiv.org/abs/2608.16885
π₀.₇: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities, Physical Intelligence, arXiv:2604.15483 (Apr 16, 2026; v2 Apr 24) — https://arxiv.org/abs/2604.15483
MEM: Multi-Scale Embodied Memory for Vision Language Action Models, arXiv:2603.03596 (Mar 4, 2026) — https://arxiv.org/abs/2603.03596
Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning, NVIDIA, arXiv:2601.16163 — https://arxiv.org/abs/2601.16163
World Model for Robot Learning: A Comprehensive Survey, arXiv:2605.00080 (Apr 30, 2026) — https://arxiv.org/abs/2605.00080
Figure and Nscale Sign Strategic Partnership For Up to 100,000 GPUs on the NVIDIA Vera Rubin Platform, Figure (Sep 3, 2026) — https://www.figure.ai/news/figure-and-nscale-sign-strategic-partnership ; Nscale press release — https://www.nscale.com/press-releases/nscale-and-figure
Figure Exceeds $1B in Series C Funding at $39B Post-Money Valuation, Figure (Sep 2025) — https://www.figure.ai/news/series-c





