Robotics Analysis
The Gap Between a Robot Demo and a Dependable Deployment
Physical AI is gaining capabilities. The harder commercial question is how much supervision remains between a successful attempt and a reliable operation.
Executive summary
A robot that completes a difficult task once has demonstrated a capability. A deployment needs to complete enough useful work, with sufficiently little supervision, to justify its cost. The distance between those two statements is where many of the important engineering and commercial questions sit.
The arrival of more general robotics models makes that distance worth examining carefully. It broadens what can be attempted, while leaving operators responsible for the conditions under which the attempts are useful.
The unfinished attempt belongs in the result
DeepMind's July 2026 release reports whole-body control and multi-robot coordination, while explicitly acknowledging remaining difficulties with multi-finger manipulation. Earlier Gemini Robotics research also examines adapting models to new robot embodiments and tasks. These are capability results, rather than evidence that every relevant workflow is ready for unattended operation. 1 2
Consider a robot unloading a mixed container. A trial that ends after a successful grasp omits the awkward package left behind, the object dropped outside its expected location and the recovery that follows. In a working facility those events consume time. Somebody must decide what happens next.
The denominator should therefore include all attempts and the time spent recovering from them. Otherwise the evaluation rewards a narrow success while externalising its operating cost.
Reliability compounds across a sequence
Suppose a workflow contains ten steps, each with a 98% chance of success. Under the simplifying assumption that failures are independent, the probability of completing all ten is approximately 82%: 0.98^10. This is an illustrative calculation, not a reported robot result.
Real failures can be correlated. A poor initial grasp can make subsequent steps harder, while a successful recovery can rescue the workflow. The point of the calculation is to show why a strong individual skill score cannot be read directly as end-to-end reliability.
For a purchasing decision, the useful measurements include completed jobs per hour, operator minutes per completed job, damage or rejection rates, and performance after a change in objects or surroundings. Record these across the operating window, not just during the best sequence.
Generality has a price and a benefit
A specialised cell can justify substantial setup work when the task is repeated enough. A more general robot may create value through faster changeovers or a wider range of useful tasks, even if it is slower at one individual operation.
That comparison depends on the business. A stable, high-volume line and a workshop with frequent product changes have different needs. Asking which robot is more intelligent does not settle which one is more useful in either setting.
Our assessment is that deployment proposals should price both adaptation and supervision. How long does a new task take to configure? What can an operator change without a model specialist? How many robots can one person supervise when several encounter exceptions at once?
A pilot should expose the operating boundary
Agree on an acceptance envelope before running the pilot: the objects, environmental variation, production rate and acceptable intervention burden. Include the cases the operator finds inconvenient, because these are often the cases that determine whether the system is worth running.
The final review should explain where performance holds and where it breaks. A restricted deployment with a well-understood boundary can be valuable. A broad promise with an undisclosed supervision requirement is difficult to budget or scale.
The strongest evidence for physical AI will be a wider range of economically useful work, with the support effort visible in the result. That is the evidence to ask for as demonstrations become more impressive.
The evidence behind the analysis