AI, Applied Benchmarks | Physical AI Wave 1
Physical AI: The Operating Reality
What 85 buyers of robotics and physical AI say about their implementations
Findings from Wave 1 of the Georgian + NewtonX Physical AI Survey, fielded 24 July – 21 August 2026 (n=85 buyers).
In July, we argued that physical AI can be understood as a systems market rather than a model market: that the hardware, data, evaluation, safety and deployment operations around the model matter as much as the model itself.
Since then, we have surveyed 85 robotics buyers ("Buyers") on the operating reality of their implementations. They are plant directors, heads of automation engineering, VPs of operations and COOs, all with decision-making authority and all already running robots. The full sample is described at the end.
Most of what the Buyers report supports our argument that physical AI is a systems market. One finding, however, challenges it.
Finding 1
Buyers say integration is the hardest part of getting physical AI to work reliably
We gave Buyers eight options and asked them to rank the three hardest. Fitting the robot into the way the plant already runs came first: 69% put it in their top three and 32% called it the single hardest thing, the highest first-place share of any option. Operating it live came second, at 55%. Taken together, 84% named workflow integration, live operations or both, and 41% named both. The AI model tied with safety and security for fifth, at 32%.
69% put it in their top three and 32% call it the single hardest thing of all.
Share of buyers ranking each of eight options among the three hardest parts of getting physical AI to work reliably. Integrating into the real workflow leads at 69%, followed by operating it live at 55%. 32% call integration the single hardest thing, the highest first-place share of any option.
Ranked in the top three
n=85 buyers
| Category | Share |
|---|---|
| Integrating it into the real workflow / environment | 69% |
| Operating it live | 55% |
| Getting enough good data and being able to simulate / test | 45% |
| Evaluation | 34% |
| Safety and security | 32% |
| The underlying AI model / autonomy | 32% |
| The robot hardware itself / on-device compute | 21% |
| Still evaluating or not applicable | 12% |
The single hardest thing of all
First-place share, n=85 buyers
| Indicator | Gap / value | Detail |
|---|---|---|
| Hardest of all | 32% | call integration the single hardest thing, the highest first-place share of any option (27 of 85 buyers) |
| Integration or operations | 84% | named workflow integration, live operations or both; 41% named both |
Asked about barriers to scaling, Buyers are more divided. Integration is still the barrier named most often, by 51%, but cost is the one Buyers most often put first.
More buyers name integration overall. When they pick just one, cost comes out ahead.
Share of buyers ranking each of twelve barriers to scaling in their top three. Integration with existing operations leads at 51%, ahead of upfront cost at 44%. When buyers name a single biggest barrier the order reverses: 20% pick upfront cost against 16% for integration.
Ranked in the top three
n=85 buyers
| Category | Share |
|---|---|
| Integration with existing operations and systems | 51% |
| Upfront cost or capital required | 44% |
| Reliability and performance in real conditions | 39% |
| Security and data concerns | 28% |
| Difficulty changing existing workflows | 24% |
| Unclear or unproven ROI | 22% |
| Workforce concerns and change management | 21% |
| Availability of high-quality training data | 20% |
| Supply chain or manufacturing capacity | 19% |
| Safety and risk to people | 14% |
| Lack of internal expertise | 13% |
| Vendor dependence or lock-in | 6% |
When buyers pick just one
Single biggest barrier, n=85 buyers
| Indicator | Gap / value | Detail |
|---|---|---|
| Upfront cost | 20% | name upfront cost or capital as their single biggest barrier (17 of 85 buyers) |
| Integration | 16% | name integration as their single biggest barrier (14 of 85 buyers) |
Finding 2
Few Buyers can measure whether their systems are improving
Asked whether they can tell if a change improved real-world performance, only 20% of Buyers say their measurement is reliable enough to trust. For how the system handles edge cases it was not built for, the figure is 19%. About half say they measure each of these, but only roughly.
The metrics Buyers track consistently are about keeping the plant running rather than autonomy. Intervention rate, for instance, is tracked by only 42%, and that barely shifts between piloting and scaling.
Share of buyers who say they measure each of four things well enough to rely on, topping out at 33% for whether a deployment is succeeding and falling to 19% for edge cases. Alongside, the metrics buyers track consistently: uptime 69% and safety incidents 65% against intervention rate 42% and recovery success rate 26%.
We measure this well and rely on it
n=85 buyers
| Category | Share |
|---|---|
| Whether a deployment is succeeding or failing | 33% |
| Whether performance is degrading over time | 27% |
| Whether a change improved real-world performance | 20% |
| How the system handles edge cases | 19% |
Metrics tracked consistently today
n=85 buyers · the two autonomy metrics accented
| Category | Share |
|---|---|
| Uptime | 69% |
| Safety incidents | 65% |
| Intervention rate | 42% |
| Recovery success rate | 26% |
Finding 3
Intervention is falling, but the reasons are complicated
76% of Buyers told us human intervention (how often a person has to stop what they're doing and go help the robot) has fallen over the past year. Asked what contributed, 88% of them said they had narrowed the range of tasks or conditions the system attempts. Other contributors were routing harder cases to people before the system tries them (75%), slowing the system down for reliability (71%) and adding supervision (63%).
Share of the 65 buyers reporting falling intervention who credit each contributor: narrowing the range of tasks or conditions 88%, routing harder cases to people first 75%, slowing things down for reliability 71%, adding more human supervision 63%.
What contributed to the drop
Base: the 65 of 85 buyers who reported intervention falling
| Category | Share |
|---|---|
| Narrowed the range of tasks or conditions | 88% |
| Routed harder cases to people first | 75% |
| Slowed things down for reliability | 71% |
| Added more human supervision | 63% |
What buyers believe about the drop
69%: all 85 buyers · 89%: the 65 who reported a drop
| Indicator | Gap / value | Detail |
|---|---|---|
| Demos are not readiness | 69% | say demonstrations are not a reliable indicator of production readiness |
| Trust their own drop | 89% | of those reporting a drop believe it reflects a genuine gain in autonomy, not a narrower job |
Each of these changes can lower intervention without necessarily changing the robot's underlying capability.
Buyers are wary of demonstrations: 69% agree they are not a reliable indicator of production readiness. Yet 89% of those who reported a drop in intervention believe it reflects a genuine gain in autonomy rather than a narrower job.
Finding 4
Buyers' budgets have not followed the difficulty
We asked Buyers to split 100 points across the three parts of a deployment. On their own estimates, robot hardware takes 41% and the AI model and autonomy 28%, roughly 70% between them. Everything else (data, evaluation, operations, integration, safety and support) takes 30%, even though it includes much of what Buyers describe as hardest.
This is not consistent with our working hypothesis that physical AI is a systems problem, with two caveats. First, the spread is wide: hardware alone runs from 10% to 80% across individual Buyers, and 36% of Buyers put at least as much on the surrounding work as on hardware. Second, the question asks Buyers to divide their own cash outlay. Where a Buyer's internal engineering team does the integration, that effort may not show up as deployment spend at all, so the surrounding share could be understated.
71% of Buyers take a complete solution from a robotics vendor, and that barely shifts as deployments mature. What seems to shift is what they add around it: 53% of Buyers at scale also buy hardware separately and 56% use a systems integrator, against 21% and 25% of those still piloting. Buyers are not dropping the packaged vendor as they grow; they are adding integration work on top of it.
Among the 43 Buyers who put at least 30 of their 100 points on the surrounding work, 70% direct the largest part of it to live operations or workflow integration.
A single stacked bar splitting deployment cost across robot hardware 41%, the AI model and autonomy 28%, and everything around them 30%. Hardware and the model together account for roughly 70% of spend, leaving 30% for the surrounding work.
| Category | Share |
|---|---|
| Robot hardware | 41% |
| The AI model and autonomy | 28% |
| Everything around them, meaning data, evaluation, operations, integration, safety and support | 30% |
Where the largest share of that spend goes
Base: the 43 buyers who put at least 30 of their 100 points on the surrounding work. Each buyer’s largest area, not a breakdown of the 30% at left.
| Category | Share |
|---|---|
| Live operations | 37% |
| Workflow and environment integration and customization | 33% |
| Data collection, simulation and evaluation | 16% |
| Safety, security and compliance | 9% |
| Field support and human supervision | 5% |
Finding 5
The economics are moving, perhaps not working
19% of Buyers say the economics of their deployments work well today, while 56% say they don't work yet but are improving over time.
By deployment maturity, 41% of Buyers who have scaled across multiple sites say their economics work well today, against 7% of those in production and 4% of those still piloting. 13 of the 16 Buyers who say the economics work are running at scale.
That does not necessarily mean scaling fixes the economics. It could also be that the deployments that work are the ones that get scaled, while the ones that don't never leave limited production.
Single-select view on deployment economics: do not work yet but improving over time 56%, work well today 19%, too early to assess 18%, do not work yet and not improving meaningfully 6%, not viable for this use case 1%. By stage, the share saying the economics work well today rises from 4% still piloting to 7% in production to 41% scaled across multiple sites.
View on economics today
n=85 buyers · single-select, so values sum to 100%
| Category | Share |
|---|---|
| Do not work yet but improving over time | 56% |
| Work well today | 19% |
| Too early to assess | 18% |
| Do not work yet and not improving meaningfully | 6% |
| Not viable for this use case | 1% |
Say the economics work well today, by stage
Group sizes are small, so treat as directional.
| Category | Share |
|---|---|
| Still piloting | 4% |
| In production in at least one workflow | 7% |
| Scaled across multiple sites | 41% |
The 48 Buyers who say things are improving credit four conditions: less human intervention (69%), the system handling more of the workflow itself (54%), faster deployments (48%) and a falling support burden (42%). Only 25% credit cheaper hardware. Asked what has actually changed over the past 12 months, more Buyers report movement on what the robot does than on what it takes to stand up a new deployment.
Payback is slow. Only 4% of Buyers expect their money back inside a year, and 55% expect it to take more than two years.
Expected payback period in scale order: under 1 year 4%, 1 to 2 years 38%, 2 to 3 years 41%, more than 3 years 14%, do not expect payback at current economics 1%, too early to tell 2%. The 2-to-3-year and 3-plus-year answers combine to 55%.
| Category | Share |
|---|---|
| Under 1 year | 4% |
| 1–2 years | 38% |
| 2–3 years | 41% |
| More than 3 years | 14% |
| Do not expect payback at current economics | 1% |
| Too early to tell | 2% |
Finding 6
Buyers want autonomy inside a boundary
85% of Buyers say human oversight is essential, and 91% stop short of full autonomy. The most common position is that robots act on their own in low-risk situations only (32%).
Overall, 57% are comfortable with autonomous operation inside an explicit boundary with an escalation path, while 34% want a person in the execution loop, either controlling every action or approving before execution.
This matches what Buyers told us about intervention: they narrow what the system attempts, route hard cases to people and add supervision.
Highest level of autonomy buyers are comfortable allowing, in scale order: a person oversees and controls every action 13%, the robot acts and a person approves first 21%, autonomous in low-risk situations only 32%, autonomous in most cases escalating critical ones 25%, fully autonomous 9%. The middle two positions combine to 57%.
| Category | Share |
|---|---|
| Person oversees and controls every action | 13% |
| Robot acts, person approves first | 21% |
| Autonomous in low-risk situations only | 32% |
| Autonomous in most cases, escalates critical ones | 25% |
| Fully autonomous | 9% |
What we take from this
Our paper argued that physical AI is a systems market rather than a model market, and on this evidence the claim holds in part. Buyers put integration and live operations ahead of both the model and the hardware, and they say much the same thing when asked what is stopping them from scaling. Their budgets, though, do not follow: most of what they spend still goes to the robot and the AI model. Buyers appear to experience physical AI as a systems problem while continuing to buy it as a hardware-and-model purchase. A single survey cannot tell us whether that gap is a lag or a correction to our argument.
Running underneath all of this is a measurement problem. A falling intervention rate is one of the numbers this category reaches for most often as evidence of progress, yet fewer than half of these Buyers track it consistently, and most of those reporting a drop achieved it partly by narrowing the robot's job. The Buyers themselves believe the gains are real. But intervention rate alone cannot tell us how much of the gain comes from better autonomy and how much from changing the operating envelope. Until that changes, claims about progress in physical AI, including ours, rest on numbers that the people reporting them cannot fully check.
