AI, Applied Benchmarks | Wave 3
The Agentic Divide
How advanced AI adopters appear to have pulled away from the rest
Findings from Wave 3 of Georgian + NewtonX AI, Applied Survey, May 2026 (n=501).
Share of Leaders in the Runner tier
The Runner-tier share of surveyed Leaders rose from 12% in Wave 1 (November 2024) to 22% in Wave 3 (May 2026).
Georgian + NewtonX AI, Applied Survey — Wave 1 (Nov. 2024, n=601), Wave 3 (May 2026, n=501).
In Wave 1 of our benchmark, 12% of Decision Makers qualified as Runners, the most AI-mature tier in Georgian’s Crawl, Walk, Run framework. In Wave 3, that share has risen to 22%.
In our view, the gap in AI maturity between companies looks less like a curve and more like a divide.
The three findings below describe:
- How Runners appear to operate differently across stack, practice and governance.
- How Runners are using agents in parallel and how spend on AI agents is growing faster than budgets.
- How the binding constraint on AI-driven engineering appears to be shifting away from productivity and reliability towards verification.
Share of surveyed Leaders classified as Runners, the most AI-mature tier in Georgian’s Crawl, Walk, Run framework. Wave 1 (Nov. 2024, n=601) to Wave 3 (May 2026, n=501).
In Wave 1, 12% of Decision Makers qualified as Runners. In Wave 3, that share reached 22%.
| Period | Value |
|---|---|
| Wave 1 (Nov. 2024) | 12% |
| Wave 3 (May 2026) | 22% |
Finding 1
A look inside advanced AI adopters
Georgian’s Crawl, Walk, Run framework segments Decision Makers into four tiers (Crawler, Walker, Jogger, Runner) based on how AI is operationalized inside the business. Runners are those who report AI projects in production at scale with significant budget commitment. For this analysis we combine Crawlers and Walkers into a single comparison group to create a large enough dataset for statistical analysis.
Across the dimensions where Wave 3 data shows the clearest Runner-vs-rest separation, four themes emerge: stack, practice, governance, and investment & ROI.
The Wave 3 dimensions with the clearest Runner-vs-rest separation cluster into four themes, a consistent gap across stack, practice, governance and investment.
| Indicator | Gap / value | Detail |
|---|---|---|
| Stack | +37-41 pts | Custom models, large reasoning models and observability deployed |
| Practice | 5× | More likely to have agentic AI live in production |
| Governance | +22 pts | Higher guardrail adoption enabling greater agent autonomy |
| Investment & ROI | 3× | Larger AI share of IT budget; revenue-led return expectations |
Stack. Runners appear to have invested in a more mature model and infrastructure stack.
Runners are more likely than Crawlers and Walkers to report a combination of custom or fine-tuned models in production, large reasoning models (LRMs) deployed and LLM observability tooling deployed.
These models and tools are associated with the multi-step reasoning, complex planning and long-horizon agentic execution that supports agentic AI.
Share of Tech Leaders reporting each stack signal. Wave 3 (May 2026). Tech Leaders n=252, segmented by AI maturity. Crawler/Walker tiers combined for sample size.
| Dimension | Crawler / Walker | Runner | Gap |
|---|---|---|---|
| Large reasoning models (LRMs) deployed | 28% | 65% | +37 pts |
| Custom or fine-tuned models in production | 24% | 65% | +41 pts |
| LLM observability | 29% | 67% | +38 pts |
Practice. Runners are 5× more likely to have Agentic AI live in production.
62% of Runners report having agentic AI live and expanding in production, against 13% of Crawlers/Walkers. Among Runners using automated coding tools, 53% report parallel-agent workflows (where most or all engineers are running multiple agents on distinct tasks concurrently) against 7% of Crawlers and Walkers.
Perhaps thanks to these agentic and infrastructure advantages, Runners are able to automate qualitatively harder problems using AI. 42% of Runners describe their internal AI workflows as mostly or very complex (multi-step processes spanning multiple tools or teams) against 10% among Crawlers and Walkers.
Share of Tech Leaders reporting each practice signal. Wave 3 (May 2026). Tech Leaders n=252, segmented by AI maturity. Crawler/Walker tiers combined.
| Dimension | Crawler / Walker | Runner | Gap |
|---|---|---|---|
| Agentic AI live in production and expanding | 13% | 62% | +49 pts |
| Parallel-agent coding workflows in use | 7% | 53% | +46 pts |
| Automating complex or very complex internal workflows | 10% | 42% | +32 pts |
Governance. Runners appear to have built more guardrails which allow for greater agentic autonomy.
Across six categories of agentic guardrails (sandboxed environments, allowlists, human approval gates, policy-as-code, role-based permissions and monitoring or audit logs), Runners report an average adoption rate of 59%, against 37% among Crawlers and Walkers.
This investment in agentic guardrails may contribute to Runners having greater comfort with higher levels of agentic autonomy. 38% of Runners report comfort letting agentic AI operate with full or near-full autonomy in production, against 13% among Crawlers and Walkers. In our view, the guardrails come first: the eval, observability and policy-as-code investment is what lets Runners hand more decisions to AI without per-action human approval.
Share of Tech Leaders reporting each governance signal. Wave 3 (May 2026). Tech Leaders n=252, segmented by AI maturity. Guardrail figure averages six control categories.
| Dimension | Crawler / Walker | Runner | Gap |
|---|---|---|---|
| Average AI guardrail implementation (across six categories) | 37% | 59% | +22 pts |
| Comfortable with full or scoped autonomy for agents | 13% | 38% | +25 pts |
Investment and ROI. The median Runner reports AI accounting for 3× more budget share.
Runners report a 3× bigger share of their IT budget being dedicated to AI than Crawlers and Walkers. This budget is, in our view, what allows for greater investment in agentic infrastructure and guardrails.
In return, Runners are more likely to be expecting revenue from AI projects than efficiency gains alone. Runners are 16 points more likely to cite new revenue, rather than efficiency alone, as a top-three reason for investing in AI.
Share of Tech Leaders reporting each investment signal. Wave 3 (May 2026). Tech Leaders n=252, segmented by AI maturity. GenAI spend figure is the median share of total IT spend.
| Dimension | Crawler / Walker | Runner | Gap |
|---|---|---|---|
| GenAI share of total IT spend (median) | 10% | 30% | +20 pts |
| New revenue as top-3 AI investment rationale | 36% | 52% | +16 pts |
Finding 2
AI coding spend set to double in 90 days as the per-seat budget model appears to be breaking
Based on Wave 3 responses, the cost of operating like a Runner appears to be rising. The data also suggests increased spend could surface as a budget question before it shows up in ROI tracking as revenue.
Tech Leaders with automated coding in production expecting per-engineer monthly agentic-coding spend to at least double within three months. Wave 3 (May 2026), n=227.
63% expect costs to double in next 3 months
0%
Georgian + NewtonX AI, Applied Survey, Wave 3, Tech Leaders with automated coding in production (n=227)
63% — 63% expect costs to double in next 3 months
63% of Tech Decision Makers report engineers now running multiple AI agents simultaneously: 80% among Runners against 45% among Crawlers and Walkers. Meanwhile, almost half of Tech Decision Makers expect flat or declining entry-level headcount over the next 12 months.
Per-seat budgeting assumes cost rises with headcount. Agentic coding tooling cost tends to rise with usage, a different cost shape entirely.
Distribution of Tech Leaders by estimated average monthly spend per engineer on AI tokens and/or monthly AI tool plans. Wave 3 (May 2026). The modal band is $251–$1,000.
The modal Tech Leader spends $251–$1,000 per engineer per month on AI tokens and tool plans today, a figure most expect to at least double within 90 days.
| Category | Share |
|---|---|
| $0–$100 | 9% |
| $101–$250 | 25% |
| $251–$1,000 (modal) | 33% |
| $1,001–$5,000 | 15% |
| More than $5,000 | 8% |
| Unsure / no answer | 10% |
The investment in coding agents is showing clear benefits.
Speed to production has accelerated across the board:
- 71% of Tech Decision Makers report moving AI features from pilot to production in under six months, and 25% in under three.
But the self-reported revenue impact is uneven:
- Runners report 70% positive AI impact on new revenue, against 39% among Crawlers and Walkers.
Speed-to-production and self-reported revenue impact among Tech Leaders. Wave 3 (May 2026), Tech Leaders n=252.
| Indicator | Gap / value | Detail |
|---|---|---|
| Ship pilot to production in under 6 months | 71% | All Tech Leaders |
| Ship pilot to production in under 3 months | 25% | All Tech Leaders |
| Positive AI impact on new revenue (Runners) | 70% | Runner-tier Tech Leaders |
| Positive AI impact on new revenue (Crawlers / Walkers) | 39% | Crawler/Walker-tier Tech Leaders |
Most Tech Decision Makers cannot yet quantify what they are getting back. 45% report actively measuring AI ROI against financial outcomes (23% against revenue, 22% against cost savings) while 24% see benefits but have not translated them into specific revenue or cost impact. A further 15% have a financial framework but do not track against it, and 14% do not know their AI ROI at all.
How Tech Leaders track AI ROI. Wave 3 (May 2026), Tech Leaders n=252. The measured share splits into 23% against revenue and 22% against cost savings.
| Category | Share |
|---|---|
| Measuring against financial outcomes | 45% |
| See benefits, cannot yet quantify | 24% |
| Have a framework, not tracking | 15% |
| Do not know AI ROI | 14% |
Finding 3
As software development speeds up, the bottleneck is shifting from developer productivity to verification
In Wave 2 (June 2025), our data showed AI delivering strong gains on developer productivity metrics but lagging well behind on software reliability metrics, a 35-point gap between the two. The Wave 3 data suggests that gap has closed substantially: reliability has caught up to within 10 points of productivity.
Year-over-year change in Tech Leaders reporting positive AI impact on reliability metrics. Wave 2 (Jun. 2025, n=634) to Wave 3 (May 2026, Tech Leaders n=252).
+37-point year-over-year rise in Tech Leaders reporting positive AI impact on reliability metrics
+0 pts
Average across four reliability metrics, Wave 2 → Wave 3
+ 37 pts — +37-point year-over-year rise in Tech Leaders reporting positive AI impact on reliability metrics
71% of Tech Decision Makers now report AI yielding positive impact on developer productivity metrics, up 13 points from Wave 2. Reliability metrics moved further still: the average across four reliability metrics rose from 23% in June 2025 to 60% in May 2026, a 37-point increase.
Share of Tech Leaders reporting positive AI impact on developer productivity metrics and on reliability metrics (average of four). Wave 2 (Jun. 2025) to Wave 3 (May 2026).
Reliability impact, which trailed productivity by 35 points in Wave 1, has closed to within roughly 10 points by Wave 3.
| Period | Productivity metrics | Reliability metrics (avg. of 4) |
|---|---|---|
| Wave 2 (Jun. 2025) | 58% | 23% |
| Wave 3 (May 2026) | 71% | 60% |
In our view, while the data paints a rosy picture for productivity and reliability, it also raises a question about whether these are still the right metrics. Thanks to the extensive use of coding agents described in Finding 2, developer productivity and reliability appear to be moving toward table-stakes.
Nahim Nasser, Head of Engineering at Georgian’s AI Lab, points to verifiability and trust as the next inflection point. As agentic AI speeds up software development and a new wave of goal-oriented vibe-coding tools comes to market, the question shifts from how fast engineers can produce code to how confidently teams can verify what was produced. In practice, that means eval suites, runtime monitoring of agent behaviour, and human review workflows for high-stakes diffs.
This connects back to Finding 1: Runners already report the LLM observability tooling and the guardrail scaffolding that kind of verification depends on.
Tying the three findings together
Runners appear to be operating AI differently, not at higher volume on the same stack, but with a different stack, a different practice and a different governance posture (Finding 1). That different way of working has a cost shape the per-seat budget model was not built for (Finding 2). And it leans on the eval, observability and guardrail layer that, in our view, will increasingly separate teams that can verify agentic output at scale from teams that cannot (Finding 3).
The report’s three findings build on one another: a different operating model (Finding 1), a usage-based cost shape (Finding 2), and a verification and guardrail layer (Finding 3). Together they form a divide between teams that can verify agentic output at scale and teams that cannot.
Each finding rests on the one below it — and the layer on top is what increasingly separates the field.