Motion

AI, Applied Benchmarks | Wave 3

The Agentic Divide

How advanced AI adopters appear to have pulled away from the rest

Findings from Wave 3 of Georgian + NewtonX AI, Applied Survey, May 2026 (n=501).

Share of Leaders in the Runner tier

The Runner-tier share of surveyed Leaders rose from 12% in Wave 1 (November 2024) to 22% in Wave 3 (May 2026).

Georgian + NewtonX AI, Applied Survey — Wave 1 (Nov. 2024, n=601), Wave 3 (May 2026, n=501).

In Wave 1 of our benchmark, 12% of Decision Makers qualified as Runners, the most AI-mature tier in Georgian’s Crawl, Walk, Run framework. In Wave 3, that share has risen to 22%.

In our view, the gap in AI maturity between companies looks less like a curve and more like a divide.

The three findings below describe:

  • How Runners appear to operate differently across stack, practice and governance.
  • How Runners are using agents in parallel and how spend on AI agents is growing faster than budgets.
  • How the binding constraint on AI-driven engineering appears to be shifting away from productivity and reliability towards verification.
Runner-tier share of Decision Makers

Share of surveyed Leaders classified as Runners, the most AI-mature tier in Georgian’s Crawl, Walk, Run framework. Wave 1 (Nov. 2024, n=601) to Wave 3 (May 2026, n=501).

In Wave 1, 12% of Decision Makers qualified as Runners. In Wave 3, that share reached 22%.

Runner-tier share of Decision Makers12%22%
Wave 1 (Nov. 2024)
Wave 3 (May 2026)
Runner-tier share
Data table for Runner-tier share of Decision Makers
PeriodValue
Wave 1 (Nov. 2024)12%
Wave 3 (May 2026)22%

Finding 1

A look inside advanced AI adopters

Georgian’s Crawl, Walk, Run framework segments Decision Makers into four tiers (Crawler, Walker, Jogger, Runner) based on how AI is operationalized inside the business. Runners are those who report AI projects in production at scale with significant budget commitment. For this analysis we combine Crawlers and Walkers into a single comparison group to create a large enough dataset for statistical analysis.

Across the dimensions where Wave 3 data shows the clearest Runner-vs-rest separation, four themes emerge: stack, practice, governance, and investment & ROI.

Four dimensions where Runners separate

The Wave 3 dimensions with the clearest Runner-vs-rest separation cluster into four themes, a consistent gap across stack, practice, governance and investment.

Stack
+37-41 pts
Custom models, large reasoning models and observability deployed
Practice
More likely to have agentic AI live in production
Governance
+22 pts
Higher guardrail adoption enabling greater agent autonomy
Investment & ROI
Larger AI share of IT budget; revenue-led return expectations
Data table for Four dimensions where Runners separate
IndicatorGap / valueDetail
Stack+37-41 ptsCustom models, large reasoning models and observability deployed
PracticeMore likely to have agentic AI live in production
Governance+22 ptsHigher guardrail adoption enabling greater agent autonomy
Investment & ROILarger AI share of IT budget; revenue-led return expectations

Stack. Runners appear to have invested in a more mature model and infrastructure stack.

Runners are more likely than Crawlers and Walkers to report a combination of custom or fine-tuned models in production, large reasoning models (LRMs) deployed and LLM observability tooling deployed.

These models and tools are associated with the multi-step reasoning, complex planning and long-horizon agentic execution that supports agentic AI.

Stack: model and infrastructure maturity

Share of Tech Leaders reporting each stack signal. Wave 3 (May 2026). Tech Leaders n=252, segmented by AI maturity. Crawler/Walker tiers combined for sample size.

Stack: model and infrastructure maturity0%25%50%75%100%
Large reasoning models (LRMs) deployed
28%65%+37 pts
Custom or fine-tuned models in production
24%65%+41 pts
LLM observability
29%67%+38 pts
Crawler / WalkerRunner
Data table for Stack: model and infrastructure maturity
DimensionCrawler / WalkerRunnerGap
Large reasoning models (LRMs) deployed28%65%+37 pts
Custom or fine-tuned models in production24%65%+41 pts
LLM observability29%67%+38 pts

Source: Georgian + NewtonX AI, Applied Survey, Wave 3. Tech Decision Makers (n=252 segmented by AI maturity). Questions: Which types of AI models is your organization currently using in production systems? (Select all that apply.) Which of the following software infrastructure components do you have in your tech stack to enable your AI products?

Practice. Runners are 5× more likely to have Agentic AI live in production.

62% of Runners report having agentic AI live and expanding in production, against 13% of Crawlers/Walkers. Among Runners using automated coding tools, 53% report parallel-agent workflows (where most or all engineers are running multiple agents on distinct tasks concurrently) against 7% of Crawlers and Walkers.

Perhaps thanks to these agentic and infrastructure advantages, Runners are able to automate qualitatively harder problems using AI. 42% of Runners describe their internal AI workflows as mostly or very complex (multi-step processes spanning multiple tools or teams) against 10% among Crawlers and Walkers.

Practice: agentic AI past the lab

Share of Tech Leaders reporting each practice signal. Wave 3 (May 2026). Tech Leaders n=252, segmented by AI maturity. Crawler/Walker tiers combined.

Practice: agentic AI past the lab0%25%50%75%100%
Agentic AI live in production and expanding
13%62%+49 pts
Parallel-agent coding workflows in use
7%53%+46 pts
Automating complex or very complex internal workflows
10%42%+32 pts
Crawler / WalkerRunner
Data table for Practice: agentic AI past the lab
DimensionCrawler / WalkerRunnerGap
Agentic AI live in production and expanding13%62%+49 pts
Parallel-agent coding workflows in use7%53%+46 pts
Automating complex or very complex internal workflows10%42%+32 pts

Source: Georgian + NewtonX AI, Applied Survey, Wave 3. Tech Decision Makers (n=252 segmented by AI maturity). Questions: To what extent is your organization implementing or planning to implement Agentic AI in the next 6 months? Which of the following best describes how your engineers are currently using agentic coding tools (e.g., Claude Code, OpenAI Codex, OpenClaw, or equivalent tools)? How would you describe the complexity of your organization's internal AI workflows?

Governance. Runners appear to have built more guardrails which allow for greater agentic autonomy.

Across six categories of agentic guardrails (sandboxed environments, allowlists, human approval gates, policy-as-code, role-based permissions and monitoring or audit logs), Runners report an average adoption rate of 59%, against 37% among Crawlers and Walkers.

This investment in agentic guardrails may contribute to Runners having greater comfort with higher levels of agentic autonomy. 38% of Runners report comfort letting agentic AI operate with full or near-full autonomy in production, against 13% among Crawlers and Walkers. In our view, the guardrails come first: the eval, observability and policy-as-code investment is what lets Runners hand more decisions to AI without per-action human approval.

Governance: guardrails enabling autonomy

Share of Tech Leaders reporting each governance signal. Wave 3 (May 2026). Tech Leaders n=252, segmented by AI maturity. Guardrail figure averages six control categories.

Governance: guardrails enabling autonomy0%25%50%75%100%
Average AI guardrail implementation (across six categories)
37%59%+22 pts
Comfortable with full or scoped autonomy for agents
13%38%+25 pts
Crawler / WalkerRunner
Data table for Governance: guardrails enabling autonomy
DimensionCrawler / WalkerRunnerGap
Average AI guardrail implementation (across six categories)37%59%+22 pts
Comfortable with full or scoped autonomy for agents13%38%+25 pts

Source: Georgian + NewtonX AI, Applied Survey, Wave 3. Tech Decision Makers (n=252 segmented by AI maturity). Questions: Which of the following guardrails or controls are currently implemented for your organization's agentic AI systems (e.g., AI agents that can take actions, call tools, or trigger workflows)? Select all that apply. Underlying categories: sandboxed environments, allowlists, human approval gates, policy-as-code, role-based permissions, monitoring or audit logs. What level of autonomy are you comfortable granting to agentic AI systems?

Investment and ROI. The median Runner reports AI accounting for 3× more budget share.

Runners report a 3× bigger share of their IT budget being dedicated to AI than Crawlers and Walkers. This budget is, in our view, what allows for greater investment in agentic infrastructure and guardrails.

In return, Runners are more likely to be expecting revenue from AI projects than efficiency gains alone. Runners are 16 points more likely to cite new revenue, rather than efficiency alone, as a top-three reason for investing in AI.

Investment & ROI: budget share and return expectations

Share of Tech Leaders reporting each investment signal. Wave 3 (May 2026). Tech Leaders n=252, segmented by AI maturity. GenAI spend figure is the median share of total IT spend.

Investment & ROI: budget share and return expectations0%25%50%75%100%
GenAI share of total IT spend (median)
10%30%+20 pts
New revenue as top-3 AI investment rationale
36%52%+16 pts
Crawler / WalkerRunner
Data table for Investment & ROI: budget share and return expectations
DimensionCrawler / WalkerRunnerGap
GenAI share of total IT spend (median)10%30%+20 pts
New revenue as top-3 AI investment rationale36%52%+16 pts

Source: Georgian + NewtonX AI, Applied Survey, Wave 3. Tech Decision Makers (n=252 segmented by AI maturity). Questions: Approximately what percentage of your organization's total IT spend in the past 12 months was allocated to AI tools, infrastructure, and related services? What is your organization's primary expected return from investing in Agentic AI? Please select up to 3 ROI reasons where #1 is the most important.

Finding 2

AI coding spend set to double in 90 days as the per-seat budget model appears to be breaking

Based on Wave 3 responses, the cost of operating like a Runner appears to be rising. The data also suggests increased spend could surface as a budget question before it shows up in ROI tracking as revenue.

The cost of operating like a Runner is rising

Tech Leaders with automated coding in production expecting per-engineer monthly agentic-coding spend to at least double within three months. Wave 3 (May 2026), n=227.

63% expect costs to double in next 3 months

0%

Georgian + NewtonX AI, Applied Survey, Wave 3, Tech Leaders with automated coding in production (n=227)

63% — 63% expect costs to double in next 3 months

63% of Tech Decision Makers report engineers now running multiple AI agents simultaneously: 80% among Runners against 45% among Crawlers and Walkers. Meanwhile, almost half of Tech Decision Makers expect flat or declining entry-level headcount over the next 12 months.

Per-seat budgeting assumes cost rises with headcount. Agentic coding tooling cost tends to rise with usage, a different cost shape entirely.

Current monthly spend per engineer on AI tokens and tool plans

Distribution of Tech Leaders by estimated average monthly spend per engineer on AI tokens and/or monthly AI tool plans. Wave 3 (May 2026). The modal band is $251–$1,000.

The modal Tech Leader spends $251–$1,000 per engineer per month on AI tokens and tool plans today, a figure most expect to at least double within 90 days.

Current monthly spend per engineer on AI tokens and tool plans
Monthly spend per engineer
9%25%33%15%8%10%
0%100%
$0–$100$101–$250$251–$1,000 (modal)$1,001–$5,000More than $5,000Unsure / no answer
Data table for Current monthly spend per engineer on AI tokens and tool plans
CategoryShare
$0–$1009%
$101–$25025%
$251–$1,000 (modal)33%
$1,001–$5,00015%
More than $5,0008%
Unsure / no answer10%

The investment in coding agents is showing clear benefits.

Speed to production has accelerated across the board:

  • 71% of Tech Decision Makers report moving AI features from pilot to production in under six months, and 25% in under three.

But the self-reported revenue impact is uneven:

  • Runners report 70% positive AI impact on new revenue, against 39% among Crawlers and Walkers.
Everyone is shipping faster, revenue impact is uneven

Speed-to-production and self-reported revenue impact among Tech Leaders. Wave 3 (May 2026), Tech Leaders n=252.

Ship pilot to production in under 6 months
71%
All Tech Leaders
Ship pilot to production in under 3 months
25%
All Tech Leaders
Positive AI impact on new revenue (Runners)
70%
Runner-tier Tech Leaders
Positive AI impact on new revenue (Crawlers / Walkers)
39%
Crawler/Walker-tier Tech Leaders
Data table for Everyone is shipping faster, revenue impact is uneven
IndicatorGap / valueDetail
Ship pilot to production in under 6 months71%All Tech Leaders
Ship pilot to production in under 3 months25%All Tech Leaders
Positive AI impact on new revenue (Runners)70%Runner-tier Tech Leaders
Positive AI impact on new revenue (Crawlers / Walkers)39%Crawler/Walker-tier Tech Leaders

Most Tech Decision Makers cannot yet quantify what they are getting back. 45% report actively measuring AI ROI against financial outcomes (23% against revenue, 22% against cost savings) while 24% see benefits but have not translated them into specific revenue or cost impact. A further 15% have a financial framework but do not track against it, and 14% do not know their AI ROI at all.

Most Tech Leaders cannot yet quantify their ROI on AI

How Tech Leaders track AI ROI. Wave 3 (May 2026), Tech Leaders n=252. The measured share splits into 23% against revenue and 22% against cost savings.

Most Tech Leaders cannot yet quantify their ROI on AI
AI ROI measurement
45%24%15%14%
0%100%
Measuring against financial outcomesSee benefits, cannot yet quantifyHave a framework, not trackingDo not know AI ROI
Data table for Most Tech Leaders cannot yet quantify their ROI on AI
CategoryShare
Measuring against financial outcomes45%
See benefits, cannot yet quantify24%
Have a framework, not tracking15%
Do not know AI ROI14%

Finding 3

As software development speeds up, the bottleneck is shifting from developer productivity to verification

In Wave 2 (June 2025), our data showed AI delivering strong gains on developer productivity metrics but lagging well behind on software reliability metrics, a 35-point gap between the two. The Wave 3 data suggests that gap has closed substantially: reliability has caught up to within 10 points of productivity.

Reliability impact has caught up to productivity

Year-over-year change in Tech Leaders reporting positive AI impact on reliability metrics. Wave 2 (Jun. 2025, n=634) to Wave 3 (May 2026, Tech Leaders n=252).

+37-point year-over-year rise in Tech Leaders reporting positive AI impact on reliability metrics

+0 pts

Average across four reliability metrics, Wave 2 → Wave 3

+ 37 pts — +37-point year-over-year rise in Tech Leaders reporting positive AI impact on reliability metrics

71% of Tech Decision Makers now report AI yielding positive impact on developer productivity metrics, up 13 points from Wave 2. Reliability metrics moved further still: the average across four reliability metrics rose from 23% in June 2025 to 60% in May 2026, a 37-point increase.

Productivity vs reliability: positive AI impact

Share of Tech Leaders reporting positive AI impact on developer productivity metrics and on reliability metrics (average of four). Wave 2 (Jun. 2025) to Wave 3 (May 2026).

Reliability impact, which trailed productivity by 35 points in Wave 1, has closed to within roughly 10 points by Wave 3.

Productivity vs reliability: positive AI impact58%71%23%60%
Wave 2 (Jun. 2025)
Wave 3 (May 2026)
% reporting positive impact
Productivity metricsReliability metrics (avg. of 4)
Data table for Productivity vs reliability: positive AI impact
PeriodProductivity metricsReliability metrics (avg. of 4)
Wave 2 (Jun. 2025)58%23%
Wave 3 (May 2026)71%60%

In our view, while the data paints a rosy picture for productivity and reliability, it also raises a question about whether these are still the right metrics. Thanks to the extensive use of coding agents described in Finding 2, developer productivity and reliability appear to be moving toward table-stakes.

Nahim Nasser, Head of Engineering at Georgian’s AI Lab, points to verifiability and trust as the next inflection point. As agentic AI speeds up software development and a new wave of goal-oriented vibe-coding tools comes to market, the question shifts from how fast engineers can produce code to how confidently teams can verify what was produced. In practice, that means eval suites, runtime monitoring of agent behaviour, and human review workflows for high-stakes diffs.

This connects back to Finding 1: Runners already report the LLM observability tooling and the guardrail scaffolding that kind of verification depends on.

Tying the three findings together

Runners appear to be operating AI differently, not at higher volume on the same stack, but with a different stack, a different practice and a different governance posture (Finding 1). That different way of working has a cost shape the per-seat budget model was not built for (Finding 2). And it leans on the eval, observability and guardrail layer that, in our view, will increasingly separate teams that can verify agentic output at scale from teams that cannot (Finding 3).

The report’s three findings build on one another: a different operating model (Finding 1), a usage-based cost shape (Finding 2), and a verification and guardrail layer (Finding 3). Together they form a divide between teams that can verify agentic output at scale and teams that cannot.

Each finding rests on the one below it — and the layer on top is what increasingly separates the field.

Methodology

Wave 3 of Georgian’s AI, Applied Benchmarks Survey was fielded in April and May 2026 in partnership with NewtonX, surveying 501 Decision Makers, including 252 Tech Decision Makers, at B2B growth-stage software and enterprise companies in 4 countries (US, Canada, UK, Israel). The survey is one of a series of waves (Wave 1: Nov. 2024, n=601; Wave 2: Jun. 2025, n=634).

Definitions

  • Tech Decision Makers: Director level and above employees who responded to any wave of the AI survey administered by NewtonX with decision-making authority at B2B technology companies, including engineering, security product and data leadership. In Wave 1 and Wave 2, we referred to this cohort as R&D Respondents, but have adjusted our terminology to more accurately reflect the make-up of these survey participants.
  • Runners: Decision Makers whose organizations qualify as Runners under Georgian's Crawl, Walk, Run framework: AI projects in production at scale, with significant budget commitment. Used for the wave-over-wave Runner share comparison only.

Key definitions:

  • Crawl, Walk, Jog, Run framework: Georgian’s four-tier AI maturity classification. Agentic AI: AI systems that take autonomous actions across multi-step workflows.
  • Automated coding: professional-developer tools (Claude Code, Cursor, GitHub Copilot).
  • Vibe coding: citizen-developer-friendly tools (Replit★, Bolt, Lovable).
  • LRMs: foundation models optimized for multi-step reasoning.
  • Durable workflow engines: orchestration infrastructure for long-running multi-step AI workflows.

★ Indicates a Georgian portfolio company.

These materials, including the AI, Applied Benchmarks Wave 3 results and any related stories (the "Information"), are prepared by Georgian Partners Growth LP Inc. and its affiliates (collectively, "Georgian") for discussion and informational purposes only. The Wave 3 Survey was administered on a blind basis by NewtonX, Inc. ("NewtonX") and sent to individuals at B2B software companies identified and selected by NewtonX. Georgian was not involved in the selection of companies or Respondents may be individuals who work for current or former portfolio companies of Georgian. Respondents were compensated by NewtonX for their work in connection with the Survey, and NewtonX was compensated by Georgian; while the anonymous nature of the Survey mitigates against any conflicts, such compensation may subject Respondents to potential conflicts of interest in connection with their responses.

While the Information may be based on third-party sources Georgian believes to be reliable, Georgian has not independently verified all such information and does not guarantee that it is accurate, complete, or up to date. The Information is not an offer to sell securities, nor should it be deemed to imply an offer of securities, and may not be relied upon for making any investment decision with respect to any fund, vehicle, or product managed by Georgian. Nothing herein is intended to be representative of Georgian's prior or current investments or of the firm's investment experience or performance as a whole.

Logos or company names of third parties are the trademarked property of the respective companies and do not suggest any affiliation, endorsement, or sponsorship of Georgian or any managed investment vehicle.

This document may contain forward-looking statements identified by words such as "believe," "anticipate," "expect," "may," "will," and similar expressions. These statements are based on Georgian's expectations and are not guarantees of future performance; actual outcomes may differ materially. Readers should not place undue reliance on forward-looking statements. Nothing herein constitutes investment, tax, financial, business, legal, or other advice. Past performance is not indicative of future results.

All currency in US dollars (USD) except as otherwise indicated. All Information is as of May 2026 except as otherwise indicated.

About NewtonX: NewtonX is the research and insights platform that empowers businesses to solve their toughest challenges with confidence. Visit newtonx.com to learn more. About Georgian: Georgian is a growth equity firm investing in B2B technology companies. Georgian’s in-house AI Lab works with portfolio companies on production AI deployment. Visit georgian.io for more information.