The cost of building software is falling fast. As my colleague Aidan Potts argued in Georgian's Developer Productivity whitepaper, AI is making software cheaper and faster to build and as a result, the people capable of building software now extend beyond professional engineers. Hobbyist developers, product managers, founders and operators are now able to “vibe code” – build software by describing what they want in plain English, with AI writing the underlying code – applications they couldn't have built two years ago.
More software, built by more people, raises an obvious question for me: how do you know if the software built by vibe-coding solutions can actually do what is asked of it? The benchmarks used to evaluate AI coding tools were designed for a narrower problem: can an AI fix a bug in an existing codebase.
But I felt that there was a gap in the ability of market benchmarks to assess whether or not an application built from a plain-English prompt can do what it was asked. That gap motivated me, along with Georgian's AI Lab to collaborate with Replit on ViBench, an open-source benchmark designed to evaluate how AI agents perform on end-to-end, vibe-coded web applications – and that collaboration started to take shape in July 2025.
In our first meeting with Replit’s president we shared the belief that the industry needed a better way to measure AI coding quality - ViBench was born from this idea.
Georgian’s relationship with Replit eventually led to Georgian participating in Replit’s competitive ~$250M Series C round in August 2025 through Georgian Growth Fund VI. We built on our relationship with Replit through our collaboration on ViBench in November 2025. The result – Replit and Georgian collaborated on research, published a paper and presented it at ACM CAIS '26, where it won a spotlight award at the conference (pictured below). This research collaboration strengthened our relationship with Replit and contributed to Georgian leading the company’s $400M Series D in March 2026 through Georgian Alignment Fund II, at a $9B valuation.

You can read the published ViBench paper and view a live leaderboard at vibench.ai.
How Georgian’s AI Lab and Replit Built ViBench
Rather than designing synthetic test cases, the collaboration between the Georgian AI Lab and Replit team (“the ViBench team”) started with real applications that Replit users had already built. The ViBench team (which included Replit engineer Peter Zhong and other members of the Georgian AI Lab) wrote plain-English product requirement documents inspired by anonymized versions of existing apps to build a set of prompts describing what each application should do, providing no technical constraints for the AI on how to build it.
The ViBench team then wrote test plans describing how a real user would interact with the finished product (i.e., clicking through a booking flow, playing a round of the game Mafia, listing an item for sale). An automated evaluator drives those test plans against whatever the AI has built, the same way a Quality Assurance engineer would test an app by using it.
What ViBench Measures
Most existing benchmarks for AI coding tools evaluate performance by asking, "did the code run? The purpose of ViBench is to ask a more meaningful question in my view: did the AI effectively build what the human actually wanted?
In order to evaluate this, ViBench tests AI models on two realistic task types that reflect how people use these tools in the real world:
- Zero-to-One Creation (Building from scratch): Can the AI build a complete, functional app (like a barber shop booking system or a clone of the Slack app) starting from nothing but plain-English description?
- Feature Augmentation (Extending an existing product): Can the AI take an existing app and successfully add a complex new feature without breaking what's already there? Based on my experience, testing feature augmentation is important as it’s more likely that a user wants to build a product feature on top of an existing product. This task is inherently more difficult because in order to add a new feature, the coding agent first has to learn about the current code base before it can build on top of it.
Rather than grading the AI on specific technical details like code structure or database choice, ViBench judges the output the way a real user would by using human-created test plans to assess the product and check whether it does what it was supposed to do. If the AI builds a "Mafia" game, the evaluation agent "plays" the game to see if the rules work, regardless of which database or framework the AI chose to use.
What the Results Show
When we ran the initial evaluation for the CAIS '26 paper, even the best AI models struggled. Building full applications remained a frontier challenge across all AI models tested. Opus 4.6 was the top performer amongst the models assessed, passing 45.7% of its tasks. GPT-5.2 was close behind at 41.9%. No freely available, open-source models exceeded 12%.
Building from scratch appeared to be harder than adding features to something that already existed. Errors also compounded: 7 of 9 models performed worse when extending their own code compared to a clean reference implementation.
Since publication of the paper, the live leaderboard at vibench.ai shows meaningful change. By May 2026, the Zero-to-One sweep across 24 applications shows Opus 4.8 passing 87.8% of its tasks and GPT-5.5 passing 86.5% — suggesting that end-to-end application development capability has advanced since our initial test.
The leaderboard also tracks cost per run alongside performance, giving teams a practical cost-performance framework rather than a capability ranking alone. ViBench updates as new models ship, so the numbers will keep moving.

Screenshot of most recent test results published at vibench.ai/results
What’s Next for ViBench?
To date, Replit is using ViBench as a gatekeeper for its own production, running the benchmark around ~10 times a week to test performance and cost. The ViBench team has engaged with frontier model companies to collaborate on how to adopt ViBench with active conversations in progress with Anthropic and Google. Ideally, we anticipate frontier model providers like Anthropic, Google and OpenAI may use ViBench to optimize the performance and cost of the models they provide users.
ViBench will continue to be a live benchmark so that as model performance improves regularly, more complex tasks can be actively added to the benchmark testing, with live results being updated and posted at vibench.ai.
For a more technical understanding of ViBench, you can read Replit’s ViBench blog here.
This blog is provided for informational purposes only and should not be relied upon as legal, business, investment, or tax advice. Nothing in this blog constitutes investment advice, nor is it intended for use by any investors or prospective investors in any Georgian funds. This blog may include links to external websites or information obtained from third-party sources. Georgian has not independently verified and makes no representations regarding the accuracy or completeness of such information, whether current or ongoing. If this content includes third-party advertisements, Georgian has not reviewed such materials and does not endorse any advertising content or the companies referenced.
Any investments or portfolio companies mentioned are for illustrative purposes only and may not be representative of all investments made by funds managed by Georgian. Please contact Georgian for more information.