HumanEval was saturated a while ago. Frontier models score in the high 90s; even small open models clear 80%. If benchmark scores translated into production software, most software engineering would already be automated.
Meanwhile, we regularly get called into engagements where a team "vibe-coded" its way to a demo in two weeks and then spent four months failing to make it production-worthy. The codebase works right up until it doesn't, nobody can say why a given design decision was made, and the test suite — written by the same model that wrote the bugs — passes cheerfully throughout.
We call the distance between those two facts the vibe-coding gap. It is worth being precise about what lives inside it, because "AI coding doesn't work" and "AI writes all our code now" are both wrong, and the truth pays better than either.
What HumanEval actually measures
HumanEval-style benchmarks (and their descendants — MBPP, LiveCodeBench, and to a lesser degree SWE-bench) measure a specific skill: producing a correct, self-contained function from a complete specification, verified by tests that exist before the code does.
Look at each property. Self-contained: no interaction with a 400k-line codebase, no legacy module with load-bearing quirks. Complete specification: the docstring tells you exactly what success means. Pre-existing tests: the oracle is independent of the author.
Production work inverts every one of these. The spec is ambiguous and half of it lives in someone's head. The code must fit an existing system, respecting conventions and invariants documented nowhere. And the tests usually get written after — by the same model, sharing the same misunderstandings. When the author of the code is also the author of the oracle, "all tests pass" mostly measures internal consistency, not correctness.
SWE-bench moved closer to reality by using real GitHub issues, and it is a genuinely better signal. But it still supplies the two luxuries production never does: a reproducible failing test and a well-scoped, known-fixable problem. Real backlogs are not curated for solvability.
What actually fails in vibe-coded systems
Across the rescue engagements we have taken, the failure modes are consistent — and almost none of them are "the model wrote a wrong function":
Architecture by accretion. Each generated change is locally plausible; the sum is a system with no load-bearing structure. Three different retry mechanisms, two competing auth flows, state duplicated across layers. Models optimize for making the current prompt's request work, and nobody was optimizing for the whole.
Error handling that placates rather than handles. Generated code loves the broad except with a log line. Failures get swallowed, systems keep running in corrupt states, and the incident, when it surfaces, is three subsystems away from its cause.
Silent dependency and version drift. Plausible-but-wrong API usage, pinned-nowhere dependencies, patterns from a library's previous major version. Compiles, runs, works in the demo, breaks under a code path the demo never exercised.
Test suites that mirror the implementation. High coverage numbers, near-zero adversarial value. The tests assert what the code does, not what the system must guarantee.
Note what these have in common: every one is invisible at the function level, which is exactly the level benchmarks measure.
How we use AI coding anyway — heavily
Here is the part that surprises people expecting a Luddite conclusion: we use AI coding assistance on essentially every engagement, and our throughput is dramatically higher for it. The gap is not an argument against AI-written code. It is an argument for engineering the layer benchmarks don't cover. What that looks like in practice:
Humans own the boundaries. Architecture, module contracts, data models, and failure semantics are decided by engineers and written down before generation starts. Models fill in structure; they do not invent it. The quality of AI-generated code tracks the quality of the constraints it is given, almost linearly in our experience.
The oracle must be independent. Acceptance tests, property-based tests, and evals are specified separately from implementation — different session, different author, ideally derived directly from requirements. An AI-written implementation checked by an AI-written-from-the-same-context test suite is one artifact wearing two hats.
CI is the immune system. Static analysis, dependency audit, type checking at maximum strictness, mutation testing where it counts. Machine-generated code volume goes up; therefore machine-enforced quality gates must go up. Review capacity is the scarce resource, so we spend automation to protect it.
Observability is non-negotiable. Vibe-coded systems fail in novel ways, so we instrument aggressively — structured logs, traces, and alerting on behavioral invariants ("queue depth never exceeds X", "this reconciliation always converges"). When accretion-style bugs surface, telemetry is the difference between a one-hour fix and a one-week archaeology dig.
A useful heuristic for the demo-to-production distance: ask what happens when a dependency is down, an input is malformed, and two requests race. If the honest answer is "we haven't looked", the benchmark score of whichever model wrote the code is not the relevant number.
The honest scoreboard
Benchmarks are fine for what they are: a coarse ranking of a narrow capability. The mistake is treating them as a proxy for "can this model build our system." The only eval that answers that question is one you build yourself, on your codebase, with your failure modes — which is the same discipline we apply to every model deployment, coding or otherwise.
Production-or-nothing applies to AI-written software exactly as it applies to AI features: the demo is the start of the work, not the evidence it is done. Models will keep climbing the benchmarks. The teams that win will be the ones engineering everything the benchmarks leave out.