Your agent POC scorecard almost certainly ends in a weighted average. A benchmark published this week shows that number and the count of deliverables your team could actually use are not the same measurement — and that the first one compresses the difference between vendors into a fraction of what the second one shows. On StartupBench, posted to arXiv on 18 August 2026, nine frontier agents averaged as high as 73.67 out of 100 across 97 workflows taken from products AI startups already sell. The best any of them managed was an acceptable deliverable on 31.27% of runs.
Those two numbers are computed from exactly the same rubric, on exactly the same runs. One of them is telling you something useful about Monday.
What StartupBench Actually Measured
StartupBench is 97 end-to-end tasks reverse-engineered from products that AI-native startups already sell to paying customers, graded against an average of 25.3 explicit requirements each. That construction is the point. Most agent benchmarks are assembled from tasks researchers think are representative; the authors instead went to products with demand attached, retaining "startups with more than USD 1M in funding" and requiring "evidence of real adoption, either through paid usage or substantial user traction," then interviewing over 30 deep users and recruiting over 50 domain experts to reconstruct the workflows.
The tasks span six domains — Medical & Healthcare (21.6%), Business & Management (19.6%), Finance (18.6%), Legal (16.5%), STEM & Computer Science (16.5%) and Education & Humanities (7.2%) — and they end in artifacts, not answers: DOCX, XLSX, PPTX, PDF, Markdown and images. Every model ran under "the same Nanobot harness with an identical tool configuration" and a maximum of 200 interaction steps.
A rubric item here is one checkable statement about the finished artifact, not a judgement about the agent's reasoning. Each one is graded satisfied or not, and each carries a weight: 5 for Core, 3 for Important, 1 for Auxiliary. Human experts agreed with the automated grader on 92.78% of rubric-level judgements and 92.84% of the pass/fail success decisions, which is the number that makes the rest of the paper worth reading.
A 73 Average Means Two Failed Deliverables in Three
Across all nine models, the average score and the share of runs that produced a usable deliverable disagree by a factor of two or more. StartupBench counts a task successful when "its final score is at least 90; otherwise, it is counted as a failure." Here is the leaderboard with both columns side by side — which is the part most leaderboards omit.
| Model | Average score | Success rate (score ≥ 90) |
|---|---|---|
| Kimi-K3 | 73.67 | 29.55% |
| GPT-5.6-sol | 73.61 | 31.27% |
| GPT-5.5 | 72.79 | 26.80% |
| Seed-2.1-Pro | 67.19 | 22.34% |
| DeepSeek-V4-Pro | 61.11 | 16.49% |
| GLM-5.1 | 60.79 | 16.49% |
| Kimi-K2.6 | 59.95 | 13.06% |
| Qwen-3.6-Max | 59.46 | 15.12% |
| Gemini-3.1-Pro | 49.73 | 6.53% |
Read the first and last rows together. Kimi-K3's average is 1.48x Gemini-3.1-Pro's; its usable-output rate is 4.53x. The average compresses the field into a narrow band where every vendor looks like a reasonable choice. The pass/fail count does not, and it is the one that maps onto your backlog: at 6.53%, an agent needs fifteen attempts to hand back one artifact a professional would accept.
The orderings also disagree outright. The paper notes that "Kimi-K3 achieves the highest average score overall but a lower success rate than GPT-5.6-sol, while Kimi-K2.6 similarly attains a higher average score than Qwen-3.6-Max but a lower success rate." Be fair about the size of that: 0.06 points separates Kimi-K3 from GPT-5.6-sol on the average, which is noise. That is exactly the problem. Where the average says "tied," the acceptance count says one vendor delivers 1.7 percentage points more finished work — and the tie is the thing your steering committee will act on.
None of this is unique to agents built on any one vendor's model. DeepSeek-V4-Pro and GLM-5.1 land on identical success rates from different averages.
Agents Miss the Requirements That Matter Most
The reason a partial-credit average flatters agents is that agents systematically bank the cheap points. Averaged across models, rubric satisfaction "decreases from 68.67% for Auxiliary rubrics to 65.89% for Important rubrics and 63.45% for Core rubrics," and eight of the nine models show that decline.
Those categories have precise definitions in the paper. Core rubrics "capture requirements that directly determine whether the primary user needs are satisfied; violating them would substantially compromise the correctness, reliability, safety, or usability of the final deliverable." Auxiliary rubrics "capture lower-level requirements whose omission primarily affects details, user experience, visual appeal, or polish rather than the fundamental usability of the deliverable."
So the model is 5.2 percentage points better at formatting than at the thing you bought it for. Layout, tone and file polish — the parts a reviewer notices in the first ten seconds — are the parts it gets right most often, and the correctness of the analysis underneath is where it slips.
This is not fixed by weighting core requirements harder. StartupBench already does: Core, Important and Auxiliary "account for an average of 59.10%, 34.33%, and 6.57% of the total task weight." Core is already the clear majority of the score, and a 73.67 average still meant fewer than three deliverables in ten were acceptable. Run the arithmetic on those published weights yourself. An agent that satisfies every Important and Auxiliary requirement and misses one core requirement in five scores 88.2 — a comfortable B+ on any scorecard in your organisation, and a failure under the paper's ≥90 bar. Miss a quarter of them and you still score 85.
That is the whole finding in one line: a weighted average lets an agent lose the deliverable and keep the grade.
Your POC Scorecard Has the Same Flaw
Published enterprise agent evaluation frameworks are built in the same shape StartupBench just exposed — a weighted percentage rollup in which the actual work is a minority of the number. One enterprise AI agent evaluation framework weights Security & Compliance at 25%, Task Performance & Accuracy at 20%, Integration & Technical Fit at 18%, Total Cost of Ownership at 15%, Vendor Stability & Roadmap at 10%, Support & Success Services at 7% and User Experience & Adoption at 5%.
Look at that second line. Whether the agent produced work your team can use is 20% of the number that goes on the slide. StartupBench gave the equivalent line 59.10% of the weight and it still produced a misleading headline. At 20%, an agent can fail most of your test cases and still clear a 75 if the SOC 2 report is clean and the connector list is long. That framework is not an outlier: a separate agentic AI vendor evaluation checklist puts Reliability and evaluation at 15% and Capability and architecture at 20%, against 20% each for security and for governance.
None of those other dimensions is wrong to score. Security, integration cost and vendor durability all belong on the sheet — that is why scoring a platform on its exit criteria rather than its feature list is the right instinct. The error is the aggregation. Security and accuracy are not substitutable goods, and averaging them together implies they are. NIST's AI Risk Management Framework, published 26 January 2023, names Measure as one of its four core functions alongside Govern, Map and Manage — but it does not hand you the aggregation rule. Choosing to collapse the measurement into one number is a decision your organisation made, and it is reversible.
The benchmark community has already moved the other way. OpenAI's GDPval, covering 44 occupations across the nine top sectors of US GDP, grades deliverables by blinded expert comparison against human work rather than by rubric average. τ-bench reports pass^k — the probability that all k attempts succeed — and found gpt-4o succeeding on under 50% of tasks with "pass^8 <25% in retail." Sierra's write-up of that result puts it plainly: the benchmark "doesn't just test whether an agent can complete a task once; it measures whether it can do so consistently multiple times."
The Agent Will Report That It Passed
Do not let the agent grade its own deliverable, because it is measurably bad at that specific job. StartupBench documents a failure mode in which the model "mistakes confidence in its own reasoning process for evidence that the generated artifact satisfies the acceptance criteria," summarising its intended plan "as evidence that the task has been verified, without independently inspecting the final deliverable."
That is not a hallucination about the world. It is a hallucination about its own output, and it survives every guardrail you have pointed at factual accuracy. It also shows up in the plainest possible way: output-format compliance ranged from 97.6% down to 86.3% across models, with violations clustering into "generating an incorrect file type" and "omitting part of the requested deliverables." At the bottom of that range, nearly one deliverable in seven broke the format contract.
This matters most for the human on the other end. Prior work has shown that vaguer AI explanations drew more trust from novice reviewers, not less. An agent that confidently narrates a verification it never performed is optimised, accidentally, for exactly that bias. If your POC acceptance step is a delivery-team member reading the agent's summary, you are measuring the summary. This is also why monitoring that captures the artifact and not just the trace is worth more than another dashboard.
What This Benchmark Does Not Prove
Take the strongest version of the objections, because several of them are real.
The harness was tested, but only on the average. The paper says every model ran under "the same Nanobot harness," and gives no citation, version or repository for it — the most prominent open-source project of that name is a 47k-star personal AI agent framework, not an evaluation harness. The authors did, however, put the obvious objection to the test: they re-ran GPT-5.5, GLM-5.1 and Qwen-3.6-Max under two other frameworks, Hermes and Claude Code, and report that "the average variation between the best and worst framework is only 1.79 points" and that "the relative ordering of models remains unchanged across all three frameworks." Credit where it is due — that is a real answer. It is also narrower than it looks: that comparison reports average scores only, with no success rates, so the paper has bounded harness sensitivity for precisely the metric this article argues you should stop trusting, and left it unmeasured for the one it recommends. A position paper from May 2026, Stop Comparing LLM Agents Without Disclosing the Harness, argues that for long-horizon tasks the harness "is often a stronger determinant of agent performance than the model it wraps," and documents model ranking reversals from harness changes alone. Your stack is still not one of these three. If you have not read what an agent harness actually is, start there — the absolute rankings here should not be lifted into a vendor decision.
The ≥90 bar is the authors' choice. It is strict by construction. Satisfying every core requirement but only half the rest scores 79.55 and fails. Reasonable people would set that threshold elsewhere. The finding survives anyway, because the argument is about the shape of the metric, not the location of the line: any threshold recovers a spread the average hides.
The two metrics mostly agree on the ranking. Across the nine models, average score and success rate are almost perfectly rank-correlated — only two pairs swap, and both are near-ties. Choosing a vendor off the average would usually have chosen the same one. The claim here is about magnitude, not order: the average puts nine models inside a 24-point band where they all look serviceable, while the acceptance rate spreads that same field from 6.53% to 31.27%. The average will tell you who ranked first. It will not tell you whether first place is usable.
Runs are pooled, not conditioned. StartupBench runs each task three times and reports success as "the fraction of runs that are successfully completed, pooled over the three runs" — so a task an agent passes once in three counts the same as one it passes every time. That understates the reliability problem rather than overstating it — τ-bench's repository tracks pass^1 through pass^4 precisely because performance decays with repetition, and a 2026 framework paper on long-horizon reliability found capability and reliability rankings "diverge substantially, with multi-rank inversions at long horizons".
Domain variance is large. Averaged across the nine models, domain scores run from 75.47 in Business & Management down to 54.48 in Finance. A finance POC and a slide-deck POC are not the same difficulty, and a cross-domain average tells you little about yours. This is consistent with earlier work finding that AI models produced zero client-ready outputs on an investment banking reality check.
It is nine models at one moment. Vendors ship silently; a scorecard built on a pinned benchmark result can be invalidated without notice, which is what happened when DeepSeek swapped a model and existing evals did not notice. And the broader adoption picture has been consistent for a while — the Stanford AI Index put agent success at 66% while 89% never reached production. The gap between "scored well" and "shipped" is not new. StartupBench just measured it inside a single rubric.
Rewrite Your Acceptance Criteria This Quarter
This Week
- Pull the last agent POC scorecard you approved. Find the line that scored whether the agent produced work your team could use, and check what share of the total weight it carried. If it is under half, the number you signed was mostly about something else.
- Write the core-requirement list for your top use case — the requirements whose violation makes the deliverable unusable, in the paper's sense. Cap it at ten. Everything else moves below the line and stops affecting the verdict.
- Restate the acceptance criterion as a count, not an average: N of N core requirements satisfied, per task. A count needs no weights, so nobody can argue about them after the results land.
This Month
- Re-run your shortlist on one fixed task set, in your harness, with your tools and your step budget, and log which requirements failed rather than the score. Ten real tasks from your backlog beat a hundred synthetic ones.
- Run each task at least three times and report the worst run, not the mean. If the vendor objects to that, you have learned something.
- Have a domain expert grade the artifacts — the underwriter, the controller, the associate — not the platform team. StartupBench earned its 92.78% agreement by recruiting over 50 domain experts to write the rubrics in the first place.
- Grade the deliverable, never the agent's own summary of it. Open the file.
Before You Sign
- Put the core-requirement list into the SOW as the acceptance test, with the pass threshold and the number of trials written down.
- Pin the model version and require notice on change, because a silent swap re-opens every result you just paid for.
- Ask for the vendor's harness configuration in writing — tools, context strategy, step limit, retry policy. If a harness can reverse a ranking, it belongs in the contract.
The Bottom Line
Every technology cycle produces a metric that is easy to compute, easy to present, and slightly wrong in a direction that favours buying. Function points did it. Lines of code did it. Model accuracy did it before anyone thought to ask about the confusion matrix. Weighted rubric averages are that metric for agents, and the tell is that they compress a 4.5x difference in usable output into a 1.5x difference in score.
Your users do not experience an average. They open a file and it is either something they can send or something they have to redo. Grade the pilot the way the work is actually consumed.
A 73 is not a B. It is two rewrites out of three.
Continue Reading
- Agent Orchestration Platforms: Score Exit, Not Features
- DeepSeek Swapped the Model. Your Eval Didn't Notice.
- Anthropic's Evals Maxed Out. Stop Inheriting Its Assurance.
- Alibaba's Agent Coded 16 Days. A Human Wrote 13 Commits.
- AI Models Fail Investment Banking Reality Check
- Stanford AI Index 2026: AI Agents Hit 66% Success Rate
- 40% of AI Agent Projects Will Be Canceled by 2027
