Your board saw the Airbnb numbers this week, and by the next planning cycle someone will ask why your engineering org isn't shipping 80% more. Airbnb put five AI productivity claims on the record at its Q2 2026 results. Four of them are activity counts that only Airbnb can define, verify, or dispute. One is a cost denominated in something the company actually sells. That distinction is the whole story, and it decides which number you should let yourself be measured on.
The five, from Airbnb's Q2 2026 release: time from concept to delivery "reduced by as much as 60%"; features and improvements shipped this year up "nearly 80%"; an AI assistant live in "more than 50 languages"; "nearly 45% of issues that begin with our AI assistant" resolved without a human agent; and customer-support-related cost per booking down "approximately 16% year-over-year." Add the Q1 2026 claim that "nearly 60 percent of the code our engineers produce is now coauthored with AI", and you have the full set your CEO has probably read.
The market liked it. ABNB rose 15.05% on August 7, on revenue of $3.6 billion (up 17%) and gross booking value of $27.2 billion (up 16%), with full-year adjusted EBITDA margin guidance raised to at least 35.5%.
What Separates an Activity Metric From a Unit Economic
An activity metric counts things your organization did. A unit economic divides a cost by a unit of business the market pays for. The first is defined entirely by the party reporting it; the second is constrained by a denominator someone outside the company can see.
"Features shipped, up nearly 80%" is an activity metric. Nobody outside Airbnb defines a feature. A copy change on the checkout page and a rebuilt payments flow can both be one. There is no auditor, no standard, and no way for a competitor — or a shareholder — to check it. The same is true of "concept to delivery down 60%," which is a stopwatch Airbnb starts and stops, on projects Airbnb selects.
"Customer support cost per booking, down about 16%" is different in kind. Bookings are a number Airbnb reports for other reasons and is legally exposed on. The cost sits inside a line item in the 10-Q. It is still management-defined at the numerator — more on that below — but it is the only one of the five where an outsider can watch the P&L and see whether the story holds.
This is not an Airbnb problem. It is the default failure of AI measurement across the industry, which is why 64% of Fortune 500 companies use AI coding agents while only a third measure return on them. Activity metrics are easy to produce and impossible to falsify. That is exactly why they proliferate.
The 60% Code Number Is a Press Distortion, and It Matters
Airbnb never said AI writes 60% of its code. It said the code is coauthored with AI. TechCrunch's May headline read "Airbnb says AI now writes 60% of its new code", and the body of that piece rendered it as "60% of the code its engineers produced in the quarter was written by AI." That is a different claim, and it is the version that reached your board.
"Coauthored" describes tool usage. If an engineer accepts a three-line completion inside a 200-line function, the file is coauthored. It tells you the assistant is installed and running. It tells you nothing about output, quality, or cost. Read that way, 60% is roughly what you would expect from a company that bought seats and told people to use them — and Airbnb's own framing, "roughly twice the industry average, by our estimate," is an estimate of a number nobody measures the same way twice.
You will see this pattern everywhere once you look for it. It is the same shape as Anthropic's claim that its engineers ship 8x more code, and Google's disclosure that an internal agent writes 30% of some code. None of these are lies. They are all unfalsifiable, and a number that cannot be wrong cannot be a target.
The trap for you is specific: an authorship percentage is trivially gameable. Mandate 60% and you will get 60%, because the metric is generated by the tool that benefits from it.
Airbnb's Own Engineering Blog Shows What a Real Number Looks Like
The most credible AI engineering result Airbnb has published did not appear in an earnings release. It appeared on Airbnb's engineering blog, describing an LLM-driven migration of nearly 3,500 Enzyme test files to React Testing Library. The team had estimated 1.5 years of engineering time to do it by hand. The pipeline finished it in six weeks — 75% of target files in the first four-hour automated run, 97% after tuning, and the last 3% by hand over another week.
Compare the two disclosures. The migration has a fixed, unambiguous scope (a file count, which — unlike a feature — cannot be redefined after the fact), a baseline set before the result was known (the 1.5-year hand estimate), and a reported failure rate (the 3% that needed humans). "Features shipped, up 80%" has none of those. One is an engineering result. The other is a communications artifact.
That is the shape of the number you should be asking your own teams for: a bounded scope, a baseline set before the work started, and an honest tail. It is also why the Alibaba agent that coded for 16 days was only assessable at all because there was a public commit trace to read.
The Evidence That Cuts the Other Way
State the strongest counter-case, because it is strong. Airbnb's margin expanded while it absorbed real AI cost, and that is hard to fake at scale. CFO Ellie Mertz told analysts the raised guidance "does assume a material increase in terms of the AI spend over the course of the year," and that "we are expanding margins while absorbing that increased cost." A company inventing its productivity story does not usually raise the margin it has to hit.
Now the other side. The best-known randomized trial on the question found the opposite of what developers report about themselves: in METR's study, 16 experienced open-source developers took 19% longer on 246 real issues when allowed to use AI tools — while predicting beforehand that AI would speed them up 24%, and still believing afterward that it had sped them up 20%. Self-reported velocity and measured velocity pointed in opposite directions, by roughly 40 percentage points. Carry two limits with it: the trial ran in early 2025 on that period's tools, and METR itself says the finding is not evidence that AI fails to speed up most developers. These were maintainers working in repositories they had known for years — the setting where an assistant has least to tell them.
And the code itself is trending badly. GitClear's analysis of 623 million changes from 2023 to 2026 found within-commit copy/paste rising from 9.4% in 2022 to 15.7% in the first half of 2026, duplicated code blocks up 81% since 2023, error-masking constructs up 47%, and two-week churn up 15% — while refactoring line moves collapsed from 21% of changed lines in 2022 to 3.8% year-to-date. Cloned lines now exceed refactored lines. Read it with the interest stated: GitClear sells the developer analytics that measure exactly these signals, and the work is correlational — it dates the deterioration to the AI era without isolating AI as its cause. Even so, shipping more, faster, while reuse and refactoring go to zero, is a description of debt accrual, not of leverage.
Google's 2025 DORA research, drawn from nearly 5,000 technology professionals, lands on the reconciling explanation: AI is an amplifier. Strong systems get faster; weak ones ship worse work sooner. Both Airbnb's result and METR's can be true, of different organizations. Which one describes yours is an empirical question you have not answered yet — and it is a large part of why 88% of AI agent pilots never reach production.
What the 10-Q Says That the Press Release Doesn't
Go to the filing and the picture gets more useful, in both directions. In Airbnb's Q2 2026 10-Q, operations and support expense was $361 million against $332 million a year earlier — up 8.7%, while nights and seats booked grew 10%. (Use bookings, not GBV: GBV's 16% is flattered by average daily rate rising 5% to $184, and bookings are the denominator Airbnb's own claim uses.) That is real operating leverage, and it moves in the direction the 16% claim implies — but note the size. On the filed line, cost per booking improved by roughly one percentage point, not sixteen. Airbnb's figure covers a support subset it defines; the line item that subset sits inside barely moved. Product development was $672 million against $610 million, up 10.2%, against 17% revenue growth. Stock-based compensation rose 14.9% to $487 million.
Read that carefully. It is not a repudiation. It is a much smaller number than the headline implies. An 80% increase in output against a 10% increase in product development spend should be visible somewhere as more than 50 basis points of guided margin. The support line is where the leverage actually shows up — and support is where the AI is doing the most mechanical, highest-volume work.
Two honest caveats, in Airbnb's favour and against it. The operations and support line also carries customer relations costs — refunds, credits and host protection payouts — so it is not a clean proxy for the support-cost claim. And the support function is largely not staffed by Airbnb at all: the FY2025 10-K discloses a global network of approximately 13,000 third-party workers supporting the majority of community support contacts, against 8,200 employees. A cost line built mostly on contracted labour can fall 16% because an AI assistant deflected the contacts, or because a vendor rate was renegotiated, or because the offshore mix shifted — and nothing filed separates those. The 16% is still management's own construction. It is just the only one of the five that a filed line item moves with.
One more thing the filings will not tell you. Mertz said Airbnb doesn't "need to grow our head count at levels that we did in the past because we're getting so much more output and speed from our existing workforce." Airbnb does not disclose headcount quarterly. The last filed figure is 8,200 employees at December 31, 2025, up 12.3% from 7,300 — and the third-party support network alongside it grew over the same stretch, from roughly 11,000 at the end of 2023 to about 13,000. "Roughly flat headcount" during the 80% period is a management characterization, not a filed number — and every version of this story that pairs "80% more features" with "same staff" is repeating a claim nobody has checked.
The Support Number Has Its Own Ceiling
The metric worth adopting still needs a quality guardrail, and the industry already ran this experiment. Klarna announced an AI assistant that "performed the work of 700 employees," later revised to "over 800 full-time roles" — then acknowledged that an overemphasis on cost-cutting had produced worse service and moved to a dual-track model with human agents back in the loop. The cost-per-contact chart looked excellent right up until it didn't.
Airbnb's "nearly 45% resolved without a human agent" is a containment rate, and containment without a satisfaction and reopen-rate constraint is not a result — it is a deflection number. As the eleven-stack voice agent benchmark showed, containment varies enormously by what you count as contained. If you adopt cost per contact as your metric, pair it with reopen rate and CSAT on escalated contacts in the same report, or you will optimize your way into Klarna's 2025.
Do This Before Your Next Planning Cycle
This Week: Write down, in one page, the metric you will be measured on next quarter and its denominator. If the denominator is engineering activity (files, PRs, features, story points, AI-authored percentage), it is not a business metric. Find the closest cost-per-transaction line your function owns — cost per ticket, per claim, per invoice processed, per release — and propose that one instead, before someone proposes 80% for you.
This Week: Get your current baseline on paper before you expand any AI tooling. A baseline established after rollout is worthless, and it is the single most common reason AI ROI cases collapse under audit. Pull the last four quarters of that cost line now.
This Month: Run one bounded, auditable pilot in the Airbnb test-migration shape — a fixed scope, a pre-registered hand estimate from the team that would have done it, and a reported failure rate. One migration, one back-office workflow, one test suite. The number it produces will be smaller than 80% and worth infinitely more, because you can defend it.
This Month: Add a quality counterweight to every efficiency metric you report upward. Cost per contact ships with reopen rate. Velocity ships with change failure rate and a duplication trend from your own repos — GitClear's signals are measurable on your codebase, not just in a research report.
Before Your Next Renewal: Reconcile seat spend against the metric. If you cannot connect Cursor or GitHub Copilot seats to a movement in a cost-per-unit line, you are renewing on an authorship percentage. Airbnb's own CFO tied AI spend to margin in public; you should be able to do the same in a budget review.
The Bottom Line
The last time an industry adopted a self-defined activity metric at this scale, it was lines of code, and it took a decade to unlearn. "Percentage of code coauthored with AI" is lines of code with better marketing — a number your vendor generates, your team can inflate without lying, and nobody outside the building can dispute. Airbnb's genuine achievement this quarter was not shipping 80% more features. It was putting a cost-per-booking figure next to a margin guide and letting both sit in the same filing.
Boards will quote the 80% at you anyway. Your job is to arrive at that meeting with a number of your own that has a denominator someone else can check.
Measure what the business buys, not what the tool produced.
Continue Reading
64% of Fortune 500 Use AI Coding Agents. 33% Measure ROI. Alibaba's Agent Coded 16 Days. A Human Wrote 13 Commits. Anthropic Engineers Ship 8x More Code: 80% AI-Written by Mid-2026 92% Trust AI Code Scanning. 70% Have Vulns in Production. Claude in Enterprise: 40% Faster, 8 Hours Saved a Week Group 1 Cut 700 Jobs. Better AI Wasn't the Reason. Why 88% of AI Agent Pilots Never Reach Production Eleven Voice Agents, One Bank Call, No Clean Winner
