Google Cloud 82% Growth Hides a Serious Supply Problem

Google Cloud revenue surged 82% in Q2 2026 to $24.8B. But the CFO's warning: demand still outpaces supply. What this means for your enterprise AI strategy.

By Rajesh Beri·July 26, 2026·10 min read
Share:
THE DAILY BRIEF
Enterprise AICloud ComputingGoogle CloudAI InfrastructureDigital Transformation
Google Cloud 82% Growth Hides a Serious Supply Problem

Google Cloud revenue surged 82% in Q2 2026 to $24.8B. But the CFO's warning: demand still outpaces supply. What this means for your enterprise AI strategy.

By Rajesh Beri·July 26, 2026·10 min read

The headline was 82% revenue growth. The real story was buried three slides into the earnings deck. Google cloud just posted $24.8 billion in Q2 2026 revenue — the fastest growth of any major hyperscaler this quarter. But CFO Anat Ashkenazi said something that every CIO, CTO, and CFO planning an AI buildout needs to hear: "We're still in a supply-constrained environment. While we have increased our capacity quite significantly over the past three years, the demand still outpaces that investment."

Alphabet is raising its 2026 capital expenditure guidance to $195–$205 billion — up from a previous estimate of $180–$190 billion — and expects spending to continue increasing significantly into 2027. They're spending more money than most countries' GDP on infrastructure. And it's still not enough.

That should change how you think about your enterprise AI timeline.

The Numbers Behind the Numbers

Cloud operating income more than tripled to $8.8 billion. Cloud backlog — the value of contracted work yet to be delivered — reached $514 billion. Nearly 90% of Fortune 100 companies are actively using Gemini Enterprise. And Alphabet's Gemini API is now processing approximately 22 billion tokens per minute, up from 16 billion just a quarter ago.

These are extraordinary metrics. But they tell two very different stories depending on which seat you're sitting in.

If you're a Google shareholder, this looks like a generational business. If you're an enterprise leader trying to scale AI workloads over the next 12–18 months, the more relevant data point is that $514 billion backlog. That's a line of enterprise customers who have committed to Google Cloud — and are waiting for capacity that doesn't fully exist yet.

Sundar Pichai told analysts that most enterprises are "barely scratching the early stages of what's possible here." Coming from a CEO on an earnings call, that's typically sales language. But in this case, the financial data backs it up. If nearly 90% of Fortune 100 companies are using Gemini Enterprise but most are still at early-stage usage, the demand growth curve is still ahead of us — not behind.

The Infrastructure Math You Need to Understand

Google spent $44.9 billion on capital expenditures in Q2 2026 alone. Annualized, that's close to $180 billion per year. And they're raising the full-year target to $205 billion — because even that level of investment isn't meeting current demand.

Synergy Research Group reported this week that total U.S. data center capacity is expected to double over the next three years. Hyperscaler-owned data center capacity specifically — the compute run by Google, Microsoft, and AWS — is expected to double in the next two years. That's an aggressive buildout by any measure.

But here's what doesn't get discussed enough: capacity buildouts at this scale face real-world constraints that capital alone can't solve. Power availability. Land permitting. Cooling infrastructure. Fiber routes. Physical construction timelines. John Dinsdale of Synergy Research was direct about it: "It is indisputable that constrained availability of power and rising local concerns over data centers are crimping many new plans for data centers."

For enterprise leaders, this creates a strategic problem that has nothing to do with your AI strategy or your vendor selection. You could pick the right model, build the right use case, and win executive buy-in — and still face capacity delays that push your timeline.

What agentic AI Is Doing to Demand

The token consumption data from this earnings call deserves more attention than it's getting. More than 2,000 enterprises consumed over 100 billion tokens in the last year. Nearly 500 Google Cloud customers processed over 1 trillion tokens each.

A trillion tokens per enterprise customer per year is a materially different scale than what most AI planning models assumed two years ago. Traditional AI integration — a chatbot here, a summarization tool there — doesn't generate this level of token consumption. Agentic AI does.

As enterprises move from point AI tools to multi-step AI agents that reason, plan, and execute across workflows, token consumption scales dramatically. A customer service agent that handles a single interaction might consume 500–2,000 tokens. An enterprise AI agent managing a procurement workflow — pulling supplier data, analyzing contracts, generating recommendations, and triggering approvals — might consume 50,000–200,000 tokens per workflow run.

Multiply that by tens of thousands of workflow executions per day, and you start to understand why Google's API is processing 22 billion tokens per minute and why that number jumped 37% in a single quarter.

The implication for enterprise planning: token consumption is no longer a developer concern. It's a budget line item that finance leaders need to understand and model. Alphabet CFO Ashkenazi explicitly flagged that greater AI use is translating to more token consumption, and that this is "starting to become a cost concern for enterprises."

The Technical Leader's View: What This Means for Your Architecture

For CIOs and CTOs, the Google Q2 data reveals three things that should influence how you architect your AI infrastructure strategy.

First, on-premises compute is back on the table. Google began deploying its Tensor Processing Unit (TPU) systems to customer data centers for the first time this quarter. That's a notable shift. For years, the hyperscaler model assumed that enterprises would move workloads to the cloud. The fact that Google is now bringing its custom silicon to enterprise data centers reflects two realities: some enterprises have data residency requirements that prevent full cloud migration, and some workloads are now generating enough compute demand that on-premises deployment is economically competitive.

If you have workloads generating high, predictable token volumes — internal AI agents, document processing pipelines, coding assistants used by large engineering teams — the ROI calculation for on-premises TPU deployment deserves a fresh look.

Second, multi-cloud positioning is no longer just a negotiating tactic. When any single cloud provider is running in a supply-constrained environment, your ability to shift workloads across providers becomes a risk management tool, not just a cost management strategy. The enterprises most exposed to capacity delays are those who've concentrated 80%+ of their AI workloads on a single provider.

Third, token cost optimization needs to be a first-class engineering discipline. When I talk to engineering leaders who've deployed AI agents at scale, the ones who avoided bill shock had one thing in common: they treated token optimization the same way they'd treat memory optimization in performance-critical systems. Caching intermediate results. Batching requests. Choosing the right model tier for the task complexity. Designing agent architectures that minimize unnecessary context accumulation. These aren't afterthoughts — they're architectural decisions made during design.

The Business Leader's View: What This Means for Your AI Investment

For CFOs, CMOs, COOs, and other business leaders evaluating AI investments, the Alphabet earnings data provides useful benchmarks for calibrating your expectations.

The ROI signals are real. Google Cloud's operating income more than tripling in a single quarter isn't a projection — it's a revealed outcome. Google is earning $8.8 billion in operating income from a business that generates $24.8 billion in revenue. That's a 35% operating margin on a cloud platform that's still in aggressive growth mode. Customers are paying a premium for AI-enhanced cloud services, and they're doing so because the business value justifies it.

The timing risk is underappreciated. Conversations with operations and finance leaders who are 12–18 months into AI deployment consistently surface the same issue: internal demand for AI capabilities scales faster than procurement and infrastructure can support. A successful pilot leads to department-wide adoption requests. Department-wide adoption leads to enterprise-wide rollout pressure. And enterprise-wide rollout runs into capacity limits — both in cloud infrastructure and in internal IT bandwidth.

The $514 billion Google Cloud backlog is a proxy for that pattern playing out at enterprise scale. Companies have contracted for capacity they need. Delivery takes time. If you're at the beginning of that adoption curve, building longer timelines into your AI roadmap isn't pessimism — it's operational realism.

The token economy needs to be in your financial models. If your enterprise is planning agentic AI deployments, your finance team needs a token consumption model before procurement, not after. Enterprises that signed cloud AI contracts without modeling token growth are experiencing meaningful cost overruns relative to initial projections. The unit economics of AI at agent-scale look very different from the unit economics of AI at chatbot-scale.

Work backward from your use cases. Estimate the number of agent runs per day. Model the average token consumption per run under different complexity scenarios. Apply a conservative buffer — agentic systems consistently consume more tokens than initial estimates suggest. That model should inform both your vendor contract structure and your internal chargeback framework.

The Strategic Play: Timing Is Everything

Here's the contrarian take that the 82% growth headline obscures: a supply-constrained AI infrastructure market is a strategic advantage for companies that plan ahead and a significant disadvantage for companies that wait.

Enterprises that committed to Google Cloud AI capacity 12–18 months ago secured favorable pricing and guaranteed availability. The $514 billion backlog suggests that window is narrowing. Enterprises entering that backlog today are doing so at higher price points and with longer delivery timelines.

This dynamic is consistent across all three major hyperscalers. Microsoft is in a similar capacity-constrained position with Azure. AWS has consistently cited AI compute demand as exceeding available capacity. The difference is that Google's earnings data — especially the 82% revenue growth and the explicit supply constraint admission — makes the dynamic visible.

For enterprises that are still in the evaluation phase: the cost of waiting isn't zero. Every quarter you delay committing to an AI infrastructure strategy is a quarter in which capacity costs more and takes longer to secure.

For enterprises that have already committed: the token consumption trajectory from Google's data suggests you should model 30–50% higher token consumption than your current projections assume, especially if you're planning agentic deployments over the next 12 months.

The Bottom Line

Google Cloud's 82% revenue growth is the attention-grabbing number. The more important number for enterprise leaders is $514 billion — the volume of contracted cloud AI work that hasn't been delivered yet. That backlog, combined with CFO Ashkenazi's frank admission that demand still outpaces supply, tells you everything you need to know about where the AI infrastructure market is right now.

Enterprise AI is not slowing down. It's accelerating. And the constraint on that acceleration isn't ambition, budget, or executive buy-in — it's physical infrastructure that takes years to build.

Three takeaways for enterprise leaders acting on this data:

  1. Lock in capacity commitments now. If you're planning significant AI workloads for H2 2026 or 2027, the time to negotiate cloud AI contracts is before the backlog grows further, not after.

  2. Build token cost modeling into your AI business cases. The enterprises running 1 trillion tokens per year are not outliers — they're the early adopters of agentic AI. Your organization will follow. Finance leaders who understand the token economics now will have significantly more control over costs when that adoption curve hits.

  3. Evaluate on-premises options for high-volume, predictable workloads. Google's decision to deploy TPUs directly to enterprise data centers signals a market shift. For workloads with consistent, high-volume token consumption, the economics of on-premises AI compute are improving faster than most infrastructure teams have modeled.

The 82% growth headline is impressive. But the supply constraint underneath it is what will define which enterprises win the AI buildout and which ones spend the next 18 months in queue.


Enterprise AI moves fast. THE DAILY BRIEF covers what matters for technical and business leaders — twice a week, no noise. Follow on X/Twitter or connect on LinkedIn.

Continue Reading

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, twice weekly.

beri.net

Subscribe at beri.net/subscribe for twice-weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Google Cloud 82% Growth Hides a Serious Supply Problem

Photo by Brett Sayles on Pexels

The headline was 82% revenue growth. The real story was buried three slides into the earnings deck. Google cloud just posted $24.8 billion in Q2 2026 revenue — the fastest growth of any major hyperscaler this quarter. But CFO Anat Ashkenazi said something that every CIO, CTO, and CFO planning an AI buildout needs to hear: "We're still in a supply-constrained environment. While we have increased our capacity quite significantly over the past three years, the demand still outpaces that investment."

Alphabet is raising its 2026 capital expenditure guidance to $195–$205 billion — up from a previous estimate of $180–$190 billion — and expects spending to continue increasing significantly into 2027. They're spending more money than most countries' GDP on infrastructure. And it's still not enough.

That should change how you think about your enterprise AI timeline.

The Numbers Behind the Numbers

Cloud operating income more than tripled to $8.8 billion. Cloud backlog — the value of contracted work yet to be delivered — reached $514 billion. Nearly 90% of Fortune 100 companies are actively using Gemini Enterprise. And Alphabet's Gemini API is now processing approximately 22 billion tokens per minute, up from 16 billion just a quarter ago.

These are extraordinary metrics. But they tell two very different stories depending on which seat you're sitting in.

If you're a Google shareholder, this looks like a generational business. If you're an enterprise leader trying to scale AI workloads over the next 12–18 months, the more relevant data point is that $514 billion backlog. That's a line of enterprise customers who have committed to Google Cloud — and are waiting for capacity that doesn't fully exist yet.

Sundar Pichai told analysts that most enterprises are "barely scratching the early stages of what's possible here." Coming from a CEO on an earnings call, that's typically sales language. But in this case, the financial data backs it up. If nearly 90% of Fortune 100 companies are using Gemini Enterprise but most are still at early-stage usage, the demand growth curve is still ahead of us — not behind.

The Infrastructure Math You Need to Understand

Google spent $44.9 billion on capital expenditures in Q2 2026 alone. Annualized, that's close to $180 billion per year. And they're raising the full-year target to $205 billion — because even that level of investment isn't meeting current demand.

Synergy Research Group reported this week that total U.S. data center capacity is expected to double over the next three years. Hyperscaler-owned data center capacity specifically — the compute run by Google, Microsoft, and AWS — is expected to double in the next two years. That's an aggressive buildout by any measure.

But here's what doesn't get discussed enough: capacity buildouts at this scale face real-world constraints that capital alone can't solve. Power availability. Land permitting. Cooling infrastructure. Fiber routes. Physical construction timelines. John Dinsdale of Synergy Research was direct about it: "It is indisputable that constrained availability of power and rising local concerns over data centers are crimping many new plans for data centers."

For enterprise leaders, this creates a strategic problem that has nothing to do with your AI strategy or your vendor selection. You could pick the right model, build the right use case, and win executive buy-in — and still face capacity delays that push your timeline.

What agentic AI Is Doing to Demand

The token consumption data from this earnings call deserves more attention than it's getting. More than 2,000 enterprises consumed over 100 billion tokens in the last year. Nearly 500 Google Cloud customers processed over 1 trillion tokens each.

A trillion tokens per enterprise customer per year is a materially different scale than what most AI planning models assumed two years ago. Traditional AI integration — a chatbot here, a summarization tool there — doesn't generate this level of token consumption. Agentic AI does.

As enterprises move from point AI tools to multi-step AI agents that reason, plan, and execute across workflows, token consumption scales dramatically. A customer service agent that handles a single interaction might consume 500–2,000 tokens. An enterprise AI agent managing a procurement workflow — pulling supplier data, analyzing contracts, generating recommendations, and triggering approvals — might consume 50,000–200,000 tokens per workflow run.

Multiply that by tens of thousands of workflow executions per day, and you start to understand why Google's API is processing 22 billion tokens per minute and why that number jumped 37% in a single quarter.

The implication for enterprise planning: token consumption is no longer a developer concern. It's a budget line item that finance leaders need to understand and model. Alphabet CFO Ashkenazi explicitly flagged that greater AI use is translating to more token consumption, and that this is "starting to become a cost concern for enterprises."

The Technical Leader's View: What This Means for Your Architecture

For CIOs and CTOs, the Google Q2 data reveals three things that should influence how you architect your AI infrastructure strategy.

First, on-premises compute is back on the table. Google began deploying its Tensor Processing Unit (TPU) systems to customer data centers for the first time this quarter. That's a notable shift. For years, the hyperscaler model assumed that enterprises would move workloads to the cloud. The fact that Google is now bringing its custom silicon to enterprise data centers reflects two realities: some enterprises have data residency requirements that prevent full cloud migration, and some workloads are now generating enough compute demand that on-premises deployment is economically competitive.

If you have workloads generating high, predictable token volumes — internal AI agents, document processing pipelines, coding assistants used by large engineering teams — the ROI calculation for on-premises TPU deployment deserves a fresh look.

Second, multi-cloud positioning is no longer just a negotiating tactic. When any single cloud provider is running in a supply-constrained environment, your ability to shift workloads across providers becomes a risk management tool, not just a cost management strategy. The enterprises most exposed to capacity delays are those who've concentrated 80%+ of their AI workloads on a single provider.

Third, token cost optimization needs to be a first-class engineering discipline. When I talk to engineering leaders who've deployed AI agents at scale, the ones who avoided bill shock had one thing in common: they treated token optimization the same way they'd treat memory optimization in performance-critical systems. Caching intermediate results. Batching requests. Choosing the right model tier for the task complexity. Designing agent architectures that minimize unnecessary context accumulation. These aren't afterthoughts — they're architectural decisions made during design.

The Business Leader's View: What This Means for Your AI Investment

For CFOs, CMOs, COOs, and other business leaders evaluating AI investments, the Alphabet earnings data provides useful benchmarks for calibrating your expectations.

The ROI signals are real. Google Cloud's operating income more than tripling in a single quarter isn't a projection — it's a revealed outcome. Google is earning $8.8 billion in operating income from a business that generates $24.8 billion in revenue. That's a 35% operating margin on a cloud platform that's still in aggressive growth mode. Customers are paying a premium for AI-enhanced cloud services, and they're doing so because the business value justifies it.

The timing risk is underappreciated. Conversations with operations and finance leaders who are 12–18 months into AI deployment consistently surface the same issue: internal demand for AI capabilities scales faster than procurement and infrastructure can support. A successful pilot leads to department-wide adoption requests. Department-wide adoption leads to enterprise-wide rollout pressure. And enterprise-wide rollout runs into capacity limits — both in cloud infrastructure and in internal IT bandwidth.

The $514 billion Google Cloud backlog is a proxy for that pattern playing out at enterprise scale. Companies have contracted for capacity they need. Delivery takes time. If you're at the beginning of that adoption curve, building longer timelines into your AI roadmap isn't pessimism — it's operational realism.

The token economy needs to be in your financial models. If your enterprise is planning agentic AI deployments, your finance team needs a token consumption model before procurement, not after. Enterprises that signed cloud AI contracts without modeling token growth are experiencing meaningful cost overruns relative to initial projections. The unit economics of AI at agent-scale look very different from the unit economics of AI at chatbot-scale.

Work backward from your use cases. Estimate the number of agent runs per day. Model the average token consumption per run under different complexity scenarios. Apply a conservative buffer — agentic systems consistently consume more tokens than initial estimates suggest. That model should inform both your vendor contract structure and your internal chargeback framework.

The Strategic Play: Timing Is Everything

Here's the contrarian take that the 82% growth headline obscures: a supply-constrained AI infrastructure market is a strategic advantage for companies that plan ahead and a significant disadvantage for companies that wait.

Enterprises that committed to Google Cloud AI capacity 12–18 months ago secured favorable pricing and guaranteed availability. The $514 billion backlog suggests that window is narrowing. Enterprises entering that backlog today are doing so at higher price points and with longer delivery timelines.

This dynamic is consistent across all three major hyperscalers. Microsoft is in a similar capacity-constrained position with Azure. AWS has consistently cited AI compute demand as exceeding available capacity. The difference is that Google's earnings data — especially the 82% revenue growth and the explicit supply constraint admission — makes the dynamic visible.

For enterprises that are still in the evaluation phase: the cost of waiting isn't zero. Every quarter you delay committing to an AI infrastructure strategy is a quarter in which capacity costs more and takes longer to secure.

For enterprises that have already committed: the token consumption trajectory from Google's data suggests you should model 30–50% higher token consumption than your current projections assume, especially if you're planning agentic deployments over the next 12 months.

The Bottom Line

Google Cloud's 82% revenue growth is the attention-grabbing number. The more important number for enterprise leaders is $514 billion — the volume of contracted cloud AI work that hasn't been delivered yet. That backlog, combined with CFO Ashkenazi's frank admission that demand still outpaces supply, tells you everything you need to know about where the AI infrastructure market is right now.

Enterprise AI is not slowing down. It's accelerating. And the constraint on that acceleration isn't ambition, budget, or executive buy-in — it's physical infrastructure that takes years to build.

Three takeaways for enterprise leaders acting on this data:

  1. Lock in capacity commitments now. If you're planning significant AI workloads for H2 2026 or 2027, the time to negotiate cloud AI contracts is before the backlog grows further, not after.

  2. Build token cost modeling into your AI business cases. The enterprises running 1 trillion tokens per year are not outliers — they're the early adopters of agentic AI. Your organization will follow. Finance leaders who understand the token economics now will have significantly more control over costs when that adoption curve hits.

  3. Evaluate on-premises options for high-volume, predictable workloads. Google's decision to deploy TPUs directly to enterprise data centers signals a market shift. For workloads with consistent, high-volume token consumption, the economics of on-premises AI compute are improving faster than most infrastructure teams have modeled.

The 82% growth headline is impressive. But the supply constraint underneath it is what will define which enterprises win the AI buildout and which ones spend the next 18 months in queue.


Enterprise AI moves fast. THE DAILY BRIEF covers what matters for technical and business leaders — twice a week, no noise. Follow on X/Twitter or connect on LinkedIn.

Continue Reading

Share:
THE DAILY BRIEF
Enterprise AICloud ComputingGoogle CloudAI InfrastructureDigital Transformation
Google Cloud 82% Growth Hides a Serious Supply Problem

Google Cloud revenue surged 82% in Q2 2026 to $24.8B. But the CFO's warning: demand still outpaces supply. What this means for your enterprise AI strategy.

By Rajesh Beri·July 26, 2026·10 min read

The headline was 82% revenue growth. The real story was buried three slides into the earnings deck. Google cloud just posted $24.8 billion in Q2 2026 revenue — the fastest growth of any major hyperscaler this quarter. But CFO Anat Ashkenazi said something that every CIO, CTO, and CFO planning an AI buildout needs to hear: "We're still in a supply-constrained environment. While we have increased our capacity quite significantly over the past three years, the demand still outpaces that investment."

Alphabet is raising its 2026 capital expenditure guidance to $195–$205 billion — up from a previous estimate of $180–$190 billion — and expects spending to continue increasing significantly into 2027. They're spending more money than most countries' GDP on infrastructure. And it's still not enough.

That should change how you think about your enterprise AI timeline.

The Numbers Behind the Numbers

Cloud operating income more than tripled to $8.8 billion. Cloud backlog — the value of contracted work yet to be delivered — reached $514 billion. Nearly 90% of Fortune 100 companies are actively using Gemini Enterprise. And Alphabet's Gemini API is now processing approximately 22 billion tokens per minute, up from 16 billion just a quarter ago.

These are extraordinary metrics. But they tell two very different stories depending on which seat you're sitting in.

If you're a Google shareholder, this looks like a generational business. If you're an enterprise leader trying to scale AI workloads over the next 12–18 months, the more relevant data point is that $514 billion backlog. That's a line of enterprise customers who have committed to Google Cloud — and are waiting for capacity that doesn't fully exist yet.

Sundar Pichai told analysts that most enterprises are "barely scratching the early stages of what's possible here." Coming from a CEO on an earnings call, that's typically sales language. But in this case, the financial data backs it up. If nearly 90% of Fortune 100 companies are using Gemini Enterprise but most are still at early-stage usage, the demand growth curve is still ahead of us — not behind.

The Infrastructure Math You Need to Understand

Google spent $44.9 billion on capital expenditures in Q2 2026 alone. Annualized, that's close to $180 billion per year. And they're raising the full-year target to $205 billion — because even that level of investment isn't meeting current demand.

Synergy Research Group reported this week that total U.S. data center capacity is expected to double over the next three years. Hyperscaler-owned data center capacity specifically — the compute run by Google, Microsoft, and AWS — is expected to double in the next two years. That's an aggressive buildout by any measure.

But here's what doesn't get discussed enough: capacity buildouts at this scale face real-world constraints that capital alone can't solve. Power availability. Land permitting. Cooling infrastructure. Fiber routes. Physical construction timelines. John Dinsdale of Synergy Research was direct about it: "It is indisputable that constrained availability of power and rising local concerns over data centers are crimping many new plans for data centers."

For enterprise leaders, this creates a strategic problem that has nothing to do with your AI strategy or your vendor selection. You could pick the right model, build the right use case, and win executive buy-in — and still face capacity delays that push your timeline.

What agentic AI Is Doing to Demand

The token consumption data from this earnings call deserves more attention than it's getting. More than 2,000 enterprises consumed over 100 billion tokens in the last year. Nearly 500 Google Cloud customers processed over 1 trillion tokens each.

A trillion tokens per enterprise customer per year is a materially different scale than what most AI planning models assumed two years ago. Traditional AI integration — a chatbot here, a summarization tool there — doesn't generate this level of token consumption. Agentic AI does.

As enterprises move from point AI tools to multi-step AI agents that reason, plan, and execute across workflows, token consumption scales dramatically. A customer service agent that handles a single interaction might consume 500–2,000 tokens. An enterprise AI agent managing a procurement workflow — pulling supplier data, analyzing contracts, generating recommendations, and triggering approvals — might consume 50,000–200,000 tokens per workflow run.

Multiply that by tens of thousands of workflow executions per day, and you start to understand why Google's API is processing 22 billion tokens per minute and why that number jumped 37% in a single quarter.

The implication for enterprise planning: token consumption is no longer a developer concern. It's a budget line item that finance leaders need to understand and model. Alphabet CFO Ashkenazi explicitly flagged that greater AI use is translating to more token consumption, and that this is "starting to become a cost concern for enterprises."

The Technical Leader's View: What This Means for Your Architecture

For CIOs and CTOs, the Google Q2 data reveals three things that should influence how you architect your AI infrastructure strategy.

First, on-premises compute is back on the table. Google began deploying its Tensor Processing Unit (TPU) systems to customer data centers for the first time this quarter. That's a notable shift. For years, the hyperscaler model assumed that enterprises would move workloads to the cloud. The fact that Google is now bringing its custom silicon to enterprise data centers reflects two realities: some enterprises have data residency requirements that prevent full cloud migration, and some workloads are now generating enough compute demand that on-premises deployment is economically competitive.

If you have workloads generating high, predictable token volumes — internal AI agents, document processing pipelines, coding assistants used by large engineering teams — the ROI calculation for on-premises TPU deployment deserves a fresh look.

Second, multi-cloud positioning is no longer just a negotiating tactic. When any single cloud provider is running in a supply-constrained environment, your ability to shift workloads across providers becomes a risk management tool, not just a cost management strategy. The enterprises most exposed to capacity delays are those who've concentrated 80%+ of their AI workloads on a single provider.

Third, token cost optimization needs to be a first-class engineering discipline. When I talk to engineering leaders who've deployed AI agents at scale, the ones who avoided bill shock had one thing in common: they treated token optimization the same way they'd treat memory optimization in performance-critical systems. Caching intermediate results. Batching requests. Choosing the right model tier for the task complexity. Designing agent architectures that minimize unnecessary context accumulation. These aren't afterthoughts — they're architectural decisions made during design.

The Business Leader's View: What This Means for Your AI Investment

For CFOs, CMOs, COOs, and other business leaders evaluating AI investments, the Alphabet earnings data provides useful benchmarks for calibrating your expectations.

The ROI signals are real. Google Cloud's operating income more than tripling in a single quarter isn't a projection — it's a revealed outcome. Google is earning $8.8 billion in operating income from a business that generates $24.8 billion in revenue. That's a 35% operating margin on a cloud platform that's still in aggressive growth mode. Customers are paying a premium for AI-enhanced cloud services, and they're doing so because the business value justifies it.

The timing risk is underappreciated. Conversations with operations and finance leaders who are 12–18 months into AI deployment consistently surface the same issue: internal demand for AI capabilities scales faster than procurement and infrastructure can support. A successful pilot leads to department-wide adoption requests. Department-wide adoption leads to enterprise-wide rollout pressure. And enterprise-wide rollout runs into capacity limits — both in cloud infrastructure and in internal IT bandwidth.

The $514 billion Google Cloud backlog is a proxy for that pattern playing out at enterprise scale. Companies have contracted for capacity they need. Delivery takes time. If you're at the beginning of that adoption curve, building longer timelines into your AI roadmap isn't pessimism — it's operational realism.

The token economy needs to be in your financial models. If your enterprise is planning agentic AI deployments, your finance team needs a token consumption model before procurement, not after. Enterprises that signed cloud AI contracts without modeling token growth are experiencing meaningful cost overruns relative to initial projections. The unit economics of AI at agent-scale look very different from the unit economics of AI at chatbot-scale.

Work backward from your use cases. Estimate the number of agent runs per day. Model the average token consumption per run under different complexity scenarios. Apply a conservative buffer — agentic systems consistently consume more tokens than initial estimates suggest. That model should inform both your vendor contract structure and your internal chargeback framework.

The Strategic Play: Timing Is Everything

Here's the contrarian take that the 82% growth headline obscures: a supply-constrained AI infrastructure market is a strategic advantage for companies that plan ahead and a significant disadvantage for companies that wait.

Enterprises that committed to Google Cloud AI capacity 12–18 months ago secured favorable pricing and guaranteed availability. The $514 billion backlog suggests that window is narrowing. Enterprises entering that backlog today are doing so at higher price points and with longer delivery timelines.

This dynamic is consistent across all three major hyperscalers. Microsoft is in a similar capacity-constrained position with Azure. AWS has consistently cited AI compute demand as exceeding available capacity. The difference is that Google's earnings data — especially the 82% revenue growth and the explicit supply constraint admission — makes the dynamic visible.

For enterprises that are still in the evaluation phase: the cost of waiting isn't zero. Every quarter you delay committing to an AI infrastructure strategy is a quarter in which capacity costs more and takes longer to secure.

For enterprises that have already committed: the token consumption trajectory from Google's data suggests you should model 30–50% higher token consumption than your current projections assume, especially if you're planning agentic deployments over the next 12 months.

The Bottom Line

Google Cloud's 82% revenue growth is the attention-grabbing number. The more important number for enterprise leaders is $514 billion — the volume of contracted cloud AI work that hasn't been delivered yet. That backlog, combined with CFO Ashkenazi's frank admission that demand still outpaces supply, tells you everything you need to know about where the AI infrastructure market is right now.

Enterprise AI is not slowing down. It's accelerating. And the constraint on that acceleration isn't ambition, budget, or executive buy-in — it's physical infrastructure that takes years to build.

Three takeaways for enterprise leaders acting on this data:

  1. Lock in capacity commitments now. If you're planning significant AI workloads for H2 2026 or 2027, the time to negotiate cloud AI contracts is before the backlog grows further, not after.

  2. Build token cost modeling into your AI business cases. The enterprises running 1 trillion tokens per year are not outliers — they're the early adopters of agentic AI. Your organization will follow. Finance leaders who understand the token economics now will have significantly more control over costs when that adoption curve hits.

  3. Evaluate on-premises options for high-volume, predictable workloads. Google's decision to deploy TPUs directly to enterprise data centers signals a market shift. For workloads with consistent, high-volume token consumption, the economics of on-premises AI compute are improving faster than most infrastructure teams have modeled.

The 82% growth headline is impressive. But the supply constraint underneath it is what will define which enterprises win the AI buildout and which ones spend the next 18 months in queue.


Enterprise AI moves fast. THE DAILY BRIEF covers what matters for technical and business leaders — twice a week, no noise. Follow on X/Twitter or connect on LinkedIn.

Continue Reading

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, twice weekly.

beri.net

Subscribe at beri.net/subscribe for twice-weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Frequently Asked Questions

How much did Google Cloud revenue grow in Q2 2026?

Google Cloud revenue grew 82% year over year to $24.8 billion in Q2 2026, the fastest growth of any major hyperscaler that quarter and well ahead of the ~$22.4 billion analysts expected. Cloud operating income more than tripled to $8.8 billion, lifting the segment's operating margin to 35.6% from 20.7% a year earlier.

What is Google Cloud's $514 billion backlog and why does it matter to enterprises?

Backlog is the value of contracted cloud work Google has signed but not yet delivered. It reached $514 billion in Q2 2026, up more than $50 billion sequentially, and Alphabet expects to recognize slightly more than half of it as revenue over the next 24 months. For enterprise buyers it is a queue: capacity is already spoken for, so new commitments tend to come at higher price points and longer delivery timelines.

How much is Alphabet spending on AI infrastructure in 2026?

Alphabet raised full-year 2026 capital expenditure guidance to $195-$205 billion, up from a prior $180-$190 billion, and spent $44.9 billion in Q2 2026 alone. CFO Anat Ashkenazi told analysts the company is 'still in a supply-constrained environment' despite that spending, and expects capex to keep rising into 2027.

Why is agentic AI driving so much token consumption?

Agentic workflows chain many model calls instead of one. A single chatbot exchange may use 500-2,000 tokens, while an AI agent running a procurement workflow — pulling supplier data, analyzing contracts, generating recommendations, triggering approvals — can consume 50,000-200,000 tokens per run. That is why Google's model APIs now process about 22 billion tokens per minute, up from 16 billion a quarter earlier, and why nearly 500 Cloud customers each processed over 1 trillion tokens in the past year.

What should enterprise leaders do about AI capacity constraints?

Three moves: lock in cloud capacity commitments before the backlog grows further; build token-consumption modeling into AI business cases before procurement rather than after; and re-evaluate on-premises options for high-volume, predictable workloads — Google began shipping TPU systems into customer data centers for the first time in Q2 2026. Multi-cloud positioning also shifts from a pricing tactic to a capacity-risk hedge.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →