Google has unveiled an agentic expansion of its Gemini model ecosystem alongside a massive infrastructure plan projected to hit $180 billion to $190 billion in capital expenditure this year. CEO Sundar Pichai revealed that Google platforms now process over 3.2 quadrillion tokens per month, driven by rapid enterprise adoption and developer integration. Developers and enterprise users will gain immediate access to lower-latency models like Gemini 3.5 Flash alongside next-generation custom hardware built for large-scale training and inference.
Key Data Points from Google's Announcement
- 3.2 quadrillion tokens per month processed across Google platforms, up 7x from 480 trillion during the previous year.
- $180B to $190B projected capex in 2026, roughly six times its $31 billion expenditure in 2022.
- 8.5 million developers building applications on Google models monthly, with APIs processing 19 billion tokens per minute.
- 4x faster token output on Gemini 3.5 Flash compared to competing frontier models, at less than half the execution price.
- 900 million monthly active users on the standalone Gemini application, more than doubling over the last 12 months.
Custom Silicon Drives Distributed AI Training
To power its growing agentic ecosystem, Google introduced its eighth-generation Tensor Processing Unit array, introducing two specialized architectures: TPU 8t for training and TPU 8i for inference. The TPU 8t chip delivers nearly three times the raw compute capacity of its predecessor.
By deploying JAX and Pathways frameworks, Google can distribute training tasks across more than 1 million TPUs globally without being bound by individual data center hardware constraints. Both processors deliver a 2x improvement in performance-per-watt efficiency, helping offset high energy demands during continuous inference cycles.
Gemini 3.5 Flash and Omni Multimodal Models
Google officially introduced Gemini 3.5 Flash, designed specifically for agentic coding and real-world workflows. On standard industry benchmarks, 3.5 Flash outperforms the older Gemini 3.1 Pro model while running four times faster than competitive frontier options. Google noted that internal development teams doubled daily internal tool token usage every few weeks using 3.5 Flash inside its Antigravity development environment.
Simultaneously, Google unveiled Gemini Omni Flash, the first model in its new Omni family capable of generating native video outputs from multi-modal inputs. The system will roll out across Google Flow, YouTube Shorts, and enterprise APIs.
However, early adopters face immediate limitations: Gemini Omni Flash initially generates only video outputs, with native image and text outputs delayed for later releases. Furthermore, conversational tools like Docs Live and automated voice integration in Gmail require paid consumer subscriptions and remain restricted to gradual summer rollouts.