Following my post yesterday on the proposed $500 billion AI infrastructure financing — and the extraordinary capex being committed by the hyperscalers — I keep coming back to another question: what happens if a meaningful percentage of AI inference eventually moves off the cloud?
Today, much of the AI investment thesis assumes rapidly growing demand for centralized compute. More users, more models, more inference, more data centers.
But technology rarely stands still. Open-weight models are becoming increasingly capable. Smaller models are improving. Apple is pushing more AI capabilities onto its devices. And Nvidia and Microsoft are now working to bring powerful local AI to a new generation of PCs that can run models and agents without sending every request to a data center.
Which raises an interesting economic question. When enterprises and consumers eventually see the full bill for cloud-based AI inference, will they start asking what can be done locally instead?
There are obvious reasons they might: cost, latency, privacy, security, data sovereignty and availability.
I don't believe cloud AI goes away. Far from it. The largest models, training workloads and many complex applications will continue to require enormous centralized compute.
But what if inference becomes increasingly hybrid? Train centrally. Run the hardest problems in the cloud. But push routine inference to phones, PCs, enterprise servers and edge devices.
If that happens at scale, it raises an interesting question for the trillions being committed to AI infrastructure: how much of today's projected inference demand ultimately materializes in the cloud? Does edge inference cannibalize cloud inference — or expand the total AI market enough that everyone wins?
Given the amount of capital now being deployed, I think it's a question worth asking.