Earlier this month, I attended the KeyBanc Technology Leadership Forum.
What I love about this event is that it provides a broad perspective across the entire technology landscape.
One topic dominated this year’s agenda: data centers.
The scale of the buildout is extraordinary. The 14 largest publicly traded data-center operators are expected to spend nearly $750B in CapEx in 2026. But the buildout is increasingly constrained by five forces: the 5 Ps:
Planet: climate and environmental constraints
Policies: regulation and permitting
Power: electricity and grid availability
People: availability of skilled construction workers
Pushback: growing resistance from local communities
What I found even more interesting was that these constraints are prompting a rethink of not just how quickly we can build computing capacity, but how much we actually need.
Furthermore, token economics, Chinese models, and open-weight models are forcing a closer look at the economics of inference and the architecture required to run AI.
I left the conference with five ideas to explore:
Deployed GPU capacity may be underutilized for inference.
Memory is becoming the constraint, forcing a rethink of how AI systems are architected.
Networks are becoming the critical connecting layer between storage and compute.
Frontier models aren’t always necessary. SLMs and open-weight models may often deliver the best economics and performance for specific tasks.
Today, we hear about agentic pilots using reasoning every step of the way, but many enterprise processes need repeatability more than reasoning, pointing toward hybrid workflows that combine deterministic automation with agentic reasoning where needed.
We are still figuring out what AI actually needs to scale. AI infrastructure constraints are forcing a rethink of the assumptions underlying the current AI scaling narrative.



