Amin Vedat, Head of Google's AI Infra, highlights the unprecedented scale of the current CapEx buildout in human history, with Google alone projected to spend over $200 billion this year, primarily on data centers. He emphasizes that while AI data centers share similarities with traditional ones, the key difference lies in "specialization." Unlike 20-30 year planning horizons for fungible general-purpose data centers, AI centers are often "purpose-built," co-designing the building with the hardware, considering power, cooling, and networking needs specific to high-density AI racks (e.g., TPUs, GPUs) which can easily hit hundreds of kilowatts, contrasting with lower-power storage racks.
Vedat introduces "good put" as a critical metric for accountability, preferring it over "flops" or other chip-centric measures. Good put measures actual workload performance, accounting for reliability and recovery from failures. At the 100,000-accelerator scale, something is failing multiple times a day or hour, necessitating rapid detection and recovery to minimize re-computation and downtime. Common failure reasons are diverse and constantly evolving, including hardware, network, and software (compiler, runtime, model) issues.
The TPU program, initiated in 2013, was a "contrarian call" against the "bitter lesson of chips" that specialization never wins. However, for a few extremely demanding applications like language translation and voice recognition, custom acceleration proved immensely successful. The program evolved significantly with the invention of transformers, expanding from inference to training, recommender systems, and now generative AI. The art of co-design involves deciding how much to specialize. For instance, Google's recent release of TPU 8i for inference and 8t for training reflects a projection of significant market growth for inference, justifying separate specialized chips that are still flexible enough to handle the other workload if needed. This contrasts with a single general-purpose chip, which might offer uniformity but sacrifices peak performance.
Google's co-design philosophy is deeply embedded in the working relationship between Vedat's team and DeepMind. This "shoulder-to-shoulder" partnership allows for real-time adjustments to chip architecture based on model optimizations, even delaying tape-outs for significant gains. While hardware planning cycles are years long, a generation of researchers has learned to influence the roadmap, ensuring deep integration from concept to production. The core TPU architecture, focusing on linear algebra primitives like matrix multiply and sparse operations, has shown remarkable durability across many generations of models and algorithms.
Vedat identifies "long-horizon agents" as a significant shift in workload, where automated interactions happen in milliseconds, compared to human-paced seconds. This dramatically increases demands not just on accelerators, but also on CPUs for reasoning and orchestration, and on networking and storage to fetch context data. This necessitates rethinking data center design, balancing high-density accelerator racks with the needs of co-located CPUs and storage, or distributing them across buildings and managing inter-building networking complexity.
Google's networking innovations, such as wave division multiplexing and optical circuit switching, are crucial. Optical circuit switching, using MEMS mirrors, transmits data entirely in the optical domain, allowing for dynamic reconfiguration of network spine and rapid redirection of traffic (e.g., rerouting light to a spare rack in milliseconds when one fails) without physically moving fiber.
Power is identified as "the single most fundamental constraint." Google's preferred model is to partner with utilities for gigawatt-scale power provisioning, involving long-term co-planning (5+ years) and ensuring Google covers the infrastructure costs to avoid impacting other ratepayers. This partnership model, leveraging "statistical multiplexing," offers mutual benefits and flexibility compared to vertical integration. Data center sizing is an art, depending on workload (training often benefits from larger, concentrated sites; inference needs distributed capacity closer to users), power availability, and local constraints.
Regarding hardware lifecycle, Vedat notes that older TPUs (7-8 years old) still run at 100% utilization, but are typically replaced after depreciation (around six years) with newer, more power-efficient systems. The discussion also touches on Google's commitment to "open standards" in software frameworks (e.g., supporting PyTorch alongside JAX) and hardware, believing it fosters growth and avoids "walled gardens."
Finally, Vedat entertains "out there" concepts like "orbital data centers." Driven by the fundamental constraint of energy, space offers 40% more solar power and near-100% sunlight coverage, eliminating batteries. While challenges like cooling, reliability, and networking (free-space optics) exist, there are no "showstoppers." Looking 10 years ahead to 2036, he envisions supercomputers characterized by extreme integration and modularity, with racks becoming multi-megawatt units manufactured centrally, requiring only power, water, and a single fiber connection upon deployment.