首页  >>  来自播客: Sequoia Capital 更新   反馈  

Sequoia Capital - Google's AI Infrastructure Chief, Amin Vahdat, on the Physics & Economics of Frontier AI

发布时间:   原节目
以下是原文的中文翻译: Amin Vedat,Google人工智能基础设施负责人,强调了人类历史上当前资本支出(CapEx)建设的空前规模,仅谷歌今年就预计将投入超过2000亿美元,主要用于数据中心。他强调,尽管人工智能数据中心与传统数据中心有相似之处,但关键区别在于“专业化”。与通用型(fungible)数据中心20-30年的规划周期不同,AI中心通常是“专用建造的”,将建筑与硬件协同设计,考虑到高密度AI机架(例如TPU、GPU)特有的电力、散热和网络需求,这些机架的功耗可以轻松达到数百千瓦,这与低功耗的存储机架形成鲜明对比。 Vedat引入了“有效吞吐量”(good put)作为衡量责任的关键指标,他倾向于使用它而非“flops”或其他以芯片为中心的衡量标准。“有效吞吐量”衡量实际的工作负载性能,并考虑了可靠性和从故障中恢复的能力。在10万加速器的规模下,每天或每小时都会有设备发生多次故障,这需要快速检测和恢复,以最大限度地减少重复计算和停机时间。常见的故障原因多种多样且不断演变,包括硬件、网络以及软件(编译器、运行时、模型)问题。 TPU项目始于2013年,是一个“逆向选择”(contrarian call),挑战了“芯片的痛苦教训”,即专业化永远无法取胜的观点。然而,对于少数极度苛刻的应用,如语言翻译和语音识别,定制加速取得了巨大成功。随着Transformer架构的发明,该项目取得了显著发展,从推理扩展到训练、推荐系统,以及现在的生成式AI。协同设计的艺术在于决定专业化到何种程度。例如,谷歌最近发布的用于推理的TPU 8i和用于训练的8t,反映了对推理市场显著增长的预测,这使得开发单独的专用芯片变得合理,这些芯片在必要时仍足够灵活以处理其他工作负载。这与单一通用芯片形成对比,尽管通用芯片可能提供统一性,但却牺牲了峰值性能。 谷歌的协同设计理念深入融入于Vedat的团队与DeepMind之间的工作关系中。这种“并肩协作”的伙伴关系使得能够基于模型优化对芯片架构进行实时调整,甚至为了获得显著收益而推迟流片(tape-outs)。尽管硬件规划周期长达数年,但一代研究人员已经学会影响其路线图,确保从概念到生产的深度整合。TPU的核心架构专注于矩阵乘法和稀疏操作等线性代数原语,在多代模型和算法中都展现出了卓越的持久性。 Vedat将“长周期智能体”(long-horizon agents)视为工作负载的重大转变,在这种情况下,自动化交互在毫秒级发生,而人类的节奏是秒级。这不仅大大增加了对加速器的需求,还增加了对用于推理和编排的CPU的需求,以及对用于获取上下文数据的网络和存储的需求。这需要重新思考数据中心设计,平衡高密度加速器机架与共置CPU和存储的需求,或者将它们分布在不同的建筑中,并管理跨建筑的网络复杂性。 谷歌的网络创新,例如波分复用(wave division multiplexing)和光电路交换(optical circuit switching),至关重要。光电路交换利用微机电系统(MEMS)镜,完全在光域中传输数据,允许对网络主干进行动态重新配置和快速流量重定向(例如,当一个机架发生故障时,在毫秒级内将光路由到备用机架),而无需物理移动光纤。 电力被认为是“最根本的单一限制”。谷歌的首选模式是与公用事业公司合作,提供千兆瓦级的电力供应,这涉及长期共同规划(5年以上),并确保谷歌承担基础设施成本,以避免影响其他电力用户。这种合作模式,利用“统计复用”(statistical multiplexing),与垂直整合相比,提供了互惠互利和灵活性。数据中心规模的确定是一门艺术,取决于工作负载(训练通常受益于更大、更集中的站点;推理需要更接近用户的分布式容量)、电力可用性和当地限制。 关于硬件生命周期,Vedat指出,较旧的TPU(7-8年前的)仍然以100%的利用率运行,但通常在折旧期后(大约六年)会被更新、更节能的系统取代。讨论还涉及谷歌对软件框架(例如,在JAX旁边支持PyTorch)和硬件的“开放标准”的承诺,认为这能促进增长并避免“封闭生态系统”(walled gardens)。 最后,Vedat提出了“异想天开”(out there)的概念,例如“轨道数据中心”。鉴于能源这一根本限制,太空能提供多40%的太阳能,并拥有近100%的日照覆盖,从而无需电池。尽管存在散热、可靠性和网络(自由空间光学)等挑战,但没有“决定性障碍”(showstoppers)。展望未来10年到2036年,他设想的超级计算机将以极致集成和模块化为特征,机架将成为集中制造的多兆瓦级单元,部署时只需电力、水和一根光纤连接。

Amin Vedat, Head of Google's AI Infra, highlights the unprecedented scale of the current CapEx buildout in human history, with Google alone projected to spend over $200 billion this year, primarily on data centers. He emphasizes that while AI data centers share similarities with traditional ones, the key difference lies in "specialization." Unlike 20-30 year planning horizons for fungible general-purpose data centers, AI centers are often "purpose-built," co-designing the building with the hardware, considering power, cooling, and networking needs specific to high-density AI racks (e.g., TPUs, GPUs) which can easily hit hundreds of kilowatts, contrasting with lower-power storage racks. Vedat introduces "good put" as a critical metric for accountability, preferring it over "flops" or other chip-centric measures. Good put measures actual workload performance, accounting for reliability and recovery from failures. At the 100,000-accelerator scale, something is failing multiple times a day or hour, necessitating rapid detection and recovery to minimize re-computation and downtime. Common failure reasons are diverse and constantly evolving, including hardware, network, and software (compiler, runtime, model) issues. The TPU program, initiated in 2013, was a "contrarian call" against the "bitter lesson of chips" that specialization never wins. However, for a few extremely demanding applications like language translation and voice recognition, custom acceleration proved immensely successful. The program evolved significantly with the invention of transformers, expanding from inference to training, recommender systems, and now generative AI. The art of co-design involves deciding how much to specialize. For instance, Google's recent release of TPU 8i for inference and 8t for training reflects a projection of significant market growth for inference, justifying separate specialized chips that are still flexible enough to handle the other workload if needed. This contrasts with a single general-purpose chip, which might offer uniformity but sacrifices peak performance. Google's co-design philosophy is deeply embedded in the working relationship between Vedat's team and DeepMind. This "shoulder-to-shoulder" partnership allows for real-time adjustments to chip architecture based on model optimizations, even delaying tape-outs for significant gains. While hardware planning cycles are years long, a generation of researchers has learned to influence the roadmap, ensuring deep integration from concept to production. The core TPU architecture, focusing on linear algebra primitives like matrix multiply and sparse operations, has shown remarkable durability across many generations of models and algorithms. Vedat identifies "long-horizon agents" as a significant shift in workload, where automated interactions happen in milliseconds, compared to human-paced seconds. This dramatically increases demands not just on accelerators, but also on CPUs for reasoning and orchestration, and on networking and storage to fetch context data. This necessitates rethinking data center design, balancing high-density accelerator racks with the needs of co-located CPUs and storage, or distributing them across buildings and managing inter-building networking complexity. Google's networking innovations, such as wave division multiplexing and optical circuit switching, are crucial. Optical circuit switching, using MEMS mirrors, transmits data entirely in the optical domain, allowing for dynamic reconfiguration of network spine and rapid redirection of traffic (e.g., rerouting light to a spare rack in milliseconds when one fails) without physically moving fiber. Power is identified as "the single most fundamental constraint." Google's preferred model is to partner with utilities for gigawatt-scale power provisioning, involving long-term co-planning (5+ years) and ensuring Google covers the infrastructure costs to avoid impacting other ratepayers. This partnership model, leveraging "statistical multiplexing," offers mutual benefits and flexibility compared to vertical integration. Data center sizing is an art, depending on workload (training often benefits from larger, concentrated sites; inference needs distributed capacity closer to users), power availability, and local constraints. Regarding hardware lifecycle, Vedat notes that older TPUs (7-8 years old) still run at 100% utilization, but are typically replaced after depreciation (around six years) with newer, more power-efficient systems. The discussion also touches on Google's commitment to "open standards" in software frameworks (e.g., supporting PyTorch alongside JAX) and hardware, believing it fosters growth and avoids "walled gardens." Finally, Vedat entertains "out there" concepts like "orbital data centers." Driven by the fundamental constraint of energy, space offers 40% more solar power and near-100% sunlight coverage, eliminating batteries. While challenges like cooling, reliability, and networking (free-space optics) exist, there are no "showstoppers." Looking 10 years ahead to 2036, he envisions supercomputers characterized by extreme integration and modularity, with racks becoming multi-megawatt units manufactured centrally, requiring only power, water, and a single fiber connection upon deployment.