Why AI Demand Is Outrunning Compute Supply
发布时间 来源
Episode 设置
讨论一开始就断言,鉴于伊隆和黄仁勋通过人工智能和太空领域的进步对人类社会产生了根本性影响,21世纪将被铭记为“伊隆和黄仁勋的时代”。尽管深刻的新技术之后常出现泡沫的历史先例,但当前的人工智能领域以加速的活动和情绪为特征。发言人之一加文指出,尽管他询问了所有人,但没有在任何业务中找到一个正在恶化的定量数据点,这凸显了人工智能领域的全面加速,从OpenAI和开源到Grok莫不如此。
文章探讨了“所有人都能赢”的局面,认为Anthropic、OpenAI、SpaceX、Meta、谷歌、开源以及各种云计算和应用公司都能取得成功。Anthropic目前正处于静默期,其特点是招聘与使命保持一致,并相对于OpenAI等竞争对手战略性地安排模型发布时机。这些人工智能实验室的财务模型是独特的;它们每吉瓦算力能产生可观的年收入,同时保留了在推理和训练之间重新分配资源的灵活性。这意味着收入可能剧烈波动(例如,通过将算力从推理转向训练,收入可以从4800亿美元降至1200亿美元),这是公众市场需要理解的一个动态。受扩展定律驱动,预计这些实验室在可预见的未来将优先考虑训练而非自由现金流。
短期回报凸显了建设人工智能基础设施的经济现实。据报道,Nebius和Core等公司在算力投资上实现了9-10个月的短期回报,而SpaceX由于其庞大且快速部署的集群,回报速度甚至更快。这些企业极易获得融资,黑石、KKR和阿波罗等实体提供了低成本资本,部分原因是这些资产的使用寿命正在延长,并且代币支出的投资回报率正在增加。
在需求侧,当前人工智能的盈利模式基于一个惊人小的用户群,很可能只有“不到1000万”的重度付费用户,而非此前暗示的3000万。全球有15亿知识工作者,人工智能有巨大的扩散空间。公司已经将人力薪酬的1-10%以上用于代币。使用GrokBot等工具的个人经历表明代币消耗和生产力迅速增加。
尽管新技术通常存在高估和过度建设的历史模式,但当前的人工智能扩张面临着大规模的供应限制。人们担心全球算力短缺会影响从铜矿开采到电力基础设施等各个行业。发言人批评美国在监管和费率方面所处的“糟糕境地”,特别是在数据中心方面。他们认为数据中心是“对美国工薪阶层来说是最好的事情”,推动了再工业化,提振了小城镇的税收,并驳斥了环境担忧(如用水量)。他们强调需要讲述一个积极、具体的人工智能对普通美国人益处的故事,超越“领先中国”或“治愈癌症”等抽象概念,并引用了Meta成功的沟通策略。如果不加以解决,这种当前“自食其果”的供应不足可能导致“算力不平等”。
文章讨论了开源人工智能的作用;尽管很重要,但开源代币并非“免费”,因为它们需要算力,而且像Kimi这样的模型甚至可能规定30%的收入分成。
展望未来,SpaceX的轨道算力概念被视为解决地面限制的现实方案。他们驳斥了“死星”意象,将轨道数据中心描述为飞机大小的芯片机架,通过太阳同步轨道上的散热器进行冷却。物理并非障碍;成本才是,但SpaceX的历史表明其能大幅降低成本(例如,星舰的可重复使用性)。伊隆·马斯克和黄仁勋正在共同设计一个“鲁宾机架”,计划于2027年第四季度发射,这预示着世界上越来越多的算力将进入轨道,最初作为“摆动容量”。SpaceX更广泛的战略包括星链移动和宽带,使其无论轨道算力的即时影响如何,都将是一个强大的参与者。提出的最具未来感的想法是小行星采矿,像Psyche小行星这样的天体所含的贵金属比地壳还多,展望了一个重工业转向太空,地球主要用于居住的未来。
在此不断演变的格局中,微软的人工智能战略被视为“更友好”。他们没有直接参与前沿模型竞赛(他们在该领域举步维艰),而是专注于“模型组合”,利用企业可以使用自己的专有数据进行微调的开源基础模型。这种方法使公司能够在路由器后面“拥有和控制自己的智能”。这种智能的“抽象层”是一个竞争激烈的领域,微软、Databricks、Palantir以及像Harvey这样的应用特定公司都在争夺地位。成功取决于执行力、成本效率和垂直整合。
英伟达,在黄仁勋的领导下,被定位为“人工智能的中央银行”。其垂直整合但横向开放的战略,结合其为算力建设融资的能力(例如,一个500亿美元的数据中心有150亿美元的股权,其余由主要机构融资),提供了巨大的竞争优势。黄仁勋对人工智能碎片化的激励被认为对美国有利。与某些看法相反,开源对英伟达的业务有利,因为它推动了更多的代币消耗,从而增加了对算力的需求。硬件开发的挑战是公认的,很少有公司能成功应对其复杂性和融资要求。伊隆·马斯克决定与英伟达合作开发Grok被视为一个“非常高明的举动”,这承认了英伟达的主导地位和生态系统。在供应受限的市场中,客户对芯片的真实偏好难以推断,但交易类型(投资、残值担保、认股权证)揭示了潜在需求。
The discussion opens with the assertion that the 21st century will be remembered as "the age of Elon and Jensen," given their fundamental impact on human society through advancements in AI and space. Despite historical precedents of bubbles following profound new technologies, the current AI landscape is characterized by accelerating activity and sentiment. Gavin, one of the speakers, notes that despite asking everyone, he hasn't found a single quantitative data point in any business that is getting worse, highlighting an across-the-board acceleration in AI, from OpenAI and open source to Grok.
The "everyone wins" scenario is explored, suggesting that Anthropic, OpenAI, SpaceX, Meta, Google, open source, and various cloud and application companies can all succeed. Anthropic, currently in a quiet period, is noted for its mission-aligned hiring and strategic timing of model releases relative to competitors like OpenAI. The financial models of these AI labs are unique; they generate significant annual revenue per gigawatt of compute but retain the flexibility to reallocate resources between inference and training. This means revenue can fluctuate dramatically (e.g., from $480 billion to $120 billion by shifting compute from inference to training), a dynamic public markets will need to understand. Labs, driven by scaling laws, are expected to prioritize training over free cash flow for the foreseeable future.
The economic reality of building AI infrastructure is highlighted by short paybacks. Companies like Nebius and Core reportedly achieve 9-10 month paybacks on compute investments, even faster for SpaceX due to its large, rapidly deployed clusters. These ventures are highly financeable, with entities like Blackstone, KKR, and Apollo providing low-cost capital, partly because the useful lives of these assets are extending, and the ROI on token spend is increasing.
On the demand side, current AI monetization is based on a surprisingly small user base, likely "sub 10 million" heavy paying users, not the 30 million suggested. With 1.5 billion knowledge workers globally, there's immense room for diffusion. Companies are already spending 1-10%+ of human compensation on tokens. Personal experiences with tools like GrokBot demonstrate rapid increases in token consumption and productivity.
Despite historical patterns of overvaluation and overbuild with new technologies, the current AI expansion faces massive supply constraints. Concerns exist about global compute shortages impacting industries from copper mining to power infrastructure. The speakers criticize the "bad place" America is in regarding regulation and rates, particularly regarding data centers. They argue that data centers are "the best thing that has ever happened to working class Americans," driving re-industrialization, boosting tax revenues in small towns, and debunking environmental concerns (like water consumption). They emphasize the need to tell a positive, tangible story of AI's benefits for everyday Americans, beyond abstract notions of "staying ahead of China" or "curing cancer," citing Meta's successful communication strategy. This current "self-inflicted" undersupply could lead to "compute inequality" if not addressed.
The role of open source AI is discussed; while important, open source tokens are not "free" as they require compute, and models like Kimi might even stipulate a 30% revenue share.
Looking to the future, the concept of orbital compute from SpaceX is presented as a realistic solution to terrestrial constraints. Dismissing "Death Star" imagery, they describe orbital data centers as airplane-sized racks of chips, cooled by radiators in sun-synchronous orbits. Physics is not the barrier; cost is, but SpaceX's history suggests dramatic cost reductions (e.g., Starship reusability). Elon Musk and Jensen Huang are co-designing a "Rubin rack" for a Q4 2027 launch, suggesting an increasing fraction of the world's compute will be in orbit, initially as "swing capacity." SpaceX's broader strategy includes Starlink mobile and broadband, making it a formidable player regardless of orbital compute's immediate impact. The most futuristic idea presented is asteroid mining, with objects like asteroid Psyche containing more precious metals than Earth's crust, envisioning a future where heavy industry shifts to space, leaving Earth primarily residential.
Microsoft's AI strategy is seen as "friendlier" in this evolving landscape. Instead of competing directly in the frontier model race (where they struggled), they focus on an "ensemble of models," leveraging open-source base models that enterprises can fine-tune with their proprietary data. This approach allows companies to "own and control their intelligence" behind a router. This "abstraction layer" for intelligence is a highly contested space, with Microsoft, Databricks, Palantir, and application-specific companies like Harvey vying for position. Success hinges on execution, cost efficiency, and vertical integration.
NVIDIA, under Jensen Huang, is positioned as the "central bank of AI." Its vertically integrated but horizontally open strategy, combined with its ability to finance compute builds (e.g., $15 billion equity for a $50 billion data center, with the rest financed by major institutions), provides a massive competitive advantage. Jensen's incentives for AI fragmentation are seen as beneficial for America. Open source, contrary to some beliefs, is good for NVIDIA's business as it drives more token consumption and thus demand for compute. The challenge of hardware development is acknowledged, with few companies successfully navigating the complexities and financing requirements. Elon Musk's decision to partner with NVIDIA for Grok is seen as a "very high ELO move," acknowledging NVIDIA's dominance and ecosystem. The true customer preferences for chips are hard to infer in a supply-constrained market, but the types of deals (investments, residual value guarantees, warrants) reveal underlying demand.
摘要
a16z’s David George sits down with Gavin Baker to unpack the state of the AI boom, why demand for intelligence may still be dramatically underestimated, and why the outcome doesn't necessarily have to be winner-take-all.
David and Gavin explore the possibility that frontier labs, open-source models, applications, clouds, and NVIDIA can all capture significant value as AI adoption expands. They dig into the economics of the infrastructure buildout, why compute investments can have unusually fast payback periods, and what happens when today's relatively small group of heavy AI users expands to hundreds of millions of people.
They also debate the risk of an AI bubble versus an AI shortage, the backlash against data centers, orbital compute, the rise of multi-model architectures, and NVIDIA's position at the center of the AI supply chain. Gavin makes the case that the AI buildout could help reindustrialize America, while David explores whether the bigger near-term risk is not overbuilding, but failing to build enough.
Timestamps:
00:00 - Intro
01:06 - Finding the Bear Case: Why Gavin Can't Find One
08:06 - How's This All Gonna Go Wrong? An "And" Thing, Not an "Or"
09:09 - Will Labs Reinvest All Their Profits Into Training Forever?
14:44 - The Demand Side: 30 Million Heavy Users & the Diffusion Question
19:16 - From Reactive Coding to Fully Autonomous Agents
30:18 - What Happens If There's a Massive Supply Shortage?
34:21 - Orbital Data Centers: The SpaceX Compute Play
40:00 - Starlink's $2 Trillion Market & the Heads-You-Win Compute Bet
44:04 - The Most Futuristic SpaceX Idea: Asteroid Mining
54:01 - Who Becomes the Abstraction Layer of Intelligence?
57:57 - Harvey, Cursor & Vertical AI Winners
01:00:26 - Jensen, Nvidia & the Central Bank of AI
01:12:16 - How Chip Deal Structures Reveal True Customer Preference
Resources:
Follow Gavin Baker on X: https://x.com/GavinSBaker
Follow David George on X: https://x.com/DavidGeorge83
Stay Updated:
If you enjoyed this episode, be sure to like, subscribe, and share with your friends!
Find a16z on X: https://twitter.com/a16z
Find a16z on LinkedIn: https://www.linkedin.com/company/a16z
Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX
Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711
Follow our host: https://x.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.
GPT-4正在为你翻译摘要中......
