20VC: "Anti-Data Centres is a Chinese Psyop" | How Many Planned Data Centers Will Actually Get Built? | Is Energy AI's Biggest Bottleneck? With Thomas Sohmers, Co-Founder @ Positron
发布时间 来源
Episode 设置
在与哈里·斯蒂宾斯(Harry Stebbings)在20VC节目中的一次讨论中,Positron AI的联合创始人兼董事长托马斯·萨默斯(Thomas Somers)深入探讨了人工智能基础设施不断演变的格局,特别侧重于推理及其更广泛的影响。Positron AI制造硬件——包括芯片、软件和系统——为生成式AI推理提供动力,而生成式AI推理是AI模型的部署阶段。
萨默斯解释说,训练AI模型是计算密集型的,受益于更多的浮点运算。相比之下,推理是内存密集型的,因为每个生成的token都需要读取模型参数,这个过程无法像训练那样大规模并行化。这凸显了日益增长的“内存墙”问题,过去十年中,GPU的浮点运算能力提升了120倍,但内存带宽仅提升了17倍,这主要是因为SRAM技术进展较慢,以及在Transformer模型出现之前,人们长期关注计算密集型卷积神经网络。
谈到token经济学,萨默斯指出供应商在缓存token上赚取“惊人”的利润,因为处理一个缓存的token的成本仅是重新计算它的“千分之一”。他驳斥了“荒谬”的说法,即像OpenAI和Anthropic这样的公司不盈利,称如果他们停止训练,他们将“一夜之间获得巨额利润”。关于“限制前沿发展”的争论,萨默斯担心暂停可能会助长“卢德主义情绪”,并有将技术能力集中在少数实体手中的风险,他称之为“对古典自由主义概念的最大攻击”。他暗示,一些支持限制发展的倡议可能在IPO(首次公开募股)之前具有战略意义,同时他也承认Anthropic的真正担忧。他还强调了西方国家限制发展而中国不限制的危险,担心出现“中国共产党(CCP)控制的超级智能AI情景”。
萨默斯颇具争议地将左右两翼日益增长的反数据中心情绪称为“几乎完全是中国的心理战(psyop)”。他认为,关于用水和用电量的常见批评是基于“明显不实”的信息,并引用数据称,“单次进出(例如,一个航班或一次出行)所耗费的水量,你知道,比美国最大的数据中心耗费的水量还要多。”他批评西方国家的监管障碍阻碍了在广阔无人区建设数据中心,并将其与中国不受限制的建设形成了对比。
KV缓存被认为是提高推理效率的关键,它将基于提示和生成token的键(Key)和值(Value)矩阵存储起来以避免重复计算。尽管它节省了计算资源,但却需要大量内存,这构成了一个复杂的挑战,因为单个用户会话可能消耗数百GB,甚至超过模型本身的权重。萨默斯详细阐述了量化技术,这些技术将数据类型从FP16压缩到FP4,在对模型性能影响最小的情况下减少内存使用。
展望未来,萨默斯预测前沿模型的规模将继续增长,其中GPT-6 Astra将是一个重大飞跃,他认为它达到了AGI(通用人工智能)的水平。他分享了Astra在复杂编码、使用通用工具(如Blender)操作电脑,以及在50多个小时内完成从RTL到GDS的完整芯片设计流程的“令人难以置信”的能力,而这项任务人类通常需要数周才能完成。他认为,尽管更小、企业专属的模型将会出现,但它们将矛盾地推动大型云端模型消耗*更多*的token,因为本地LLM(大型语言模型)会不断为其更智能的云端对应模型生成任务。
token的成本已从五年前的每百万60美元大幅下降到如今的不到1美元。萨默斯认为,尽管价格下降了,但由于模型能力的提升,每个token的价值却增加了“一百到一千倍”。他预计定价模式将超越基于token的成本,演变为“每个有用结果的成本”或虚拟代理的订阅模式。
最后,萨默斯对“人类对齐”表示乐观,他强调解决地缘政治和社会问题对于驾驭人工智能的未来至关重要。他强调了考虑经济因素的重要性,他认为经济因素最终推动了AI格局中的真实决策和进步。
In a discussion with Harry Stebbings on 20VC, Thomas Somers, co-founder and chairman at Positron AI, delved into the evolving landscape of AI infrastructure, particularly focusing on inference, and its broader implications. Positron AI builds hardware—chips, software, and systems—to power generative AI inference, the deployment phase of AI models.
Somers explained that training AI models is compute-bound, benefiting from more floating-point operations. In contrast, inference is heavily memory-bound because each generated token requires reading model parameters, a process that cannot be massively parallelized like training. This highlights a growing "memory wall," where GPU flops have improved 120x in a decade, but memory bandwidth only 17x, largely due to slower advancements in SRAM technology and a historical focus on compute-bound convolutional neural networks before the transformer model.
On token economics, Somers noted the "insane" margins providers make on cached tokens, where processing a cached token is "one one thousandth of the cost" of recomputing it. He dismissed the "absurd" meme that companies like OpenAI and Anthropic are unprofitable, stating they would be "massively profitable overnight" if they stopped training. Regarding the "pacing the frontier" debate, Somers expressed concern that a pause could empower "luddite sentiment" and risk concentrating technological capability among a few entities, which he called "the biggest attack on classical, like liberal freedom concepts." He suggested that some advocacy for pacing might be strategic ahead of IPOs, while acknowledging Anthropic's genuine concerns. He also highlighted the danger of Western nations pacing while China does not, fearing a "Chinese CCP controlled, super intelligent AI scenario."
Somers controversially labeled the growing anti-data center sentiment on both the left and right as "almost entirely a Chinese psyop." He argued that common criticisms about water and power usage are based on "patently false" information, citing that "a single in and out uses, you know, more water than, you know, the largest data centers in the United States." He criticized regulatory hurdles in the West that prevent data center construction in vast, unpopulated areas, contrasting it with China's unconstrained build-out.
KV caching, where Key and Value matrices based on prompts and generated tokens are stored to avoid recomputation, was identified as crucial for inference efficiency. While it saves compute, it demands significant memory, creating a complex challenge as individual user sessions can consume hundreds of gigabytes, exceeding model weights. Somers elaborated on quantization techniques that compress data types from FP16 to FP4, reducing memory usage with minimal model performance degradation.
Looking ahead, Somers predicted frontier model sizes would continue to grow, with GPT-6 Astra being a significant leap he considers AGI. He shared personal experiences of Astra's "mind boggling" capabilities in complex coding, computer use with generic tools (like Blender), and completing full RTL to GDS chip design flows in just over 50 hours, a task typically taking weeks for humans. He believes that while smaller, enterprise-specific models will emerge, they will paradoxically drive *more* token consumption from large cloud models, as local LLMs constantly generate tasks for their more intelligent cloud counterparts.
The cost of tokens has drastically fallen from $60 per million five years ago to under $1 today. Somers argued that while the price decreased, the value per token has increased "a hundred or a thousand fold" due to improved model capabilities. He expects pricing models to evolve beyond per-token costs to "cost per useful result" or subscription models for virtual agents.
Finally, Somers expressed optimism about "human alignment," emphasizing that addressing geopolitical and social problems is crucial for navigating AI's future. He stressed the importance of considering economic factors, which he believes ultimately drive real decision-making and progress in the AI landscape.
摘要
Thomas Sohmers is the co-founder and chairman of Positron AI, building chips to make running AI dramatically cheaper and more energy-efficient. The company recently announced an $875 million Series C at a $5 billion valuation, backed by investors including Gavin Baker's Atreides Management, NEA, Valor Equity Partners and Netscape co-founder Jim Clark. AGENDA: 04:30 Why Does AI Inference Need Different Hardware from Training? 10:25 What Is Nobody Telling You About AI's Token Economics? 13:10 Should We Really Slow Down the AI Frontier? 16:55 What Happens If America Slows Down—and China Doesn't? 19:30 Should We Restrict China's Access to AI Chips? 20:15 Why Aren't Zuckerberg and Jensen Backing an AI Slowdown? 21:25 Are We Being Misled About Data Centers? 23:25 Can the West Beat China While Drowning in Regulation? 26:20 How Many Planned Data Centers Will Actually Get Built? 27:15 Data Centers in Space: Real Opportunity or Elon Hype? 28:10 Is Energy AI's Biggest Bottleneck? 29:55 Could the Debt Boom Derail AI? 31:25 What Is KV Caching—and Why Does It Change AI Economics? 35:25 Does Compressing AI's Memory Make It Less Intelligent? 41:05 How Do We Solve AI's Exploding Memory Demands? 43:45 Will Every Company Own Its Own AI Model? 47:10 Why Bet on Ever-Bigger Models? 48:20 If We Already Have AGI, What Comes Next? 53:25 What Happens When Every AI Lab Builds Its Own Chips? 57:30 If DeepSeek Can Slash Costs with Software, Why Build New Hardware? 59:50 How Cheap Will AI Tokens Be by 2028? 1:03:15 What Would Be the First Warning Sign of an AI Bust? 1:03:55 Could Open Models Break OpenAI and Anthropic's Growth? 1:05:55 Could Mercor and Surge Become $200 Billion Companies?
GPT-4正在为你翻译摘要中......
