Cerebras CEO: Why GPUs Can't Do Fast Inference

发布时间    来源
Episode 设置


登录已过期或未登录,无法修改。请先登录后再试。

摘要

AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down with Andrew Feldman, co-founder and CEO of Cerebras, to explain why fast inference may define the next era of AI. Cerebras is known for building a chip the size of a silicon wafer. But this conversation is not just about one company or one chip. It is a deep dive into the AI infrastructure stack: GPUs, ASICs, memory, HBM, SRAM, data centers, power, TSMC, AWS, OpenAI, agents, reasoning models, and why speed changes what AI products can become. Andrew explains why “tokens per second per user” matters, why generating a single word can require moving the equivalent of 100 HD movies through memory, why agents amplify latency, why GPUs struggle with certain inference workloads, and why fast AI may eventually reshape SaaS itself. This is a reference conversation on fast inference, AI chips, and the next compute bottleneck. Andrew Feldman LinkedIn - https://www.linkedin.com/in/andrewdfeldman X/Twitter - https://x.com/andrewdfeldman Cerebras Website - https://www.cerebras.ai/ X/Twitter - https://x.com/cerebras Matt Turck (Managing Director) Blog - https://mattturck.com LinkedIn - https://www.linkedin.com/in/turck/ X/Twitter - https://x.com/mattturck FirstMark Capital Website - https://firstmark.com X/Twitter - https://x.com/FirstMarkCap Listen on: Spotify - https://open.spotify.com/show/7yLATDSaFvgJG80ACcRJtq Apple - https://podcasts.apple.com/us/podcast/the-mad-podcast-with-matt-turck/id168623872 00:00 Cold open & Intro 01:31 Why speed became the AI bottleneck 02:32 Tokens per second per user, explained 03:16 AI’s broadband moment and the Netflix analogy 04:35 The AI chip landscape: GPUs, TPUs, Trainium, ASICs 06:36 What is an ASIC? 08:08 Nvidia, Groq, and the fast inference war 09:16 OpenAI, Broadcom, and specialized silicon 12:10 China, power, and sovereign AI infrastructure 15:05 Is the AI infrastructure boom a bubble? 18:56 The hidden bottlenecks: HBM, CoWoS, and 3nm 22:57 Why agents are creating CPU demand 25:36 Andrew Feldman’s path from SeaMicro to Cerebras 26:13 Why Cerebras bet on AI in 2016 31:14 SRAM vs. HBM: why inference is a memory problem 33:19 What wafer-scale computing actually means 34:28 The deep-tech “Everest” problem 36:07 The moment the first Cerebras system worked 36:49 Ringing the bell and surviving deep tech 39:08 How a giant chip handles failure 41:22 Why GPUs struggle with decode 42:17 Prefill vs. decode explained 44:01 The “100 HD movies” problem in AI inference 45:04 How fast inference changes RL and training 48:08 Reasoning models and why they cost more compute 50:08 Verification, guardrails, and small models checking big models 52:37 Multimodal AI and the path to video 53:51 Cerebras’ business model: hardware, cloud, and API 55:14 OpenAI’s 750MW inference deal 55:36 Why data centers are measured in megawatts 58:01 AWS Trainium + Cerebras decode 59:29 Fast tokens as a cloud product 01:00:52 Is CUDA still a moat? 01:03:53 How TSMC helped Cerebras build the giant chip 01:07:41 Why nobody cared in 2020 01:08:15 Why chip supply chains are hard to diversify 01:09:54 Why today’s AI models will be the worst you ever use 01:10:38 What fast AI could do to SaaS

GPT-4正在为你翻译摘要中......

中英文字稿