Cerebras CEO: Why GPUs Can't Do Fast Inference
发布时间 来源
Episode 设置
摘要
AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down with Andrew Feldman, co-founder and CEO of Cerebras, to explain why fast inference may define the next era of AI.
Cerebras is known for building a chip the size of a silicon wafer. But this conversation is not just about one company or one chip. It is a deep dive into the AI infrastructure stack: GPUs, ASICs, memory, HBM, SRAM, data centers, power, TSMC, AWS, OpenAI, agents, reasoning models, and why speed changes what AI products can become. Andrew explains why “tokens per second per user” matters, why generating a single word can require moving the equivalent of 100 HD movies through memory, why agents amplify latency, why GPUs struggle with certain inference workloads, and why fast AI may eventually reshape SaaS itself.
This is a reference conversation on fast inference, AI chips, and the next compute bottleneck.
Andrew Feldman
LinkedIn - https://www.linkedin.com/in/andrewdfeldman
X/Twitter - https://x.com/andrewdfeldman
Cerebras
Website - https://www.cerebras.ai/
X/Twitter - https://x.com/cerebras
Matt Turck (Managing Director)
Blog - https://mattturck.com
LinkedIn - https://www.linkedin.com/in/turck/
X/Twitter - https://x.com/mattturck
FirstMark Capital
Website - https://firstmark.com
X/Twitter - https://x.com/FirstMarkCap
Listen on:
Spotify - https://open.spotify.com/show/7yLATDSaFvgJG80ACcRJtq
Apple - https://podcasts.apple.com/us/podcast/the-mad-podcast-with-matt-turck/id168623872
00:00 Cold open & Intro
01:31 Why speed became the AI bottleneck
02:32 Tokens per second per user, explained
03:16 AI’s broadband moment and the Netflix analogy
04:35 The AI chip landscape: GPUs, TPUs, Trainium, ASICs
06:36 What is an ASIC?
08:08 Nvidia, Groq, and the fast inference war
09:16 OpenAI, Broadcom, and specialized silicon
12:10 China, power, and sovereign AI infrastructure
15:05 Is the AI infrastructure boom a bubble?
18:56 The hidden bottlenecks: HBM, CoWoS, and 3nm
22:57 Why agents are creating CPU demand
25:36 Andrew Feldman’s path from SeaMicro to Cerebras
26:13 Why Cerebras bet on AI in 2016
31:14 SRAM vs. HBM: why inference is a memory problem
33:19 What wafer-scale computing actually means
34:28 The deep-tech “Everest” problem
36:07 The moment the first Cerebras system worked
36:49 Ringing the bell and surviving deep tech
39:08 How a giant chip handles failure
41:22 Why GPUs struggle with decode
42:17 Prefill vs. decode explained
44:01 The “100 HD movies” problem in AI inference
45:04 How fast inference changes RL and training
48:08 Reasoning models and why they cost more compute
50:08 Verification, guardrails, and small models checking big models
52:37 Multimodal AI and the path to video
53:51 Cerebras’ business model: hardware, cloud, and API
55:14 OpenAI’s 750MW inference deal
55:36 Why data centers are measured in megawatts
58:01 AWS Trainium + Cerebras decode
59:29 Fast tokens as a cloud product
01:00:52 Is CUDA still a moat?
01:03:53 How TSMC helped Cerebras build the giant chip
01:07:41 Why nobody cared in 2020
01:08:15 Why chip supply chains are hard to diversify
01:09:54 Why today’s AI models will be the worst you ever use
01:10:38 What fast AI could do to SaaS
GPT-4正在为你翻译摘要中......
