在一场引人入胜的对话中,Anthropic AI 研究与实验室产品负责人 Diane Penn 深入探讨了公司的发展历程、不断演变的产品格局,以及构建前沿人工智能所面临的独特挑战与机遇。三年多前,当 Anthropic 仅有五名工程师时,Penn 作为公司的首位技术产品经理加入。她亲眼见证了 Anthropic 从一个不被看好的“弱者”飞速崛起,成为领先的人工智能巨头。
**从“弱者”到前沿人工智能领导者**
Penn 回忆起 Anthropic 的早期,当时包括她在内的许多人都觉得,面对像 OpenAI 这样的老牌巨头,他们“毫无胜算”。公司最初的定位是一场探索,以强烈的“创业”文化和对自身使命的坚定承诺为特色。一个早期但小众的转折点是 2024 年的“金门克劳德”(Golden Gate Claude)实验,研究人员在其中展示了模型“痴迷”于某一主题的能力。这种涵盖工程、产品和研究的快速原型开发,增强了 Anthropic 能够构建独特、真实产品的信心。
主要里程碑包括 Opus 3 的开发,尽管以今天的标准来看它“并不出色”,但对当时这家小公司而言,它是一个关键时刻,巩固了其训练前沿模型的能力。一年后的 Opus 4.5 标志着另一个重要的转折点,尤其当它与 Claude Code 等创新产品结合使用时。Penn 强调:“你需要前沿产品才能拥有前沿模型,才能让人们感受到前沿模型的魔力。” 模型能力和产品体验之间的这种协同作用,对产品的普及至关重要。
**驾驭人工智能指数级增长**
Penn 将当前时代描述为身处人工智能进步的“指数曲线之内”,每一次改进都带来能力的巨大飞跃。这种快速的节奏要求具备适应性、第一性原理思维,以及拥抱“涌现能力”的意愿——即模型功能中常常不可预测的、非连续性的飞跃。团队必须不断发现模型解锁的新可能性,这使得“共同发现”和共享实验变得至关重要。尽管 Penn 认可 Gary Tan 的“最大化 token 利用”(token-maxing)理念,但她将其重新定义为实现实验的一种手段,强调了内部协作和思想共享的价值。
Anthropic 的“实验室”(Labs)团队,由 Penn 负责产品,专注于识别和追求“可能不在核心路线图中的非连续性大赌注”。这个孵化器已经产生了 Claude Code、Skills、Claude Design 和 MCP 等重要产品。Labs 团队的文化鼓励对主题持有强烈见解,但不过度执着于精确的原型,从而允许在不同模型世代之间进行探索和重新审视想法。
**产品角色演变:“评估集是新的产品需求文档”**
Penn 强调了人工智能领域产品管理角色发生的重大转变,即“评估集是新的产品需求文档(PRD)”。产品经理不再仅仅依赖传统的产品需求文档,而是将细致的用户反馈(例如,“Claude 出现了幻觉”)转化为可操作、可衡量的评估集供研究人员使用。例如,早期的 Claude 模型在 JSON 输出方面存在问题,这个痛点被提炼成一个包含 30-40 个示例的“评估集”,以便持续衡量和改进性能。这种方法类似于产品经理的测试驱动开发,缩短了模型改进的可操作性距离。尽管产品需求文档仍用于更广泛的对齐和解决模糊问题,但重点已转向对模型能力的直接、可衡量的影响。
Penn 还强烈主张产品负责人要亲力亲为,积极地亲自尝试和部署人工智能。这种“亲力亲为的管理者”方法有助于与快速演进的模型保持“心智理论”的理解,并促进更好的决策。
**人工智能与人类能力**
关于 Fable/Mythos 等高级模型,Penn 承认它们面临更严格的审查和限制。尽管这为内部团队创造了暂时的优势,但 Anthropic 的目标仍然是让这些强大的技术广泛可用。
也许最引人入胜的是,Penn 分享了她如何个人使用 Claude 进行“情商增强”——将其作为“教练”,借鉴《关键对话》(Crucial Conversations)等书籍的经验,帮助她准备艰难的对话。这表明人工智能的潜力超越了智力任务,能帮助人类成为更好的沟通者和管理者。她主张将人工智能用作能够反驳并提炼想法的“思维伙伴”,而不仅仅是顺从的助手,从而保护和增强人类的思维过程。
**保持步调**
展望未来,Penn 认为人类的判断力、毅力和积极主动性将继续发挥无可估量的价值。对于她自己的孩子,她鼓励他们培养好奇心、毅力以及发展他们的“内在声音”。为了避免在这种快节奏环境下的倦怠,Penn 强调团队协作、彻底的所有权(radical ownership)和相互支持。Anthropic 的“蜂巢思维”(hive mind)文化,即同事之间互相照应并在复杂问题上进行“思想融合”,让个人能够得到休息,并防止重担落在任何一个人身上。
Penn 总结道,人工智能领域始终需要以用户为中心的产品人才——那些充满好奇心、运用第一性原理思维,并拥有“修修补补、探索创新(tinkering, hackery)精神”的人,才能利用这项变革性技术产生积极影响。
In a captivating discussion, Diane Penn, Head of Product for AI Research and Labs at Anthropic, offered an insightful look into the company's journey, the evolving product landscape, and the unique challenges and opportunities of building frontier AI. Joining Anthropic as its first technical product manager over three years ago when it had only five engineers, Penn has witnessed its meteoric rise from an underdog to a leading AI powerhouse.
**From Underdog to Frontier AI Leader**
Penn recalls the early days of Anthropic when many, including herself, felt they "had no chance" against established giants like OpenAI. The company's initial identity was a quest, marked by a strong "startup-y" culture and a commitment to its mission. An early, albeit niche, inflection point was the "Golden Gate Claude" experiment in 2024, where researchers showcased the model's ability to "obsess" over a theme. This rapid prototyping, spanning engineering, product, and research, instilled confidence that Anthropic could build unique, authentic products.
Major milestones included the development of Opus 3, which, despite being "not great" by today's standards, was a pivotal moment for the then-small company, solidifying its ability to train frontier models. Opus 4.5, a year later, marked another significant inflection, especially when paired with innovative products like Claude Code. Penn emphasizes, "You need frontier products in order to have frontier models and for people to feel the magic of frontier models." This synergy between model capability and product experience proved critical for adoption.
**Navigating the AI Exponential**
Penn describes the current era as being "inside the exponential curve" of AI advancement, where every improvement brings a massive jump in capabilities. This rapid pace necessitates adaptability, first-principles thinking, and a willingness to embrace "emerging capabilities"—discontinuous jumps in model function that are often unpredictable. The team must constantly discover what new possibilities a model unlock, making "communal discovery" and shared experimentation vital. While acknowledging Gary Tan's "token-maxing" idea, Penn reframes it as a means to achieve experimentation, stressing the value of internal collaboration and sharing ideas.
Anthropic's "Labs" team, where Penn leads product, focuses on identifying and pursuing "discontinuous large bets that might not be in the core roadmap." This incubator has yielded major products like Claude Code, Skills, Claude Design, and MCP. Labs thrives on a culture that embraces strong opinions on themes but weak attachments to exact prototypes, allowing for exploration and revisiting ideas across model generations.
**The Evolving Product Role: "Evals are the New PRDs"**
Penn highlights a significant shift in the product management role within AI, where "evals are the new PRDs." Instead of traditional Product Requirements Documents, PMs now translate nuanced user feedback (e.g., "Claude hallucinated") into actionable, measurable evaluation sets for researchers. For instance, early Claude models struggled with JSON output; this pain point was distilled into an "eval set" of 30-40 examples to consistently measure and improve performance. This approach, akin to test-driven development for PMs, shortens the distance to actionability for model improvements. While PRDs still serve for broader alignment and ambiguous problems, the emphasis has shifted to direct, measurable impact on model capabilities.
Penn also strongly advocates for product leaders to be hands-on, actively tinkering and shipping with AI themselves. This "hands-on manager" approach helps maintain a "theory of mind" with the rapidly evolving models and fosters better decision-making.
**AI and Human Capabilities**
On the topic of advanced models like Fable/Mythos, Penn acknowledges the increased scrutiny and restrictions. While this creates a temporary advantage for internal teams, Anthropic's goal remains to make these powerful technologies broadly accessible.
Perhaps most intriguingly, Penn shared how she personally uses Claude for "EQ augmentation"—as a "coach" to prepare for difficult conversations, drawing lessons from books like *Crucial Conversations*. This demonstrates AI's potential beyond just IQ tasks, helping humans become better communicators and managers. She advocates for using AI as a "thinking partner" that can push back and refine ideas, rather than just a compliant assistant, thus protecting and enhancing human thought processes.
**Sustaining the Pace**
Looking ahead, Penn believes human judgment, persistence, and proactivity will remain invaluable. For her own children, she encourages curiosity, persistence, and the development of their "inner voice." To avoid burnout in such a high-paced environment, Penn emphasizes team collaboration, radical ownership, and mutual support. The "hive mind" culture at Anthropic, where colleagues look out for each other and mind-meld on complex problems, allows individuals to recharge and prevents the burden from falling on any single person.
Penn concludes by asserting the enduring need for user-centric product people in AI—individuals who are deeply curious, apply first-principles thinking, and possess a "tinkering, hackery spirit" to harness this transformative technology for positive impact.