Where AI products go next: voice, agents, and self-driving software | Tara Sesha and Nan Yu (OpenAI)
发布时间 来源
Episode 设置
讨论由来自OpenAI的Claire拉开序幕,她承认交付不完美产品所面临的挑战,这与她在Stripe精益求精是常态的过往经历形成了鲜明对比。她强调,在当前的AI时代,紧迫性至关重要,用户在产品上迭代产生的经验证据胜过理论上的完美。“切换”功能(ChatGPT十亿多用户最初的代理控制工具)就是例证:一个被认为对即时功能和未来迭代至关重要的不完美解决方案。
随后,Non深入探讨了决定哪些功能应该废弃、哪些应该保留的问题。他强调了为用户提供一个“连贯的故事”的重要性,这能让他们能够了解产品的演变过程。如果用户理解这个过程,他们会更接受改变,即使这意味着在此过程中构建和废弃大量的代码或产品。
对话随后转向了尽管快速开发,仍能保持高品质标准所面临的制约。Claire概述了三个关键要素:首先,通过释放真正的价值,确保产品“对用户有增益”;其次,达到“内部标准”,即产品在内部进行试用,以评估其接受度、用户满意度和新颖用例;第三,以模型“两到三个月后”的状态为目标,避免过分拘泥于当下或过于未来化而无法使用。Non补充说,最大的制约是用户对新功能的“理解和吸收能力”,他指出存在“能力过剩”现象,即模型能做的比用户目前实际利用的要多。
他们探讨了普遍存在的观点,即B2B客户无法适应变化。Claire认为,尽管速度可能感觉“快得惊人”,但不推出前沿进步成果可能会让企业面临“被超越”的风险。她举例说明了代理革命,在这场革命中,企业需要快速部署代理,即使这会打破现有流程,以便释放更大的价值。这再次强调了构建用户 *真正需要* 的东西,而不仅仅是他们 *声称* 需要的东西这一真理。
一场关于单一身份代理与多身份代理的辩论展开。Non认为,成功的设计符合人性,他指出管理“40个代理确实很多”。他观察到用户创建了“幕僚长代理”来整合管理,这表明了将活动分组的自然倾向。然而,Claire强调了数据访问、权限和分段内存等实际的战术考量,她总结道,“用例在此类决策中至关重要。”
关于产品构建者所需的能力,两人都同意经典产品经理的特质,如用户同理心和系统思维,如今这些特质被应用于一个新的技术宇宙中。Claire补充说,需要“不懈地尝试和持续迭代”并忍受“很多痛苦”,而Non则强调了从这些反馈循环中学习的能力。
讨论转向了更广泛的生态系统和生产力工具。Claire将ChatGPT描述为一个提供多层功能的平台,从原生功能到第三方插件,以及一个“计算机使用”的备用方案。这种分层方法确保了如果一个工具不起作用,用户仍然有途径完成他们的任务。Non强调了“完成99%”的产品与“完成全部任务”的产品之间的区别,他赞扬“计算机使用”能够始终“完全完成”任务的能力,即使速度较慢。他将理想的用户体验比作成为“家中的客人”,能预见到所有需求。
Claire,作为OpenAI研究部门的新成员,分享了她的学习曲线。她建议提供具体的用例、清晰的用户目标,最重要的是,学习编写“评估(evals)”以推动模型训练的迭代循环。
关于规划的时间范围,Claire坚持她“两到三个月”的理想期限,警告不要进行更长期的预测,因为它们固有的不准确性,同时强调需要为未来而非仅仅为当下构建。
在讨论经典产品经理技艺与新入门门槛时,Non强调,由于“能力过剩”以及用户需要吸收新产品,现在“入职引导”变得尤为关键。Claire补充说,用户理解“隐私和他们的数据”以及确保半自主代理行为的可预测性变得更加重要。她还指出,经典的产品经理技能仍然适用,但“通过经验检验事物和更快行动的重要性只增不减”。两人都同意,“直接的用户关系”,或者成为一个“可私信的产品经理”,变得越来越重要,特别是在理解细微的用户反馈方面,比如代理表现出“态度”等。
最后,对于2027年的预测:Claire“非常看好语音”,她引用了其自然界面和改进的模型能力。Non预测,“自动驾驶”产品将会兴起,它们能够利用自身解决“空白输入框问题”并提供温和的入职引导。主持人Claire预测将出现“取代我们随身携带笔记本电脑的硬件”,她构想了一个“婴儿机器人”。
The discussion kicked off with Claire from OpenAI acknowledging the challenge of shipping imperfect products, a stark contrast to her previous experience at Stripe where deep polish was the norm. She emphasized that in the current AI era, urgency is paramount, and empirical evidence from users iterating on products outweighs theoretical perfection. The "toggle," an initial agentic harness for ChatGPT's billion-plus users, exemplifies this: an imperfect solution deemed necessary for immediate functionality and future iteration.
Non then delved into the decision of what to deprecate versus what to keep. He highlighted the importance of a "coherent story" for users, allowing them to follow the product's evolution. If users understand the journey, they are more accepting of changes, even if it means building and discarding a lot of code or product along the way.
The conversation then shifted to the constraints that maintain a high-quality bar despite rapid development. Claire outlined three key elements: first, ensuring the product is "additive for users" by unlocking genuine value; second, meeting an "internal bar" where products are trialed internally for uptake, delight, and novel use cases; and third, aiming for where the models will be in "two to three months," avoiding being overly anchored in the present or too futuristic to be usable. Non added that the biggest constraint is users' "understanding and ability to absorb" new capabilities, noting a "capability overhang" where models can do more than users currently leverage.
They addressed the common belief that B2B customers cannot absorb change. Claire argued that while the pace might feel "beyond breakneck," not shipping frontier advancements risks enterprises being "leapfrogged." She used the example of the agent revolution, where enterprises needed agents fast, even if it broke existing processes, to unlock greater value. This reinforces the truism of building what users *need*, not just what they *say* they need.
A debate emerged around single versus multi-identity agents. Non suggested that successful designs map onto human nature, noting that managing "40 agents is quite a lot." He observed users creating "chief of staff agents" to consolidate management, indicating a natural inclination towards grouping activities. Claire, however, emphasized the practical, tactical considerations like data access, permissions, and segmented memory, concluding that the "use case matters so much in deciding this."
Regarding the necessary skill sets for product builders, both agreed on classic PM traits like user empathy and systems thinking, now applied to a new technological universe. Claire added "a relentlessness to like try it out and keep iterating" and enduring "a lot of pain," while Non highlighted the ability to learn from those feedback loops.
The discussion moved to the broader ecosystem and tools for productivity. Claire described ChatGPT as a platform offering layers of capability, from native functions to third-party plugins, and a "computer use" fallback. This layered approach ensures that if one tool doesn't work, users still have a path to accomplish their tasks. Non underscored the difference between a product that gets "99% there" and one that finishes the job, praising "computer use" for its ability to always get tasks "all the way done," even if it's slower. He likened the ideal user experience to being a "guest in your home," anticipating all needs.
Claire, new to working with research at OpenAI, shared her learning curve. She advised providing specific use cases, clear user goals, and, most importantly, learning to write "evals" to drive the iterative loop of model training.
On the topic of time horizons for planning, Claire maintained her ideal of "two to three months out," cautioning against longer-term predictions due to their inherent inaccuracy, while also stressing the need to build for the future, not just the present.
Closing with classic PM craft versus new table stakes, Non highlighted "onboarding" as significantly more critical now due to the "capability overhang" and the need for users to absorb new offerings. Claire added the increased importance of users understanding "privacy and their data" and ensuring predictable behavior from semi-autonomous agents. She also noted that classic PM skills still apply, but "the importance of testing things empirically and moving faster has only increased." Both agreed that "direct user relationships," or being a "dmable PM," has become increasingly vital, especially for understanding subtle user feedback like an agent having "attitude."
Finally, for predictions for 2027: Claire is "so bullish on voice," citing its natural interface and improved model capabilities. Non predicted the rise of "self-driving" products that can use themselves to solve the "empty input box problem" and provide gentle onboarding. The host, Claire, predicted "hardware that replaces carrying around our laptop," envisioning a "baby bot."
摘要
How do you build AI products when the technology and users’ expectations are changing so quickly? At the Lenny and Friends Summit, OpenAI’s Tara Seshan and Nan Yu join Claire Vo to discuss shipping imperfect features, designing agents people can understand, and deciding how those agents should work with other tools. They explain why rapid testing, direct conversations with users, and completing the last mile of a task matter more than ever.
Recorded live at Lenny and Friends Summit on September 10, 2026, in San Francisco.
GPT-4正在为你翻译摘要中......
