Solving the Hardest Problem in Robotics | World Labs with a16z
发布时间 来源
Episode 设置
World Labs,一家前沿人工智能实验室,正在引领开发“空间智能”——一种全新的人工智能范式,该范式能够生成、理解、推理并与物理及虚拟空间进行交互。这一愿景的核心在于构建大型世界模型。该公司最近宣布收购了 Cynix,该公司由李飞飞的前博士后研究员允舒共同创立。此次整合被视为关键一步,旨在使 World Labs 的人工智能能够通过机器人技术在物理空间中行动。
Cynix 的核心创新是一种“真实到模拟再到真实”(real-to-sim-to-real)的管道。这项技术能将真实环境精确映射到高度对齐的数字孪生体中,确保模拟事件能紧密反映真实世界的成果。这种能力直接解决了机器人技术的一个主要瓶颈:即用于训练和评估的可扩展、高质量数据的稀缺性,这与语言模型所拥有的海量数据形成了鲜明对比。
World Labs 和 Cynix 之间的协同效应是深远的。World Labs 的生成模型 Marble(Cynix 之前已是其客户)能够根据各种提示词创建几何一致的 3D 世界。Cynix 利用这一点进行对机器人技术至关重要的环境密集重建和动态建模。它们共同旨在构建一个“一致的世界”——在空间、时间、视角和交互方面保持一致——这对可靠的机器人操作至关重要。
支撑他们方法的关键哲学洞察是模拟的关键作用,它不仅在于数据生成,更在于“反事实推理”。人类使用心理模拟来探索假设情景,并从真实世界中尚未发生或不可能发生的事件中学习。这种能力对机器人技术来说是不可或缺的,因为现实世界中的数据收集往往成本高昂、耗时且危险。李飞飞指出,Waymo 为自动驾驶汽车进行的数十亿小时模拟,证明了模拟的重要性,即使是对“最简单”的机器人来说也是如此。
允舒强调了模拟的两个主要优势:可靠性和效率。模拟允许对环境变量进行系统性随机化(如光照、摩擦、几何形状、物体类型),从而提供状态空间的全面覆盖,以训练出强大的机器人策略。在效率方面,模拟能够实现超越人类实时速度的加速训练,并考虑到动态变化,这对于机器人更快、更有效地运行至关重要。
他们的平台旨在实现“具身无关”(embodiment agnostic)和“模型无关”(model agnostic)。这意味着它能够与各式各样的机器人硬件(从单臂机器人到移动机械手)集成,并促进各种人工智能模型的训练,包括视觉-语言-动作模型或世界-动作模型。目标是为机器人提供底层基础设施,以便它们在数字世界中可靠地学习和评估,最终将这种能力转化为现实世界中的部署。
该团队采取务实策略,最初专注于“半结构化环境”(例如仓库、受控生产线),然后再解决完全“非结构化环境”(如家庭环境)的复杂性。他们承认,虽然人形机器人是为通用非结构化环境设计的,但针对特定任务的专业机器人通常是更可行、更高效的商业方案。李飞飞和允舒都认为,机器人在功耗方面达到人类水平是一个遥远的目标,并强调即使是先进的人工智能模型目前也对能源有很高的需求。
World Labs 和 Cynix 正在规划一次深思熟虑的整合,充分利用彼此的互补优势。他们已经收到了大量来自寻求其技术用于训练和评估的机器人公司的咨询,尤其是一些即将把机器人部署到实际高价值任务中的客户。他们鼓励所有机器人公司与 World Labs 互动,探索其空间智能基础设施的潜在应用。
World Labs, a frontier AI lab, is spearheading the development of "spatial intelligence," a new paradigm for AI that can generate, understand, reason with, and interact with spaces, both physical and virtual. Central to this vision is building large world models. The company recently announced the acquisition of Cynix, co-founded by Yunshu, a former postdoctoral researcher of Fei-Fei Li. This integration is seen as a crucial step in enabling World Labs' AI to act within physical spaces through robotics.
Cynix's core innovation is a "real-to-sim-to-real" pipeline. This technology accurately maps real environments into highly aligned digital twins, ensuring that simulated events closely mirror real-world outcomes. This capability directly addresses a major bottleneck in robotics: the scarcity of scalable, high-quality data for training and evaluation, a stark contrast to the abundant data available for language models.
The synergy between World Labs and Cynix is profound. World Labs' generative model, Marble (which Cynix was already using as a customer), can create geometrically consistent 3D worlds from various prompts. Cynix leverages this for dense reconstruction and dynamic modeling of environments essential for robotics. Together, they aim to build a "consistent world"—consistent across space, time, viewpoints, and interactions—which is foundational for reliable robotic operation.
A key philosophical insight underpinning their approach is the crucial role of simulation, not just for data generation, but for "counterfactual reasoning." Humans use mental simulations to explore hypothetical scenarios and learn from events that haven't or couldn't happen in the real world. This capability is indispensable for robotics, where real-world data collection is often costly, slow, and dangerous. Fei-Fei Li points to Waymo's billions of simulation hours for self-driving cars as an example of simulation's importance, even for the "simplest" robots.
Yunshu highlights two primary benefits of simulation: reliability and efficiency. Simulation allows for systematic randomization of environmental variables—lighting, friction, geometry, object types—providing comprehensive coverage of state space to train robust robotic policies. For efficiency, simulation enables accelerated training beyond real-time human speed, accounting for dynamic changes that are critical for robots to operate faster and more effectively.
Their platform is designed to be "embodiment agnostic" and "model agnostic." This means it can integrate with diverse robotic hardware, from single arms to mobile manipulators, and facilitate the training of various AI models, including vision-language-action models or world-action models. The goal is to provide the underlying infrastructure for robots to learn and evaluate reliably in digital worlds, ultimately translating that capability to real-world deployment.
The team adopts a pragmatic strategy, focusing initially on "semi-structured environments" (e.g., warehouses, controlled manufacturing lines) before tackling the complexities of fully "unstructured environments" (like homes). They acknowledge that while humanoids are designed for general unstructured environments, specialized robots for specific tasks are often a more viable and efficient business approach. Both Fei-Fei and Yunshu agree that achieving human-level power efficiency in robots is a distant goal, emphasizing the current energy demands of even advanced AI models.
World Labs and Cynix are planning a thoughtful integration, leveraging their complementary strengths. They are already experiencing significant inbound interest from robotics companies seeking their technology for both training and evaluation, particularly from clients nearing deployment in practical, high-value-adding tasks. They encourage all robotics companies to engage with World Labs to explore potential applications of their spatial intelligence infrastructure.
摘要
Last week, World Labs announced its acquisition of SceniX, bringing together two teams working on one of AI's biggest unsolved problems: how to give machines a true understanding of the physical world.
Martin Casado sits down with Fei-Fei Li, co-founder and CEO of World Labs, creator of ImageNet, and pioneer of spatial intelligence, alongside Yunzhu Li, co-founder of SceniX and assistant professor at Columbia University. They discuss why World Labs acquired SceniX, how simulation can unlock the next generation of robotics, and why training robots may require a fundamentally different approach than training language models.
The conversation explores real-to-sim-to-real pipelines, world models, robotics foundation models, evaluation, synthetic data, and why the future of AI depends not just on understanding language—but on understanding and interacting with the physical world.
Timestamps:
00:00 - Intro
01:08 - World Labs & SceniX
06:04 - Marble & the Data Bottleneck in Robotics
07:13 - How the Two Teams Come Together
10:55 - Building a Foundation Model for Robotics
12:35 - Video Models vs Real-to-Sim-to-Real
19:38 - Why Simulation is Essential for Robot Learning
23:01 - Training, Evaluation & Real Customer Use Cases
29:17 - Humanoids, Semi-Structured Environments & the Grand Challenge
36:56 - Integration Plans & What Success Looks Like in Two Years
Resources:
Follow Fei-Fei Li on X: https://x.com/drfeifei
Follow Yunzhu Li on X: https://x.com/YunzhuLiYZ
Follow Martin Casado on X: https://x.com/martin_casado
Stay Updated:
If you enjoyed this episode, be sure to like, subscribe, and share with your friends!
Find a16z on X: https://twitter.com/a16z
Find a16z on LinkedIn: https://www.linkedin.com/company/a16z
Listen to the a16z Show on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX
Listen to the a16z Show on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711
Follow our host: https://x.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see http://a16z.com/disclosures.
GPT-4正在为你翻译摘要中......
