首页  >>  来自播客: Sequoia Capital 更新   反馈  

Sequoia Capital - Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil

发布时间:   原节目
Jerry,前OpenAI副总裁,以及Rohan,谷歌大脑和Gemini的关键人物,共同创立了Core Automation,怀揣着一个宏大的愿景:取代Transformer架构,开创一个真正能从经验中学习的人工智能新时代。他们认为,尽管Transformer对当前的人工智能热潮具有奠基性作用,但其局限性使得彻底的变革成为必然。 Jerry承认Transformer的深远影响和经济价值,它使得可扩展的预训练和强化学习(RL)成为可能。然而,他认为该领域目前专注于渐进式改进(即让它们更便宜/更高效),而非从根本上更强大或更具表现力的发展。他相信,核心瓶颈是架构本身。Transformer尽管能力强大,但在现实世界的适应性方面表现不佳。它们在受控的实验室环境中进行训练,但在部署时却面临“更混乱的”数据分布。它们的“上下文学习”(in-context learning)在范围和持续时间上都很有限(例如,Codecs只有20分钟),而持续微调则面临灾难性遗忘和数据效率低下的问题。Jerry强调,当前的人工智能系统是有效的“人-LLM混合体”,但未能达到真正的通用人工智能(AGI)——一种能够在无需人类干预的情况下自我改进的模型。他强调“从经验中学习”比单纯的强化学习范围更广,他将足球(类似强化学习)与深层数学推理进行类比。Core Automation旨在开发新的算法,以架构的方式呈现,使模型能够在测试时、基于真实世界用户数据、在更长的时间跨度内持续学习和适应。 Rohan表示赞同,指出Transformer尽管在预训练(数据压缩)方面效率很高,但在推理方面却效率低下,尤其是因为其“浅层计算深度”和逐令牌生成的特点。他认为思维链推理(chain-of-thought reasoning)是一种效率低下的权宜之计,旨在弥补这种架构上的局限性。Rohan提倡一种整体方法,将预训练和强化学习相结合,以创建更高效的学习过程。他强调优化方法的关键作用,认为更强大的优化器可以释放更复杂、更深层架构的潜力,而这些架构目前很难训练。他还将当前的数字硬件与生物学习进行对比,暗示要实现显著的效率提升可能需要更具模拟特性、更注重硬件的设计。 两位创始人解释道,他们决定创办Core Automation源于他们在大型实验室中观察到的停滞,这些实验室目前由于竞争压力专注于现有Transformer技术的短期扩展。相比之下,Core Automation旨在成为“世界上最自动化的实验室”,旨在促进快速实验并突破架构研究的界限。他们计划自动化“内核生成”(即优化硬件计算的低级代码)等瓶颈,这目前需要稀有的人类专业知识和大量时间。例如,优化QR内核可以带来60倍的加速,但超出了LLM的当前能力。通过大幅加速迭代周期——可能每天执行几十甚至几百个实验——他们相信他们可以有效地探索巨大的可能架构空间。他们的最终目标是找到一种能够展现有意义的、长期适应性的“卓越架构”,甚至可能通过观察他们自己的模型每天在Core Automation科学家的工作中变得更出色来实现。

Jerry, former VP at OpenAI, and Rohan, a key figure at Google Brain and Gemini, have founded Core Automation with a bold vision: to replace the Transformer architecture and pioneer a new era of AI that truly learns from experience. They argue that while Transformers have been foundational to the current AI boom, their limitations necessitate a radical shift. Jerry acknowledges the profound impact and economic value of Transformers, enabling scalable pre-training and reinforcement learning (RL). However, he contends that the field is now focused on incremental improvements (making them cheaper/more efficient) rather than fundamentally more powerful or expressive. The core bottleneck, he believes, is the architecture itself. Transformers, despite their capabilities, struggle with real-world adaptability. They are trained in controlled lab environments but face "messier" distributions in deployment. Their "in-context learning" is limited in scope and duration (e.g., 20 minutes for Codecs), while continuous fine-tuning suffers from catastrophic forgetting and data inefficiency. Jerry highlights that current AI systems are effective "human-LLM hybrids" but fall short of true AGI – a model capable of improving itself without human intervention. He emphasizes that "learning from experience" is broader than just RL, drawing parallels between football (RL-like) and deep mathematical reasoning. Core Automation aims to develop new algorithms, expressed architecturally, that enable models to learn and adapt continuously at test time, on real-world user data, over much longer horizons. Rohan concurs, pointing out that Transformers, while efficient for pre-training (compressing data), are inefficient for inference, particularly due to their "shallow computational depth" and token-by-token generation. He sees chain-of-thought reasoning as an inefficient workaround to compensate for this architectural limitation. Rohan advocates for a holistic approach, combining pre-training and RL to create a more efficient learning procedure. He stresses the critical role of optimization methods, arguing that stronger optimizers can unlock the potential of more complex, deeper architectures that are currently difficult to train. He also contrasts current digital hardware with biological learning, suggesting that significant efficiency gains might require more analog, hardware-aware designs. Both founders explain their decision to start Core Automation stems from a perceived stagnation in larger labs, which are currently focused on short-term scaling of existing Transformer technology due to competitive pressures. Core Automation, in contrast, aims to be the "most automated lab in the world," designed to facilitate rapid experimentation and push the boundaries of architectural research. They plan to automate bottlenecks like "kernel generation" (the low-level code that optimizes computations on hardware), which currently requires rare human expertise and significant time. For example, optimizing a QR kernel can yield a 60x speedup but is beyond the current capabilities of LLMs. By dramatically accelerating the iteration cycle – potentially executing dozens or even hundreds of experiments daily – they believe they can efficiently search the vast space of possible architectures. Their ultimate goal is to find a "superior architecture" that demonstrates meaningful, long-term adaptability, perhaps even by observing their own models becoming better at the core automation scientists' work each day.