Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem
发布时间 来源
Episode 设置
以下是内容的中文翻译:
CHI Discovery 由 Josh 和 Matt 共同创立,正走在利用人工智能将药物发现从试错过程转变为工程学科的最前沿。他们的核心使命是创建一个“生物学基础模型实验室”,这类似于大型语言模型 (LLM) 彻底改变了代码生成的方式。他们展望未来,分子(尤其是药用分子)将通过精准设计而诞生,而非偶然“发现”。
创始人强调药物开发领域正在发生边界转移:现在越来越多的部分可以通过工程化实现,而不再仅仅依赖于现实世界的测试。这一范式转变随着深度学习的出现获得了显著的势头。历史上,像每两年一次的学术团队预测蛋白质结构竞赛,其性能在大约 2018 年出现了“巨大飞跃”,最终在 2020 年诞生了 AlphaFold 2。最初,深度学习模型预测氨基酸之间的距离,然后将其转化为蛋白质结构。这一进展从单纯预测结构演变为设计能折叠成所需形状的序列,从而越来越接近药物设计问题。
2021-2022 年,随着扩散模型的兴起,迎来了一个关键时刻。与之前试图压缩输入分布(如变分自编码器)的生成方法不同,扩散模型通过迭代优化嘈杂、破损的结构,在多个步骤中学习“使其稍微更好”。这种方法被证明对生物学领域非常有效,它使得蛋白质结构和序列能够同时生成,并结合了越来越真实的“提示”和现实世界约束。
2024 年成立 CHI Discovery 的决定,源于一项突破:即预测和设计抗体的能力。在此之前,由于数据有限,这长期以来一直被认为是一个过于困难的问题。这项成就,在新的人工智能架构(如扩散模型)的推动下,为解决具有治疗意义的蛋白质问题打开了大门。
CHI 的核心指导原则是“简洁”。像他们的 CHI-1 这样拥有 23 个不同子模块的复杂模型,迭代起来会变得困难。通过简化并识别核心要素,扩展和改进的研究过程变得更加清晰。这种精神也延伸到他们的模型构建中:他们从头开始创建所有模型,而不是微调现有模型,这种理念他们称之为“扩展计算、数据和模型”的“惨痛教训”。
他们的跨学科团队反映了这一复杂的挑战。该团队由人工智能研究员、化学家和生物学家(一个“四语”团队)组成,其中包括抗体工程师 Andy Young 和来自 David Baker 实验室的蛋白质设计专家 Nathan Rollins 等杰出人才。尽管背景各异,他们对高质量代码库的关注突显了其作为一家提供人工智能驱动解决方案的“软件公司”的身份。
CHI 的模型在分子质量方面取得了显著进展。当他们开始时,*从头*抗体设计的最新技术结合率仅为 0.1%。凭借他们的 CHI-2 模型,这一成功率跃升至 15%,为迭代提供了具有统计学意义的数据,并使他们能够从一开始就“融入”类药性。这一成功不仅是孤立的突破,还在于将扩展定律应用于生物学,理解如何将生物序列“标记化”,以及利用庞大的蛋白质结构(如蛋白质数据库)和序列数据库来持续改进他们的模型。基于结构和基于序列(如 Josh 在 ESM 上的工作)方法之间的相互作用是关键,两者相互启发和增强。
他们认为生物学是一个高度可验证的领域。尽管代码生成可能难以处理“混乱代码”等定性方面的问题,但分子设计可以实现特定的、可测量的特性(结合力、可制造性)。CHI 内部的不同团队对“硬目标”的定义各不相同——从科学家关注的不可成药的 GPCRs,到机器学习研究人员关注的复杂膜蛋白——这种多学科的推动力促进了平稳、持续的进展。
CHI 的最终目标是构建一个“分子计算机辅助设计 (CAD) 套件”。该套件将使研究人员能够预先指定设计原则,从而将从构思到可测试假设的循环时间从数月大幅缩短至数周甚至数天,彻底改变更优质药物的发现。
展望未来 10-20 年,创始人预见药物将更具意图性,能为特定疾病量身定制,且副作用更少。由于更快、更便宜的药物开发,以前被认为过于罕见或难以靶向的疾病将变得可治疗。他们选择了一种“基础设施”业务模式,与礼来 (Eli Lilly)、诺华 (Novartis)、Argenx 和辉瑞 (Pfizer) 等主要制药公司合作,而不是自己开发药物。这一战略使他们能够将影响力扩展到整个行业,确保他们的模型在实际项目中得到严格验证和部署。这种竞争格局要求他们的模型“真正有效”,从而创造了一种激励结构,使更优的模型能够直接转化为合作伙伴乃至患者的价值。
尽管扩展 GPU 基础设施和维护生产级代码存在复杂性,但团队受到一个共同的、使命驱动的目标的驱使:通过创造突破性药物来深刻改变世界。CHI 这个名字代表着“化学与人工智能”,突显了他们简洁而强大的愿景。他们最兴奋的是看到他们的模型在患者治疗中取得切实的成果,并继续保持快速的进步,这可能很快就会看到数十甚至数百种由 CHI 设计的分子进入临床试验。
CHI Discovery, co-founded by Josh and Matt, is at the forefront of transforming drug discovery from a trial-and-error process into an engineering discipline using AI. Their core mission is to create a "foundation model lab for biology," akin to how large language models (LLMs) have revolutionized code generation. They envision a future where molecules, especially for medicines, are designed with precision rather than "discovered" serendipitously.
The founders highlight a shifting boundary in drug development: an increasing proportion can now be engineered rather than solely tested in the real world. This paradigm shift gained significant momentum with the advent of deep learning. Historically, protein folding competitions, like the biannual event where academic groups predicted protein structures, saw a "big step change" in performance around 2018, culminating in AlphaFold 2 in 2020. Initially, deep learning models predicted amino acid distances, which were then rendered into protein structures. This evolved from simply predicting structures to designing sequences that would fold into desired shapes, moving closer to the drug design problem.
A pivotal moment came in 2021-2022 with the rise of diffusion models. Unlike previous generative approaches that tried to compress input distributions (like variational autoencoders), diffusion models iteratively refine noisy, broken structures, learning to "make it slightly better" over many steps. This approach proved remarkably effective for biology, allowing for the simultaneous generation of protein structure and sequence with increasingly realistic "prompts" and real-world constraints.
The decision to start CHI Discovery in 2024 was driven by a breakthrough: the ability to predict and design antibodies, a problem long considered too hard due to limited data. This achievement, fueled by new AI architectures like diffusion models, opened the door to tackling therapeutically significant proteins.
A central guiding principle at CHI is "simplicity." With complex models like their CHI-1 having 23 distinct sub-modules, iterating becomes difficult. By simplifying and identifying core elements, the research process for scaling and improvement becomes much clearer. This ethos extends to their model building: they create all models from scratch rather than fine-tuning existing ones, a philosophy they call the "bitter lesson" of scaling compute, data, and models.
Their interdisciplinary team reflects this complex challenge. Comprising AI researchers, chemists, and biologists (a "quadrilingual" group), the team includes luminaries like antibody engineer Andy Young and protein design expert Nathan Rollins from David Baker's lab. Despite diverse backgrounds, their focus on a high-quality codebase underlines their identity as a "software company" delivering AI-driven solutions.
CHI's models have demonstrated remarkable progress in molecule quality. When they started, the state-of-the-art for *de novo* antibody design had a mere 0.1% binding rate. With their CHI-2 model, this success rate jumped to 15%, providing statistically significant data for iteration and allowing them to "bake in" drug-like properties from the start. This success is not just about isolated breakthroughs but also about applying scaling laws to biology, understanding how to "tokenize" biological sequences, and leveraging vast databases of protein structures (Protein Data Bank) and sequences to continuously improve their models. The interplay between structure-based and sequence-based (like Josh's work on ESM) approaches is key, with each informing and enhancing the other.
They believe biology is a highly verifiable domain. While code generation might struggle with qualitative aspects like "sloppy code," molecule design allows for specific, measurable properties (binding, manufacturability). Different teams within CHI define "hard targets" differently – from undruggable GPCRs for scientists to complex membrane proteins for ML researchers – and this multidisciplinary push drives smooth, continuous progress.
CHI's ultimate goal is to build a "computer-aided design (CAD) suite for molecules." This suite would enable researchers to specify design principles upfront, drastically accelerating the loop from idea to testable hypothesis from months to weeks or even days, thereby revolutionizing the discovery of better medicines.
Looking 10-20 years ahead, the founders envision medicines with increased intentionality, tailored for specific diseases with fewer side effects. Diseases previously deemed too rare or difficult to target could become addressable due to faster, less expensive drug development. They chose an "infrastructure" business model, partnering with major pharmaceutical companies like Eli Lilly, Novartis, Argenx, and Pfizer, rather than developing drugs themselves. This strategy allows them to scale their impact across the industry, ensuring their models are rigorously validated and deployed on real programs. This competitive landscape demands that their models "really work," creating an incentive structure where better models translate directly into value for partners and, ultimately, patients.
Despite the complexities of scaling GPU infrastructure and maintaining production-level code, the team is driven by a shared, mission-driven goal: to make a profound difference in the world by creating breakthrough medicines. The name CHI, representing "Chemistry and AI," underscores their simple yet powerful vision. They are most excited about seeing their models lead to tangible results in patient treatments and continuing the rapid pace of progress that could soon see dozens, if not hundreds, of CHI-designed molecules entering clinical trials.
摘要
Most people treat biology as a bespoke, messy science. Josh Meier and Matt McPartlon, co-founders of Chai Discovery, treat it as an engineering problem. They make the case that drug design obeys the bitter lesson: scale data, models, and compute, and the model can learn what a hand-built pipeline simply couldn't capture. The results are concrete: Chai-2 pushed de novo antibody design from a sub 0.1% hit rate to 16%, turning a needle-in-a-haystack search into something more like designing a key to fit a lock. Josh argues, counterintuitively, that biology is more verifiable than code, and explains why the goal should be more lab experiments, not fewer. Their bet: a design suite that collapses drug discovery from nine months to nine days, and arms the pharma industry rather than competing with it.
Hosted by Pat Grady and Sonali Singh, Sequoia Capital
00:00 Introduction
01:52 From Discovery to Design
03:25 Protein AI Breakthroughs Timeline
06:04 Why Start in 2024
10:13 Diffusion Models Intuition
11:41 Building the Avengers Team
15:22 Hit Rates and Scaling Laws
25:01 Molecular CAD Vision
25:24 Faster Design Loops
26:32 Future Drug Discovery
28:37 Platform Business Model
31:14 Partnering Reality Check
33:44 Data Flywheel Explained
37:16 Staying Ahead at Scale
39:44 Culture and What's Next
GPT-4正在为你翻译摘要中......
