Why The Harness Matters More Than The Model | YC Paper Club

发布时间    来源
Episode 设置


登录已过期或未登录,无法修改。请先登录后再试。

摘要

Harnesses get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be farther from the truth. The same model weights that score 30% on ARC-AGI score 95% with a better harness. So we gathered a group of researchers and founders working at the frontier to do a deep dive into the state of harnesses. We'll cover how we got to this point, the case for making your harness as expressive as possible, and what YC learned building an agent for every employee in the company. Sign up for the next Paper Club: Alternative Compute Paradigms https://events.ycombinator.com/yc-paperclub-Sep23 Chapters: 00:00 - Why harnesses matter 04:27 - Building an auto-researcher by accident 07:13 - A five minute history of harnesses 13:56 - Self-improving harnesses 17:22 - Tonight's speakers 18:35 - Seth Karten: Prime Agent, a self-improving RLM harness 21:50 - Context as an L1, L2, L3 cache 24:51 - From Turing machine to von Neumann computer 28:33 - Messaging between agents 30:04 - ARC-AGI results 33:09 - Emulator Bench and GPU kernels 37:30 - Jon Saad-Falcon: OpenJarvis, personal AI on personal devices 38:26 - How far behind are local models 39:21 - The five primitives of a personal AI stack 42:47 - Letting cloud models optimize your local stack 43:53 - 800x cheaper than the cloud 45:58 - Josh France and Regan Bell: QM, YC's agent harness for work 47:29 - A history of YC's internal agents 49:24 - OpenClaw and a fleet of 50 agents 51:04 - Pulling the brain out of the sandbox 54:43 - Letting the agent choose its own sandbox and model 57:16 - The grind tool: budgets on goals 58:50 - Agents don't understand social context Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs

GPT-4正在为你翻译摘要中......

中英文字稿