How to Build Long-Horizon AI Agents — Mitch Troyanovsky, Basis

发布时间    来源
Episode 设置


登录已过期或未登录,无法修改。请先登录后再试。

摘要

AI agents can write code for hours, but ask them to do real work in the real economy, and they break. Mitch Troyanovsky is co-founder of Basis, a unicorn AI company whose agents run autonomously for hours — sometimes days — completing complex tax returns end to end. His answer to the reliability problem: stop grading outcomes, and start supervising the process. This is a definitive, reference-style conversation on building long-horizon AI agents. Mitch walks through the full history — from ReAct and the AutoGPT crash to reasoning models and RLVR — and explains why the industry abandoned process supervision in 2023, and why it's now coming back at a completely different scale. We go deep on behavior specs, the open standard Basis just released with Braintrust for defining and evaluating how agents behave across entire trajectories, with no ground truth required. Along the way: why context is really runtime training data, why your documentation must be treated like a codebase, ontologies as "worlds for agents to live in," the judge-as-agent architecture, why Basis hires philosophy majors as Language Architects, deploying agents as "onboarding 300 brilliant alien employees," and Mitch's prediction for when the bitter lesson swallows the harness. Mitchell Troyanovsky LinkedIn - https://www.linkedin.com/in/mitchelltroyanovsky X/Twitter - hhttps://x.com/mitch_troy Basis Website - https://www.getbasis.ai X/Twitter - https://x.com/trybasis Matt Turck (General Partner) Blog - https://mattturck.com LinkedIn - https://www.linkedin.com/in/turck/ X/Twitter - https://x.com/mattturck FirstMark Capital Website - https://firstmark.com X/Twitter - https://x.com/FirstMarkCap Listen on: Spotify - https://open.spotify.com/show/7yLATDSaFvgJG80ACcRJtq Apple - https://podcasts.apple.com/us/podcast/the-mad-podcast-with-matt-turck/id168623872 Watch the Data Driven NYC Session with Mithcell Troyanovsky - https://youtu.be/xH_b6iwrASc 01:09 Why Basis Engineers Whisper to Their Agents 04:12 Accounting as Compression: an Intelligence Layer Over the Economy 06:11 Defining Long-Horizon: When You Exceed the Context Window 08:24 Anatomy of a Multi-Day Autonomous Trajectory 10:19 Handoff Design: Optimizing Output for the Reviewer 11:17 ReAct and Why Reasoning Must Regulate Its Own State 12:33 Large Working Memory, No Long-Term Memory 14:13 Compounding Errors: Why AutoGPT and BabyAGI Broke 15:51 Opus 3, o1, o3: the Three Real Paradigm Shifts 17:07 Titrating Inference Compute Across Easy and Hard Steps 18:23 Process Reward vs. Outcome Reward: "Let's Verify Step by Step" 20:32 RLVR and Why the METR Curve Overstates Reliability 22:09 Verifiable at Runtime: the Real Reason Coding Won 25:14 No Ground Truth, No Cheap Verification, No Data 26:55 Encoding Deterministic Checks From Human Review Process 29:18 Synthetic Data Limits: Generating Artifacts, Not Text 33:16 100 Evals Pass — Does It Generalize to Production? 35:53 Primary Sources vs. Pre-Training Knowledge 36:37 Behavior Specs: Markdown, Judges, and True/False/N.A. 39:58 Specificity vs. Brittleness in Spec Authoring 42:18 Context as Runtime Training Data 44:21 Judge-as-Agent: Trajectory Maps and Sub-Agent Attribution 46:45 The Move 37 Objection: Reliability Over Optimality 50:02 The Magic Box Model: Building Without Weights Access 52:41 "Nothing Paradigm-Shifting Has Changed Since o3" 54:56 Open-Sourcing the Behavior Spec Standard With Braintrust 01:02:54 Ontology Design: Virtual Filesystems, Graphs, Embeddings 01:04:20 Canonical vs. Non-Canonical: Docs as Codebase 01:06:33 Language Architects and Writing for Runtime Interpretation 01:09:05 Deployed Intelligence: 300 Alien Employees With No Context 01:11:10 Closing the Loop: Signal → Context, Tools, Harness 01:12:50 Context Slop: the Mistake Most Agent Builders Make 01:14:29 Reward Function Design and Credit Assignment Over Trajectories 01:17:01 Will the Bitter Lesson Swallow the Harness? 01:18:46 Business Moats vs. Technical Moats 01:21:03 Paradigm Thinking Over Timeline ADHD

GPT-4正在为你翻译摘要中......

中英文字稿