How to Build Long-Horizon AI Agents — Mitch Troyanovsky, Basis
发布时间 来源
Episode 设置
摘要
AI agents can write code for hours, but ask them to do real work in the real economy, and they break. Mitch Troyanovsky is co-founder of Basis, a unicorn AI company whose agents run autonomously for hours — sometimes days — completing complex tax returns end to end. His answer to the reliability problem: stop grading outcomes, and start supervising the process.
This is a definitive, reference-style conversation on building long-horizon AI agents. Mitch walks through the full history — from ReAct and the AutoGPT crash to reasoning models and RLVR — and explains why the industry abandoned process supervision in 2023, and why it's now coming back at a completely different scale. We go deep on behavior specs, the open standard Basis just released with Braintrust for defining and evaluating how agents behave across entire trajectories, with no ground truth required.
Along the way: why context is really runtime training data, why your documentation must be treated like a codebase, ontologies as "worlds for agents to live in," the judge-as-agent architecture, why Basis hires philosophy majors as Language Architects, deploying agents as "onboarding 300 brilliant alien employees," and Mitch's prediction for when the bitter lesson swallows the harness.
Mitchell Troyanovsky
LinkedIn - https://www.linkedin.com/in/mitchelltroyanovsky
X/Twitter - hhttps://x.com/mitch_troy
Basis
Website - https://www.getbasis.ai
X/Twitter - https://x.com/trybasis
Matt Turck (General Partner)
Blog - https://mattturck.com
LinkedIn - https://www.linkedin.com/in/turck/
X/Twitter - https://x.com/mattturck
FirstMark Capital
Website - https://firstmark.com
X/Twitter - https://x.com/FirstMarkCap
Listen on:
Spotify - https://open.spotify.com/show/7yLATDSaFvgJG80ACcRJtq
Apple - https://podcasts.apple.com/us/podcast/the-mad-podcast-with-matt-turck/id168623872
Watch the Data Driven NYC Session with Mithcell Troyanovsky - https://youtu.be/xH_b6iwrASc
01:09 Why Basis Engineers Whisper to Their Agents
04:12 Accounting as Compression: an Intelligence Layer Over the Economy
06:11 Defining Long-Horizon: When You Exceed the Context Window
08:24 Anatomy of a Multi-Day Autonomous Trajectory
10:19 Handoff Design: Optimizing Output for the Reviewer
11:17 ReAct and Why Reasoning Must Regulate Its Own State
12:33 Large Working Memory, No Long-Term Memory
14:13 Compounding Errors: Why AutoGPT and BabyAGI Broke
15:51 Opus 3, o1, o3: the Three Real Paradigm Shifts
17:07 Titrating Inference Compute Across Easy and Hard Steps
18:23 Process Reward vs. Outcome Reward: "Let's Verify Step by Step"
20:32 RLVR and Why the METR Curve Overstates Reliability
22:09 Verifiable at Runtime: the Real Reason Coding Won
25:14 No Ground Truth, No Cheap Verification, No Data
26:55 Encoding Deterministic Checks From Human Review Process
29:18 Synthetic Data Limits: Generating Artifacts, Not Text
33:16 100 Evals Pass — Does It Generalize to Production?
35:53 Primary Sources vs. Pre-Training Knowledge
36:37 Behavior Specs: Markdown, Judges, and True/False/N.A.
39:58 Specificity vs. Brittleness in Spec Authoring
42:18 Context as Runtime Training Data
44:21 Judge-as-Agent: Trajectory Maps and Sub-Agent Attribution
46:45 The Move 37 Objection: Reliability Over Optimality
50:02 The Magic Box Model: Building Without Weights Access
52:41 "Nothing Paradigm-Shifting Has Changed Since o3"
54:56 Open-Sourcing the Behavior Spec Standard With Braintrust
01:02:54 Ontology Design: Virtual Filesystems, Graphs, Embeddings
01:04:20 Canonical vs. Non-Canonical: Docs as Codebase
01:06:33 Language Architects and Writing for Runtime Interpretation
01:09:05 Deployed Intelligence: 300 Alien Employees With No Context
01:11:10 Closing the Loop: Signal → Context, Tools, Harness
01:12:50 Context Slop: the Mistake Most Agent Builders Make
01:14:29 Reward Function Design and Credit Assignment Over Trajectories
01:17:01 Will the Bitter Lesson Swallow the Harness?
01:18:46 Business Moats vs. Technical Moats
01:21:03 Paradigm Thinking Over Timeline ADHD
GPT-4正在为你翻译摘要中......
