RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor
发布时间 来源
Episode 设置
以下是原文的中文翻译:
Mercore的Brendan讨论了AI数据市场正在经历的变革性转变,即从低技能众包转向一个以高技能专家和复杂的强化学习(RL)环境为中心的“智能体时代”。Mercore凭借对更复杂数据解决方案的需求,实现了快速增长。
他将**RL环境**定义为一个用于AI模型评估和训练的三部分系统:
1. **世界(Worlds)**:现实、全面的场景,包含消息、文档、电子表格以及其他数字资料,模拟真实世界的项目。
2. **应用(Apps)**:Salesforce、Microsoft 365或Google Workspace等流行应用的高保真克隆,允许AI智能体通过API、CLI或UI进行真实交互。
3. **任务(Tasks)**:与强大的验证器(如评分标准或单元测试)配对的特定提示,规定了智能体的目标并衡量其性能。
Brendan解释说,根本挑战在于覆盖所有经济领域的人类任务的整个分布范围。Mercore的策略是利用广泛的人类专业知识,仅第二季度就记录了250万专家工时,这充分证明了这一点。人类至关重要,因为与数学等易于量化的领域不同,AI模型无法在复杂、开放式任务(例如起草幻灯片演示文稿或法律备忘录)中可靠地自我评估。专家设计的评分标准对于准确评估和有效的模型学习至关重要。
Brendan展示了一个法律RL环境作为例子。顶级律师事务所合作概述真实的法律项目场景,用相关信息填充详细的“数据室”。然后这些内容在克隆的应用环境(例如Google Workspace)中呈现。智能体被分配复杂的分析任务,而人类创建的、具有特定标准的评分标准确保了精确的验证并防止了“奖励作弊”(reward hacking),最终形成了排行榜和聚合模型分数。
这种方法的影响是巨大的。在“Apex Agents”等数据集上进行训练后(例如GLM 4.7在1800个任务上进行了50万次计算),模型表现出显著的性能提升和对其他基准测试的显著泛化能力。这项最初为前沿实验室开发的技术,现在正变得可供应用层公司使用,使他们能够“拥有自己的智能”——这是当今AI领域的一个关键差异化因素。
Mercore提供三种主要的数据策划模型:
1. **按任务(By Task)**:为特定客户需求量身定制的、高度复杂的任务,由于需要大量人类专家参与,通常价值很高(每个任务50到10,000美元)。
2. **现成数据(Off-the-shelf Data)**:预构建的高质量数据集,出售给多个客户,特别有利于寻求规模经济的新实验室。
3. **按小时计费的专家(Hourly Experts)**:按小时提供专家服务,尽管这种模式正逐渐被更结构化的数据产品所取代。
在问答环节,Brendan详细阐述了数据定价,它同时考虑了客户模型改进的目标和Mercore的成本结构。数据质量通过注重真实性(由专家概要保证)和验证器的准确性来维持,后者通常通过“轨迹分析”和人类反馈进行完善。他澄清说,虽然使用了“合成数据”(模型生成的轨迹或环境填充),但人类专业知识对于将AI能力推向当前AI的“前沿之外”仍然不可或缺。
展望未来,Brendan指出了RL环境的两个主要趋势:开发能够执行**超长周期任务**(需要数百甚至数千小时)的智能体,以及整合**虚拟同事**来训练智能体进行社交互动和协作,这是人类工作中一个关键但很大程度上未被衡量的方面。他最后重申,虽然AI提供帮助,但模型无法可靠地为前沿任务创建自己的准确评分标准,强调了人类专业知识在任务和验证器创建方面对竞争优势的独特价值。
Brendan from Mercore discussed the transformative shift in the AI data market, moving from low-skilled crowdsourcing to an "agentic era" focused on high-skilled experts and sophisticated Reinforcement Learning (RL) environments. Mercore has experienced rapid growth, fueled by this demand for more complex data solutions.
He defined an **RL environment** as a three-part system designed for AI model evaluation and training:
1. **Worlds:** Realistic, comprehensive scenarios encompassing messages, documents, spreadsheets, and other digital artifacts that mimic real-world projects.
2. **Apps:** High-fidelity clones of popular applications like Salesforce, Microsoft 365, or Google Workspace, allowing AI agents to interact authentically via APIs, CLI, or UI.
3. **Tasks:** Specific prompts paired with robust verifiers (rubrics or unit tests) that dictate the agent's objective and measure its performance.
The fundamental challenge, Brendan explained, is to span the entire distribution of human tasks across all economic sectors. Mercore's strategy involves leveraging extensive human expertise, evidenced by 2.5 million expert hours recorded in Q2 alone. Humans are crucial because, unlike in easily quantifiable domains like math, AI models cannot reliably self-evaluate in complex, open-ended tasks (e.g., drafting a slide deck or legal memo). Expert-designed rubrics are essential for accurate assessment and effective model learning.
Brendan showcased a legal RL environment as an example. Top law firms collaborate to outline real legal project scenarios, populating detailed "data rooms" with relevant information. These are then rendered within cloned application environments (e.g., Google Workspace). Agents are tasked with complex analyses, and human-created rubrics, with specific criteria, ensure precise verification and prevent "reward hacking," ultimately leading to leaderboards and aggregate model scores.
The impact of this approach is substantial. Post-training on datasets like "Apex Agents" (e.g., GLM 4.7 with 500k compute on 1800 tasks) demonstrated dramatic performance gains and significant generalization to other benchmarks. This powerful technology, initially developed for frontier labs, is now becoming accessible to application-layer companies, enabling them to "own their own intelligence" – a crucial differentiator in today's AI landscape.
Mercore offers three primary data curation models:
1. **By Task:** Custom, highly complex tasks tailored to specific client needs, often commanding high value ($50 to $10,000 per task) due to the extensive human expert involvement required.
2. **Off-the-shelf Data:** Pre-built, high-quality datasets sold to multiple customers, especially beneficial for new labs seeking to leverage economies of scale.
3. **Hourly Experts:** Providing experts on an hourly basis, though this is gradually being superseded by more structured data offerings.
During the Q&A, Brendan elaborated on data pricing, which considers both the customer's goal for model improvement and Mercore's cost structure. Data quality is maintained through a focus on realism (ensured by expert outlines) and the accuracy of verifiers, often refined through "trajectory analysis" and human feedback. He clarified that while "synthetic data" (model-generated trajectories or environment population) is utilized, human expertise remains indispensable for pushing capabilities "beyond the frontier" of current AI.
Looking ahead, Brendan identified two major future trends for RL environments: developing agents capable of **ultra-long horizon tasks** (requiring hundreds or even thousands of hours) and integrating **virtual co-workers** to train agents on social interaction and collaboration, a critical yet largely unmeasured aspect of human work. He concluded by reiterating that while AI assists, models cannot reliably create their own accurate rubrics for frontier tasks, emphasizing the unique value of human expertise in task and verifier creation for competitive advantage.
摘要
RL environments have become the hottest topic in AI training data. But what are they, exactly? At Sequoia Capital’s Own Your Intelligence event, Mercor CEO Brendan Foody breaks down the three components: worlds (the messages, docs, and files of a real project), apps (high-fidelity clones of tools like Salesforce and Google Workspace), and tasks (prompts paired with rubric-based verifiers). He walks through a real legal environment built with lawyers from top firms, and shares post-training results showing dramatic gains on domain-specific tasks from modest compute.
Brendan also covers why humans remain essential for measuring the frontier, how Mercor prices data, the shift from crowdsourced labeling to expert-built environments, and what's next: ultra-long-horizon tasks and virtual coworkers. He argues that technology once limited to frontier labs is now reaching application companies, and the datasets they build are becoming their moat.
00:00 Introduction
00:47 A short history of the data market: crowdsourcing to agentic data
02:29 What an RL environment is: worlds, apps, tasks
03:57 Why only humans can measure the frontier
05:35 Building verifiers is the hard part
06:44 Walkthrough: a real legal RL environment
08:18 Leaderboards — and what open weights change
09:45 Post-training results on Apex Agents
11:17 Three ways companies buy data
12:49 Q&A: How do you price data?
14:17 Q&A: What "data quality" actually means
16:42 Q&A: The misunderstanding about synthetic data
18:17 Q&A: Why RL environments now — and what comes after
21:20 Q&A: Can you scale rubric generation with models?
23:00 Q&A: RL environments for cyber defense
25:33 Q&A: Build data in-house or partner?
GPT-4正在为你翻译摘要中......
