Brendan from Mercore discussed the transformative shift in the AI data market, moving from low-skilled crowdsourcing to an "agentic era" focused on high-skilled experts and sophisticated Reinforcement Learning (RL) environments. Mercore has experienced rapid growth, fueled by this demand for more complex data solutions.
He defined an **RL environment** as a three-part system designed for AI model evaluation and training:
1. **Worlds:** Realistic, comprehensive scenarios encompassing messages, documents, spreadsheets, and other digital artifacts that mimic real-world projects.
2. **Apps:** High-fidelity clones of popular applications like Salesforce, Microsoft 365, or Google Workspace, allowing AI agents to interact authentically via APIs, CLI, or UI.
3. **Tasks:** Specific prompts paired with robust verifiers (rubrics or unit tests) that dictate the agent's objective and measure its performance.
The fundamental challenge, Brendan explained, is to span the entire distribution of human tasks across all economic sectors. Mercore's strategy involves leveraging extensive human expertise, evidenced by 2.5 million expert hours recorded in Q2 alone. Humans are crucial because, unlike in easily quantifiable domains like math, AI models cannot reliably self-evaluate in complex, open-ended tasks (e.g., drafting a slide deck or legal memo). Expert-designed rubrics are essential for accurate assessment and effective model learning.
Brendan showcased a legal RL environment as an example. Top law firms collaborate to outline real legal project scenarios, populating detailed "data rooms" with relevant information. These are then rendered within cloned application environments (e.g., Google Workspace). Agents are tasked with complex analyses, and human-created rubrics, with specific criteria, ensure precise verification and prevent "reward hacking," ultimately leading to leaderboards and aggregate model scores.
The impact of this approach is substantial. Post-training on datasets like "Apex Agents" (e.g., GLM 4.7 with 500k compute on 1800 tasks) demonstrated dramatic performance gains and significant generalization to other benchmarks. This powerful technology, initially developed for frontier labs, is now becoming accessible to application-layer companies, enabling them to "own their own intelligence" – a crucial differentiator in today's AI landscape.
Mercore offers three primary data curation models:
1. **By Task:** Custom, highly complex tasks tailored to specific client needs, often commanding high value ($50 to $10,000 per task) due to the extensive human expert involvement required.
2. **Off-the-shelf Data:** Pre-built, high-quality datasets sold to multiple customers, especially beneficial for new labs seeking to leverage economies of scale.
3. **Hourly Experts:** Providing experts on an hourly basis, though this is gradually being superseded by more structured data offerings.
During the Q&A, Brendan elaborated on data pricing, which considers both the customer's goal for model improvement and Mercore's cost structure. Data quality is maintained through a focus on realism (ensured by expert outlines) and the accuracy of verifiers, often refined through "trajectory analysis" and human feedback. He clarified that while "synthetic data" (model-generated trajectories or environment population) is utilized, human expertise remains indispensable for pushing capabilities "beyond the frontier" of current AI.
Looking ahead, Brendan identified two major future trends for RL environments: developing agents capable of **ultra-long horizon tasks** (requiring hundreds or even thousands of hours) and integrating **virtual co-workers** to train agents on social interaction and collaboration, a critical yet largely unmeasured aspect of human work. He concluded by reiterating that while AI assists, models cannot reliably create their own accurate rubrics for frontier tasks, emphasizing the unique value of human expertise in task and verifier creation for competitive advantage.