This episode of "Flirting with Models," hosted by Corey Hofstein, features Ben Wellington, Head of Complex Feature Engines at Two Sigma. The conversation delves into the human factor behind quantitative strategy, specifically focusing on feature generation and the transformative impact of AI in this space.
Here's a detailed summary of the discussion:
**1. Ben Wellington's Background and Role:**
* **Education:** Ben has a PhD in Natural Language Processing (NLP) from New York University, specializing in machine translation during an era (pre-Siri/Alexa) when such applications were novel. He chose the field because it "sounded cool."
* **Career at Two Sigma:** He joined Two Sigma in 2007 when it was a smaller firm (125 people). His initial role was on the data engineering team, focused on collecting, ingesting, and cleaning data.
* **Shift to Modeling:** Ben discovered interesting datasets and performed an internal "political analysis of news places," which caught the attention of a modeler at Two Sigma. This led him to shift his focus to studying how textual data predicts markets, a role he's held for over 10 years, aiming to predict asset movements (up/down) over various time horizons.
* **Current Title:** Head of Complex Feature Engines, highlighting his core focus.
**2. Defining "Feature" in Quant Investing:**
* **Core Concept:** A "feature" is a "fact about the world that is worth itself noting," and for Two Sigma, it must be "economically meaningful."
* **Examples:**
* **Familiar:** Change in stock price over the last week, stock price movement since last year, volatility.
* **More Interesting:** An analyst recently upgrading a stock.
* **Textual/Alternative (Ben's expertise):** Amount of news coverage about a company compared to its history, the day/time a press release is issued (e.g., Friday after hours).
* **Purpose:** Features encapsulate hypotheses. They layer human intelligence and intuition onto raw data, helping computers focus on what's potentially meaningful, rather than just throwing all raw data at an algorithm. This makes the job tractable and leverages human knowledge.
**3. Two Sigma's Operating Model vs. Pod Shops:**
* **Shared Platform Model:** Unlike pod shops where teams (pods) compete, Two Sigma operates on a model where modelers' predictions are centralized and aggregated for the firm's larger portfolios. Their goal is not to control a specific portfolio, but to make the best possible predictions.
* **Global Optimization:** This structure encourages "global optimization" rather than "local optimization" (as in pod shops). Teams are incentivized to elevate their best work (features, algorithms) to a shared platform for others to use, making the overall firm better. This is described as a "cyclical" process where teams constantly push improvements to the platform.
**4. Where Alpha Lives (Data, Features, Forecasting):**
* Ben identifies three potential silos for alpha:
1. **Data Access:** Its proportional weight has "dropped" due to increased commoditization over the last 15-20 years. While Two Sigma still has proprietary data (long history, crowdsourced trading ideas), the edge from *just* data access is less than before.
2. **Feature Generation (Ben's Focus):** This is "just huge." The ability to creatively and uniquely transform data into insightful observations is a massive edge, driven by "human ingenuity." Ben gives the example of studying analysts: beyond simple upgrades, considering their location, school, connection to CEOs, etc. Each tiny, clever question (e.g., 200 questions = 200 features) collectively creates a significant alpha signal.
3. **Forecasting/ML Algorithms:** This is also a "very big part," with a dedicated "techniques" team focusing on advanced algorithms to extract maximum information from features.
* **Synergy:** The intersection of feature-focused teams (identifying clever data) and techniques teams (applying advanced algorithms) is where "amazing things happen," driven by diversity of approach.
**5. Ideating Features and Managing Challenges:**
* **Ideation Process:** It requires "persistence" and a human layer to discern what's meaningful (e.g., analyst's eye color is likely irrelevant, but their training or location might be). It involves diversity of thought from various scientific backgrounds.
* **Generalizing Instances:** Taking real-world events (e.g., a train derailment affecting a company) and asking "Is this generalizable?" to create broader features.
* **Scientific Method:** Adhering to the scientific method by stating hypotheses, testing them, and not post-hoc rationalizing results that contradict priors.
* **Collinearity:** This is "the death of a signal." Researchers constantly search for *orthogonality* to avoid features that merely re-explain existing signals (e.g., news sentiment that just tracks post-earnings reports). Tools and thresholds are used to ensure new alpha signals have sufficient orthogonality.
* **Forecast Horizon:** Researchers are mindful of the time horizon relevant to a feature (e.g., store visit data might predict sales months out).
* **Feature Lifecycle:** As more market participants act on a signal, its forecast horizon can speed up, potentially eroding its viability. Two Sigma aims for creative features that resist this pressure and monitors for "alpha decay."
**6. Hiring for Creativity and Ingenuity:**
* **Beyond Technical Skills:** While math, stats, and coding are important, feature engineering requires ingenuity, creative, and lateral thinking.
* **Interview Approach:** Open-ended questions are used to assess creativity, such as "If you had access to a social media firehose, what could you predict?" The goal is to see how candidates generate novel, interesting ideas without a fixed "right" answer.
**7. AI/LLMs and the Future of Feature Generation:**
* **Ben's Core Insight:** "NLP was about turning language into numbers, with LLMs I can turn numbers into language. Anything can be language now."
* **Impact:** LLMs *generate textual data*, which is a "revolution of opportunity" for Two Sigma, a firm with 15+ years of experience trading textual data. Proprietary data now includes *generated* data.
* **Cost Collapse:** LLMs dramatically reduce the cost and time investment for exploring ideas. A hypothetical example: previously, studying CEO blinking for predictive value (requiring computer vision, 6 months of work) had low ROI. Now, LLMs can make such explorations 100x faster, making previously "too expensive" ideas viable. This "frees the mind from having to worry so much about implementation."
* **Democratization of Edge (Critique):** Ben acknowledges that if LLMs make feature creation cheaper for everyone, the edge could erode faster.
* **Two Sigma's Counter:** The firm is "shovel ready" due to its deep experience and systematic approach to using such technologies. While the ability to *create* data might be democratized, the "then what" (how to use it effectively) remains a key differentiator rooted in experience.
**8. Risk of Lowering Entropy of Output (Homogenization):**
* **The Danger:** Automation for a production line (e.g., cardboard boxes) requires identical output. However, for alpha generation, *orthogonality* and *originality* are crucial. If AI tools lead everyone to generate the same signals, it homogenizes portfolios and erodes alpha.
* **Two Sigma's Strategy:** View AI as an *amplification tool* for originality. Tools should empower researchers to multiply their unique insights. If two researchers with different backgrounds use the same tool, it should produce different, context-aware answers, reinforcing diversity of thought rather than homogenizing it.
**9. Trade-off: Leaning into Difficulty vs. Waiting for Capability:**
* **Pushing Boundaries:** It's important to "push the boundaries faster than the tools come" to understand limitations ("see the wall"). This way, when new tools emerge, researchers know exactly where to apply them.
* **Competitive Advantage:** Waiting 12-18 months for AI capabilities to mature might save engineering effort, but being first in the market with new alpha signals provides a significant, potentially short-lived, advantage.
**10. Infrastructure Choices (On-Premise vs. Vendor Models):**
* **Diversification:** The key is to avoid being held hostage by any single approach.
* **Flexibility:** The technical stack must allow for swapping between vendor, external, or internal models as needed.
* **Control over Lifecycle:** A major concern is vendor models being discontinued or upgraded, forcing continuous adaptation. Internal models offer more control.
* **Look-Ahead Bias:** Not understanding a model's training data (especially external ones) can lead to spurious correlations or biases (e.g., an LLM trained on post-Enron data might inherently label anything about Enron as "bad," even when querying historical data).
* **Balance:** Balancing the control of internal models with the impressive pace of innovation from frontier labs is crucial.
**11. P-Hacking and Overfitting with AI:**
* **Natural Safeguards:** Researchers failing to create models that predict the future (due to overfitting the past) will naturally face career setbacks.
* **Institutional Measures:** Two Sigma has long-standing tools, statistical tests, and best practices to combat p-hacking and overfitting in scaled experiments.
* **Nuance of Overfitting:** All models "overfit by definition" because the future is never exactly like the past. The goal is to overfit to a "comfortable and measurable" degree, ensuring robustness to trading environments.
* **AI's Impact:** AI *exacerbates* the risk of p-hacking due to the sheer volume of experiments possible. Therefore, increased vigilance in education, checks, and robustness is necessary as scale increases.
**12. Data Upgrades and Model Rebuilding:**
* **Academic vs. Business Mindset:** In academia, the goal is often 100% perfection on one project. In a diverse portfolio context like Two Sigma, "the last 10% of gains in that journey is 90% of the work."
* **Diminishing Returns:** It's often more beneficial to pursue 10 different ideas at 90% completion than one idea at 95% completion. A good quant knows "when to move on."
* **ROI-Driven Decisions:** When data sources (e.g., vendor V3 to V4) or algorithms improve, the decision to rebuild or retrain an existing model is made with a business ROI lens. Often, the time is better spent exploring new, orthogonal signals rather than marginally improving an already well-performing model.
**13. Idiosyncrasy at Scale (AI Frontier):**
* **Inspiring Specificity:** Ben uses "idiosyncrasy at scale" to encourage detailed, company-specific analysis, traditionally avoided due to small sample sizes.
* **AI's Role:** AI offers the unique ability to dive deeply into a specific company's data, building highly tailored features. These bespoke insights can then be generalized across thousands of other companies simultaneously.
* **Mimicking Discretionary Investing:** AI enables a scale-oriented firm like Two Sigma to model the "deep knowledge" and "company-specific opportunities" often identified by discretionary investors, which typically don't show up in broad, aggregated data. This allows for deeper exploration of "bespoke and hard-to-model parts of company data."
**14. Skills for the Future for Junior Engineers:**
* **Shift in Value:** Technical coding skills will decrease in relative value. The ability to articulate ideas and "tell the AI what to do" (prompting) will increase.
* **Key Advice:**
1. Get comfortable with AI and its capabilities; become "AI native."
2. View AI as an *amplification tool* for unique personal skills and ideas, rather than a replacement.
3. Focus on developing originality, orthogonality, and creativity.
4. Ideas and the effective use of AI to implement them will be paramount.
**15. Ben's Personal Obsession:**
* **Public Data and Government Releases:** Ben is fascinated by the wealth of public data (parking tickets, speeding stops, restaurant health inspections) and its potential for improving cities and society.
* **AI's Impact on Public Data:** He finds it inspiring that AI can democratize access to this information, removing the need for specialized data science skills. Urban planners, lawyers, and the general public can ask questions directly, making government data more accessible and useful.
* **Corey's Humorous Counter:** As this data becomes more accessible, it might highlight government inefficiencies, potentially leading to the data being pulled.