Jerry, former VP at OpenAI, and Rohan, a key figure at Google Brain and Gemini, have founded Core Automation with a bold vision: to replace the Transformer architecture and pioneer a new era of AI that truly learns from experience. They argue that while Transformers have been foundational to the current AI boom, their limitations necessitate a radical shift.
Jerry acknowledges the profound impact and economic value of Transformers, enabling scalable pre-training and reinforcement learning (RL). However, he contends that the field is now focused on incremental improvements (making them cheaper/more efficient) rather than fundamentally more powerful or expressive. The core bottleneck, he believes, is the architecture itself. Transformers, despite their capabilities, struggle with real-world adaptability. They are trained in controlled lab environments but face "messier" distributions in deployment. Their "in-context learning" is limited in scope and duration (e.g., 20 minutes for Codecs), while continuous fine-tuning suffers from catastrophic forgetting and data inefficiency. Jerry highlights that current AI systems are effective "human-LLM hybrids" but fall short of true AGI – a model capable of improving itself without human intervention. He emphasizes that "learning from experience" is broader than just RL, drawing parallels between football (RL-like) and deep mathematical reasoning. Core Automation aims to develop new algorithms, expressed architecturally, that enable models to learn and adapt continuously at test time, on real-world user data, over much longer horizons.
Rohan concurs, pointing out that Transformers, while efficient for pre-training (compressing data), are inefficient for inference, particularly due to their "shallow computational depth" and token-by-token generation. He sees chain-of-thought reasoning as an inefficient workaround to compensate for this architectural limitation. Rohan advocates for a holistic approach, combining pre-training and RL to create a more efficient learning procedure. He stresses the critical role of optimization methods, arguing that stronger optimizers can unlock the potential of more complex, deeper architectures that are currently difficult to train. He also contrasts current digital hardware with biological learning, suggesting that significant efficiency gains might require more analog, hardware-aware designs.
Both founders explain their decision to start Core Automation stems from a perceived stagnation in larger labs, which are currently focused on short-term scaling of existing Transformer technology due to competitive pressures. Core Automation, in contrast, aims to be the "most automated lab in the world," designed to facilitate rapid experimentation and push the boundaries of architectural research. They plan to automate bottlenecks like "kernel generation" (the low-level code that optimizes computations on hardware), which currently requires rare human expertise and significant time. For example, optimizing a QR kernel can yield a 60x speedup but is beyond the current capabilities of LLMs. By dramatically accelerating the iteration cycle – potentially executing dozens or even hundreds of experiments daily – they believe they can efficiently search the vast space of possible architectures. Their ultimate goal is to find a "superior architecture" that demonstrates meaningful, long-term adaptability, perhaps even by observing their own models becoming better at the core automation scientists' work each day.