The Father of Reinforcement Learning Leaves Carmack to Start His Own AI Lab
Richard Sutton, the Turing Award winner widely known as the father of reinforcement learning, has left John Carmack’s Keen Technologies to establish his own lab. Its mission is unconventional: developing an agent capable of continuous real-time learning with approximately 20 watts of power, mirroring the functionality of a human brain rather than relying on a data center full of frozen models.
July 14, 2026 – 11:54 am
Image by: Xuthoria
One of AI’s most celebrated minds is embarking on a solo journey. Richard Sutton shared the 2024 Turing Award for pioneering modern reinforcement learning. On Monday, he announced his departure from John Carmack’s startup Keen Technologies to found a new venture, Oak Lab, via a post on X. He expressed gratitude towards Carmack and Keen while revealing that he and collaborator Khurram Javed are launching their own startup to pursue "a slightly different path toward understanding intelligence."
Sutton’s critique of current deep-learning methods is direct. In his view, they are “weak and inefficient, and need not more tweaks, but fundamentally new ideas and a thorough reworking.”
Learning from Experience, Not Datasets
Oak Lab’s core argument revolves around the source of intelligence. Sutton has consistently maintained that intelligence arises from runtime experience rather than being distilled from curated human datasets. This distinction holds significant weight. Today’s models learn from data meticulously collected, cleaned, and filtered by humans. Real-world experiences are messier, comprising predictable signals and abundant noise.
In Oak Lab’s first research post, Sutton and Javed demonstrate that standard optimizers like SGD struggle to differentiate between the two. They propose a solution based on an old idea of Sutton’s, refining it with a new neural approach they call NetworkIDBD. This algorithm selectively assigns credit to meaningful signals, ignoring noise, which improves learning efficiency.
An Agent That Runs on 20 Watts
Oak Lab’s ultimate goal is to achieve trillion-parameter agent that learns and plans in real time using just 20 watts of power—the approximate energy consumption of the human brain. Today’s models, by contrast, undergo initial training in data centers consuming significant megawatts before remaining dormant.
A Contrarian Bet
Sutton’s career has been defined by challenging the mainstream. His 2019 essay “The Bitter Lesson” is a landmark in AI literature, and his textbook with Andrew Barto has trained countless researchers. Yet, he expresses skepticism about scaling up pre-trained language models as the path to genuine intelligence, aligning with perspectives from figures like Yann LeCun, who has bet $1bn on world models, and David Silver, creator of AlphaGo.