← Research
Research · Essay

Reinforcement Learning and Sequential Decision-Making

The people who taught machines to learn from consequences, and to master games once thought to need human intuition.

Last reviewed: 26 August 2026 · Next review due: 26 August 2027

Part of the AI People directory: a structured reference to the individuals shaping artificial intelligence across research, industry, governance, ethics and public discourse.

Richard Sutton, reinforcement learning theory, Canada. Co-authored the definitive textbook on reinforcement learning with Andrew Barto and developed temporal-difference learning, the foundation of modern RL. His 2019 essay The Bitter Lesson argued that general methods leveraging computation outperform approaches that encode human knowledge, a principle vindicated by recent AI advances. Key works: Reinforcement Learning: An Introduction (2018); The Bitter Lesson (2019).

Andrew Barto, reinforcement learning, United States. Co-developed actor-critic methods and co-authored the canonical RL textbook with Sutton, establishing reinforcement learning as a rigorous discipline connecting machine learning, control theory and neuroscience. Key works: Reinforcement Learning: An Introduction (2018); actor-critic methods.

Richard Bellman, dynamic programming, United States. Created dynamic programming and the Bellman equation in the 1950s, mathematical foundations that underpin all modern reinforcement learning, and whose work on optimal sequential decision-making became essential fifty years later as RL matured. Key works: dynamic programming; the Bellman equations.

David Silver, game AI and deep RL, United Kingdom. Led the teams that created AlphaGo, AlphaZero and AlphaFold at DeepMind. AlphaGo's 2016 victory over Lee Sedol showed that deep reinforcement learning could master domains previously thought to require human intuition. Key works: AlphaGo; AlphaZero.

John Schulman, policy optimisation, United States. Created TRPO and PPO (Proximal Policy Optimisation), algorithms that made reinforcement learning practical and stable. PPO became the standard for training RL systems and was central to the RLHF techniques used to align language models. Key works: TRPO; PPO.

Pieter Abbeel, robot learning and deep RL, United States. Pioneered using deep reinforcement learning for robotics, enabling robots to learn complex manipulation tasks, and co-founded Covariant and Berkeley's robot learning lab, bridging academic research and practical robotics. Key works: robot learning; inverse reinforcement learning; apprenticeship learning.

Sergey Levine, robot learning and model-based RL, United States. Leads research on enabling robots to learn from real-world experience rather than simulation, addressing key challenges in deploying learning systems in physical environments. Key works: robot RL; offline RL; the decision transformer.

Doina Precup, hierarchical RL and the options framework, Canada. Developed the options framework for hierarchical reinforcement learning, enabling agents to learn and reuse skills, and leads DeepMind's Montreal lab while maintaining academic research on temporal abstraction. Key works: the options framework; hierarchical RL.

Shane Legg, intelligence measurement and RL, United Kingdom. Co-founded DeepMind and developed formal definitions of machine intelligence. His thesis on universal intelligence measures provided theoretical grounding for comparing different AI systems' capabilities. Key works: the Machine Super Intelligence thesis; universal intelligence.

Noam Brown, game theory and strategic AI, United States. Created Libratus and Pluribus, AI systems that defeated top humans at poker, a game of imperfect information, and now works on reasoning and planning at OpenAI, demonstrating that AI can master strategic decision-making under uncertainty. Key works: Libratus; Pluribus.

Reference · SS-2026-026

Cite this page

Hirji, R. (2026). Reinforcement Learning and Sequential Decision-Making. The SuperSkills evidence base, SS-2026-026. https://thesuperskills.com/research/ai-people/reinforcement-learning-leader-of-ai. Last reviewed 26 August 2026.

An evidence review by Rahim Hirji, not peer-reviewed research. For a material claim, cite the underlying study as well; every study here carries its own permanent link.

How citations and IDs work
Ask the evidence
What does the evidence actually show?What should our board be asking about this?Where does Rahim disagree with the consensus?
Bring this into your organisation

If this describes something happening in your teams, say so.

Keynotes, board sessions and advisory work, drawing on research across more than 200 organisations in 30 countries. Tell me the room, the date and the shift you need. A reply within 24 hours.

Start a conversation

Topics and audiences  ·  All research

Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory  ·  Boards  ·  Enquire