← Research
Research

Reinforcement Learning and Sequential Decision-Making

The people who taught machines to learn from consequences, and to master games once thought to need human intuition.

Last reviewed: 26 August 2026

Part of the AI People directory: a structured reference to the individuals shaping artificial intelligence across research, industry, governance, ethics and public discourse.

Richard Sutton, reinforcement learning theory, Canada. Co-authored the definitive textbook on reinforcement learning with Andrew Barto and developed temporal-difference learning, the foundation of modern RL. His 2019 essay The Bitter Lesson argued that general methods leveraging computation outperform approaches that encode human knowledge, a principle vindicated by recent AI advances. Key works: Reinforcement Learning: An Introduction (2018); The Bitter Lesson (2019).

Andrew Barto, reinforcement learning, United States. Co-developed actor-critic methods and co-authored the canonical RL textbook with Sutton, establishing reinforcement learning as a rigorous discipline connecting machine learning, control theory and neuroscience. Key works: Reinforcement Learning: An Introduction (2018); actor-critic methods.

Richard Bellman, dynamic programming, United States. Created dynamic programming and the Bellman equation in the 1950s, mathematical foundations that underpin all modern reinforcement learning, and whose work on optimal sequential decision-making became essential fifty years later as RL matured. Key works: dynamic programming; the Bellman equations.

David Silver, game AI and deep RL, United Kingdom. Led the teams that created AlphaGo, AlphaZero and AlphaFold at DeepMind. AlphaGo's 2016 victory over Lee Sedol showed that deep reinforcement learning could master domains previously thought to require human intuition. Key works: AlphaGo; AlphaZero.

John Schulman, policy optimisation, United States. Created TRPO and PPO (Proximal Policy Optimisation), algorithms that made reinforcement learning practical and stable. PPO became the standard for training RL systems and was central to the RLHF techniques used to align language models. Key works: TRPO; PPO.

Pieter Abbeel, robot learning and deep RL, United States. Pioneered using deep reinforcement learning for robotics, enabling robots to learn complex manipulation tasks, and co-founded Covariant and Berkeley's robot learning lab, bridging academic research and practical robotics. Key works: robot learning; inverse reinforcement learning; apprenticeship learning.

Sergey Levine, robot learning and model-based RL, United States. Leads research on enabling robots to learn from real-world experience rather than simulation, addressing key challenges in deploying learning systems in physical environments. Key works: robot RL; offline RL; the decision transformer.

Doina Precup, hierarchical RL and the options framework, Canada. Developed the options framework for hierarchical reinforcement learning, enabling agents to learn and reuse skills, and leads DeepMind's Montreal lab while maintaining academic research on temporal abstraction. Key works: the options framework; hierarchical RL.

Shane Legg, intelligence measurement and RL, United Kingdom. Co-founded DeepMind and developed formal definitions of machine intelligence. His thesis on universal intelligence measures provided theoretical grounding for comparing different AI systems' capabilities. Key works: the Machine Super Intelligence thesis; universal intelligence.

Noam Brown, game theory and strategic AI, United States. Created Libratus and Pluribus, AI systems that defeated top humans at poker, a game of imperfect information, and now works on reasoning and planning at OpenAI, demonstrating that AI can master strategic decision-making under uncertainty. Key works: Libratus; Pluribus.

Cite this

Hirji, R. (2026). Reinforcement Learning and Sequential Decision-Making. The SuperSkills Intelligence Company. Last reviewed 26 August 2026. thesuperskills.com/research/ai-people/reinforcement-learning-leader-of-ai

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →
Box of Amazing

Rahim’s free weekly letter on AI and human capability

If this was useful, the weekly letter is where the thinking happens first. Most of what ends up on this site starts there. Weekly essays on AI, capability and the future of work. Read by 25,000 people, every week since 2017. Free, and one click to stop.

Opens Substack to confirm. No pitch in it, unsubscribe in one click, and nobody follows up because you read something.

Running an event, or responsible for how AI arrives in your organisation? Keynotes  ·  Advisory and coaching  ·  Enquire