← Research
Research

AI Safety, Alignment and Existential Risk

The people asking whether we can keep powerful AI systems doing what humans actually want.

Part of the AI People directory: a structured reference to the individuals shaping artificial intelligence across research, industry, governance, ethics and public discourse.

Stuart Russell, value alignment and AI foundations, United States. Co-author of the definitive AI textbook, Artificial Intelligence: A Modern Approach, who became a leading voice on alignment. His book Human Compatible reframes AI development around uncertainty over human preferences rather than fixed objectives. *Key works: AI: A Modern Approach (2020); Human Compatible (2019).*

Nick Bostrom, existential risk and superintelligence, United Kingdom. Philosopher whose Superintelligence (2014) framed AI alignment as a civilisational challenge, and who founded Oxford's Future of Humanity Institute, making AI safety a mainstream concern before the current wave of progress. *Key works: Superintelligence (2014).*

Eliezer Yudkowsky, alignment theory and risk communication, United States. Co-founded the Machine Intelligence Research Institute and has written extensively on alignment since 2001. His detailed scenarios of misaligned AI have shaped how the field thinks about failure modes, and he is a persistent voice for taking catastrophic risks seriously. *Key works: long-form alignment essays; MIRI research direction.*

Paul Christiano, alignment techniques and RLHF, United States. Developed foundational techniques for aligning AI systems with human preferences, contributing to the RLHF methods now used to train ChatGPT and Claude, and founded the Alignment Research Center. *Key works: the RLHF lineage; scalable-oversight proposals; ARC.*

Jan Leike, scalable alignment, United Kingdom. Led alignment research at OpenAI before departing in 2024 over concerns about safety prioritisation, and co-leads superalignment efforts focused on using AI systems to help align more powerful AI. *Key works: scalable-alignment research; superalignment.*

Chris Olah, interpretability and mechanistic understanding, United States. Pioneered neural network interpretability, developing techniques to understand what happens inside neural networks, and co-founded Anthropic's interpretability team, making the field's work accessible through visual explanations. *Key works: neural-network visualisations; circuits research.*

Dan Hendrycks, safety benchmarks and risk research, United States. Created benchmarks for measuring AI robustness, safety and dangerous capabilities, providing concrete metrics for tracking safety progress, and leads the Center for AI Safety. *Key works: the MMLU benchmark; robustness benchmarks; CAIS.*

Toby Ord, existential risk and ethics, United Kingdom. Philosopher who synthesised existential risks in The Precipice, arguing that AI poses significant risks this century, and co-founded Giving What We Can and the Centre for Effective Altruism. *Key works: The Precipice (2020).*

Max Tegmark, AI risk communication and physics perspectives, United States. MIT physicist who founded the Future of Life Institute and co-organised the Asilomar AI Principles. His book Life 3.0 brought AI safety concerns to popular audiences. *Key works: Life 3.0 (2017); the Future of Life Institute.*

Ajeya Cotra, AI forecasting, United States. Authored influential reports forecasting AI timelines using biological anchors and compute trends. Her work at Open Philanthropy shapes how funders and researchers think about when transformative AI might arrive. *Key works: the biological-anchors report; AI forecasting.*

Jack Clark, AI policy and safety advocacy, United States. Co-founded Anthropic after serving as Policy Director at OpenAI, and his Import AI newsletter has tracked AI progress for years, bridging technical development and policy implications. *Key works: the Import AI newsletter; Anthropic co-founding; policy advocacy.*

Helen Toner, AI governance and security, United States. Director at Georgetown's Center for Security and Emerging Technology and a former OpenAI board member whose tenure included the November 2023 crisis. Her research on AI governance and international competition informs policy discussions. *Key works: AI governance research; CSET; policy analysis.*

The work

Where the writing comes from.

These essays draw on research across more than 200 organisations in 30 countries. See the wider body of work, or bring it into your organisation.

All research →