Transformative AI
Transformative Artificial Intelligence
Progress in the field of artificial intelligence (AI) has been extraordinarily rapid in recent years. Although the future impacts of AI remain highly uncertain and are subject to deep disagreement, there is a very real possibility that further advances in the field could lead to radical economic and social transformations and that navigating these changes could be among the most significant challenges we face in coming decades. Part of our research therefore aims to inform our understanding of the potential for transformative societal changes due to the development of advanced AI systems, the largest-scale challenges likely to arise as a result, and the best means of addressing them.
Catastrophic Risk
It has been argued that advanced AI systems pose significant catastrophic risks, up to and including the risk of existential catastrophe (Yudkowsky 2008; Bostrom 2014; Russell 2019; Cotra 2022; Carlsmith 2025; Hendrycks 2025; Kulveit et al. 2025). Other large-scale risks include extreme power concentration and democratic collapse (Drago and Laine 2025; Acemoglu, Gitmez, and Shadmehr 2026). Many arguments for catastrophic risk appear to turn on philosophical assumptions, such as instrumental convergence (Omohundro 2008; Bostrom 2014: 105-114; Gallow 2025). How should we approach and assess these arguments? Which threat models are worth taking seriously, and how should we prioritize among the different catastrophic risks they spotlight? (Note that here – as elsewhere – we assume no particular conclusions: we are interested in work arguing that risks from AI are overstated, as well as work that might strengthen the case for catastrophic risk.)
Many have argued that one of our key strategies for mitigating catastrophic risks associated with the potential for loss-of-control should be to ensure that AI systems are aligned with human values. However, there remains deep disagreement about how best to conceptualize alignment: in particular, to what extent alignment should be understood in terms of deferring to human preferences and commands as opposed to the internalization of appropriate ethical values capable of guiding autonomous choice (Leike et al. 2018; Gabriel 2020; Askell et al. 2021; Zhi-Xuan 2025). A wide range of further technical and governance-based mitigation strategies have been proposed for addressing catastrophic risks from AI (Dafoe 2018; D’Alessandro and Kirk-Giannini 2025). What are their strengths and weaknesses? What novel mitigation strategies remain to be analysed in depth?
The Trajectory of Civilisation
Even setting aside the possibility of catastrophe, the development of transformative artificial intelligence could potentially bring about profound societal changes, including extremely rapid economic growth (Trammell and Korinek 2023; Davidson et al. 2026) and large-scale technological unemployment (Susskind 2020; Korinek and Juelfs 2024). Social norms and institutions may be upended due to the rapid pace of change, before potentially crystallizing into some new, highly durable configuration (MacAskill 2022: 75-102).
Ideally, we’d like to know not only how plausible it is that advanced AI might bring about different possible societal transformations, but also how best to navigate the different risks and opportunities this might pose. This may include the desirability and feasibility of broad-based improvements in the ability of key societal institutions to respond to rapid changes, as well as preparatory efforts to map out key challenges and best responses beforehand, reducing the risk that we’ll be caught off-guard (MacAskill and Moorhouse 2025). Last but not least, it may involve efforts to map out realistic, positive visions of a society transformed by the impacts of advanced machine intelligence, providing a sense of what to aim and hope for.
Digital Minds
The topics outlined above are most naturally thought of as concerned with the potential positive and negative impacts of advanced AI on human flourishing. Impacts from transformative AI will stretch far beyond the human realm, however. Non-human animals will be impacted in myriad ways, such as through advances in precision livestock farming (Simoneau-Gilbert and Birch 2024). AI systems may themselves become moral patients, meriting our concern and respect (Long et al. 2024). We are interested in research at the intersection of ethics and the philosophy of mind and cognitive science relevant to understanding what kind of moral standing, if any, AI systems could have, and how we should respond to different forms of evidence relevant to settling that question. For more details, see here.
- Acemoglu, D., Gitmez, A. A. and Shadmehr, M. 2026. "Automation and Repression." NBER Working Paper 35336. URL: https://www.nber.org/papers/w35336
- Askell, A., Bai, Y., Chen, A., et al. 2021. "A General Language Assistant as a Laboratory for Alignment." arXiv:2112.00861. URL: https://arxiv.org/abs/2112.00861
- Bostrom, N. 2014. Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press.Carlsmith, J. 2025. "Existential Risk from Power-Seeking AI." In H. Greaves, J. Barrett and D. Thorstad (eds.), Essays on Longtermism, 383–409. Oxford: Oxford University Press.
- Cotra, A. 2022. "Without Specific Countermeasures, the Easiest Path to Transformative AI Likely Leads to AI Takeover." LessWrong. URL: https://www.lesswrong.com/posts/pRkFkzwKZ2zfa3R6H/without-specific-countermeasures-the-easiest-path-to
- Dafoe, A. 2018. AI Governance: A Research Agenda. GovAI. URL: https://www.governance.ai/research-paper/agenda
- D'Alessandro, W. and Kirk-Giannini, C. D. 2025. "Artificial Intelligence: Approaches to Safety." Philosophy Compass 20(5): e70039.
- Davidson, T., Halperin, B., Houlden, T. and Korinek, A. 2026. "When Does Automating AI Research Produce Explosive Growth? Feedback Loops in Innovation Networks." NBER Working Paper 35155. URL: https://www.nber.org/papers/w35155
- Drago, L. and Laine, R. 2025. The Intelligence Curse. URL: https://intelligence-curse.ai/
- Gabriel, I. 2020. "Artificial Intelligence, Values, and Alignment." Minds and Machines 30: 411–437.
- Gallow, J. D. 2025. "Instrumental Divergence." Philosophical Studies 182(7): 1581–1607.
- Hendrycks, D. 2025. Introduction to AI Safety, Ethics, and Society. Boca Raton, FL: CRC Press.
- Korinek, A. and Juelfs, M. 2024. "Preparing for the (Non-Existent?) Future of Work." In J. Bullock et al. (eds.), The Oxford Handbook of AI Governance, 746 - 776. Oxford: Oxford University Press.
- Kulveit, J., Douglas, R., Ammann, N., Turan, D., Krueger, D. and Duvenaud, D. 2025. "Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development." arXiv:2501.16946. URL: https://arxiv.org/abs/2501.16946
- Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V. and Legg, S. 2018. "Scalable Agent Alignment via Reward Modeling: A Research Direction." arXiv:1811.07871. URL: https://arxiv.org/abs/1811.07871
- Long, R., Sebo, J., Butlin, P., et al. 2024. "Taking AI Welfare Seriously." arXiv:2411.00986. URL: https://arxiv.org/abs/2411.00986
- MacAskill, W. 2022. What We Owe the Future. New York: Basic Books.
- MacAskill, W. and Moorhouse, F. 2025. "Preparing for the Intelligence Explosion." Forethought. URL: https://www.forethought.org/research/preparing-for-the-intelligence-explosion
- Omohundro, S. M. 2008. "The Basic AI Drives." In P. Wang, B. Goertzel and S. Franklin (eds.), Proceedings of the 2008 conference on Artificial General Intelligence 2008: Proceedings of the First AGI Conference, 483–492. Amsterdam: IOS Press.
- Russell, S. 2019. Human Compatible: Artificial Intelligence and the Problem of Control. New York, NY: Viking.
- Simoneau-Gilbert, V. and Birch, J. 2024. “The Dangers of AI Farming”. Aeon. URL: https://aeon.co/essays/how-to-reduce-the-ethical-dangers-of-ai-assisted-farming
- Susskind, D. 2020. A World Without Work. London: Allen Lane.
- Trammell, P. and Korinek, A. 2023. "Economic Growth under Transformative AI." NBER Working Paper 31815. URL: https://www.nber.org/papers/w31815
- Yudkowsky, E. 2008. "Artificial Intelligence as a Positive and Negative Factor in Global Risk." In N. Bostrom and M. M. Ćirković (eds.), Global Catastrophic Risks, 308–345. Oxford: Oxford University Press.
- Zhi-Xuan, T., Carroll, M., Franklin, M. and Ashton, H. 2025. "Beyond Preferences in AI Alignment." Philosophical Studies 182: 1813–1863.