Fetching the paper…
Reading the bibliography…
In coming years or decades, artificial general intelligence (AGI) may surpass human capabilities across many critical domains.
An investigation of model-free planning, May 2019
Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, Sébastien Racanière, Théophane Weber, David Raposo, Adam Santoro, Laurent Orseau, Tom Eccles, Greg Wayne, David Silver, and Timothy Lillicrap · 1901
Earlier work this paper cites.
Towards automatic concept-based explanations, 2019
Amirata Ghorbani, James Wexler, James Zou, and Been Kim · 1902
Earlier work this paper cites.
Risks from Learned Optimization in Advanced Machine Learning Systems, December 2021
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 1906
Earlier work this paper cites.
Reinforcement Learning Upside Down: Don’t Predict Rewards – Just Map Them to Actions, June 2020
Juergen Schmidhuber · 1912
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
A neural substrate of prediction and reward
Wolfram Schultz, Peter Dayan, and P Read Montague · 1997
Earlier work this paper cites.
No free lunch theorems for optimization
David H Wolpert and William G Macready · 1997
Earlier work this paper cites.
Guns, germs, and steel , volume 521
Jared M Diamond and Doug Ordunio · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
The p versus np problem
Stephen Cook · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart Russell · 2000
Earlier work this paper cites.
The tragedy of great power politics
John J Mearsheimer, Glenn Alterman, et al · 2001
Earlier work this paper cites.
Evolution of digital organisms at high mutation rates leads to survival of the flattest
Claus O Wilke, Jia Lan Wang, Charles Ofria, Richard E Lenski, and Christoph Adami · 2001
Earlier work this paper cites.
The evolved radio and its implications for modelling the evolution of novel sensors
Jon Bird and Paul Layzell · 2002
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Marcus Hutter · 2004
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
Shane Legg and Marcus Hutter · 2007
Earlier work this paper cites.
The basic AI drives
Stephen M Omohundro · 2008
Earlier work this paper cites.
Artificial intelligence as a positive and negative factor in global risk
Eliezer Yudkowsky et al · 2008
Earlier work this paper cites.
The human brain in numbers: a linearly scaled-up primate brain
Suzana Herculano-Houzel · 2009
Earlier work this paper cites.
Hidden incentives for auto-induced distributional shift, 2020
David Krueger, Tegan Maharaj, and Jan Leike · 2009
Earlier work this paper cites.
Learning to summarize from human feedback, 2020
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2009
Earlier work this paper cites.
The superintelligent will: Motivation and instrumental rationality in advanced artificial agents
Nick Bostrom · 2012
Earlier work this paper cites.
Existential risk prevention as global priority
Nick Bostrom · 2013
Earlier work this paper cites.
Evolutionary perspectives on interpersonal acceptance and
Mark R Leary and Catherine A Cottrell · 2013
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom · 2014
Earlier work this paper cites.
Artificial general intelligence: concept, state of the art, and future prospects
Ben Goertzel · 2014
Earlier work this paper cites.
Confirmation bias: Roles of search engines and search contexts, 2015
Varol Kayhan · 2015
Earlier work this paper cites.
Nearest unblocked strategy, 2015
Eliezer Yudkowsky · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Racing to the precipice: a model of artificial intelligence development
Stuart Armstrong, Nick Bostrom, and Carl Shulman · 2016
Earlier work this paper cites.
Scott Garrabrant, Tsvi Benson-Tilsen, Andrew Critch, Nate Soares, and Jessica Taylor · 2016
Earlier work this paper cites.
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2016
Earlier work this paper cites.
The AI alignment problem: why it is hard, and where to start
Eliezer Yudkowsky · 2016
Earlier work this paper cites.
Learning from human preferences, 2017
Dario Amodei, Paul Christiano, and Alex Ray · 2017
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Cyclegan, a master of steganography, 2017
Casey Chu, Andrey Zhmoginov, and Mark Sandler · 2017
Earlier work this paper cites.
The off-switch game
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2017
Earlier work this paper cites.
The flash crash: High-frequency trading in an electronic market
Andrei Kirilenko, Albert S Kyle, Mehrdad Samadi, and Tugkan Tuzun · 2017
Earlier work this paper cites.
Prediction machines: the simple economics of artificial intelligence
Ajay Agrawal, Joshua Gans, and Avi Goldfarb · 2018
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
David Alvarez Melis and Tommi Jaakkola · 2018
Earlier work this paper cites.
Vector-based navigation using grid-like representations in artificial agents
Andrea Banino, Caswell Barry, Benigno Uria, Charles Blundell, Timothy Lillicrap, Piotr Mirowski, Alexander Pritzel, Martin J Chadwick, Thomas Degris, Joseph Modayil, et al · 2018
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts, October 2018
Paul Christiano, Buck Shlegeris, and Dario Amodei · 2018
Earlier work this paper cites.
AI governance: a research agenda
Allan Dafoe · 2018
Earlier work this paper cites.
Tom Everitt, Gary Lea, and Marcus Hutter · 2018
Earlier work this paper cites.
Embedded Agents, October 2018
Scott Garrabrant · 2018
Earlier work this paper cites.
When will AI exceed human performance? evidence from AI experts
Katja Grace, John Salvatier, Allan Dafoe, Baobao Zhang, and Owain Evans · 2018
Earlier work this paper cites.
Learning latent dynamics for planning from pixels, 2018
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2018
Earlier work this paper cites.
AI safety via debate, May 2018
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Categorizing variants of goodhart’s law, 2018
David Manheim and Scott Garrabrant · 2018
Earlier work this paper cites.
AI and Compute, May 2018
OpenAI · 2018
Earlier work this paper cites.
Constructing unrestricted adversarial examples with generative models
Yang Song, Rui Shu, Nate Kushman, and Stefano Ermon · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle-Perez, Chico Q Camargo, and Ard A Louis · 2018
Earlier work this paper cites.
A parametric, resource-bounded generalization of löb’s theorem, and a robust cooperation criterion for open-source game theory
Andrew Critch · 2019
Earlier work this paper cites.
Neural architecture search: A survey
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter · 2019
Earlier work this paper cites.
Learning to predict without looking ahead: World models without forward prediction
Daniel Freeman, David Ha, and Luke Metz · 2019
Earlier work this paper cites.
Mor Geva, Yoav Goldberg, and Jonathan Berant · 2019
Cited alongside, same era.
Capture the Flag: the emergence of complex cooperative agents, May 2019
Max Jaderberg, Wojciech Marian Czarnecki, Iain Dunning, Thore Graepel, and Luke Marris · 2019
Cited alongside, same era.
Human compatible: Artificial intelligence and the problem of control
Stuart Russell · 2019
Cited alongside, same era.
Observational overfitting in reinforcement learning
Xingyou Song, Yiding Jiang, Stephen Tu, Yilun Du, and Behnam Neyshabur · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Locating and Editing Factual Associations in GPT, June 2022
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Closest in time.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Closest in time.
The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models, February 2022
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Closest in time.
Mapping language models to grounded conceptual spaces
Roma Patel and Ellie Pavlick · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Debate update: Obfuscated arguments problem - AI Alignment Forum, December 2020
Beth Barnes and Paul Christiano · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Toward trustworthy AI development: mechanisms for supporting verifiable claims
Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, et al · 2020
Cited alongside, same era.
Forecasting TAI with biological anchors, 2020
Ajeya Cotra · 2020
Cited alongside, same era.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Cited alongside, same era.
Aligning AI with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
Specification gaming: the flip side of AI ingenuity, April 2020
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Cited alongside, same era.
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike · 2022
Closest in time.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton · 2022
Closest in time.
Defining and characterizing reward gaming
Joar Max Viktor Skalse, Nikolaus HR Howe, Dmitrii Krasheninnikov, and David Krueger · 2022
Closest in time.
2022 Expert Survey on Progress in AI, August 2022
Zach Stein-Perlman, Benjamin Weinstein-Raun, and Katja Grace · 2022
Closest in time.
ML Systems Will Have Weird Failure Modes, January 2022
Jacob Steinhardt · 2022
Closest in time.
Parametrically retargetable decision-makers tend to seek power
Alexander Matt Turner and Prasad Tadepalli · 2022
Closest in time.
Dual use of artificial-intelligence-powered drug discovery
Fabio Urbina, Filippa Lentzos, Cédric Invernizzi, and Sean Ekins · 2022
Closest in time.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Closest in time.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Closest in time.
The people in intimate relationships with AI chatbots, 2022
Chiara Wilkinson · 2022
Closest in time.
Least-to-most prompting enables complex reasoning in large language models, 2022
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, and Ed Chi · 2022
Closest in time.
Adversarial training for high-stakes reliability
Daniel M Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Ben Weinstein-Raun, Daniel de Haas, et al · 2022
Closest in time.
Claude’s constitution, 2023
Anthropic · 2023
Closest in time.
Managing ai risks in an era of rapid progress
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, et al · 2023
Closest in time.
Taken out of context: On measuring situational awareness in llms
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans · 2023
Closest in time.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al · 2023
Closest in time.
About, January 2023
DeepMind · 2023
Closest in time.
Gpts are gpts: An early look at the labor market impact potential of large language models, 2023
Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock · 2023
Closest in time.
Overthinking the truth: Understanding how language models process false demonstrations, 2023
Danny Halawi, Jean-Stanislas Denain, and Jacob Steinhardt · 2023
Closest in time.
An overview of catastrophic ai risks
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside · 2023
Closest in time.
Bing chat is blatantly, aggressively misaligned, 2023
Evan Hubinger · 2023
Closest in time.
Ai alignment: A comprehensive survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al · 2023
Closest in time.
Power-seeking can be probable and predictive for trained agents
Victoria Krakovna and Janos Kramar · 2023
Closest in time.
Towards a situational awareness benchmark for llms
Rudolf Laine, Alexander Meinke, and Owain Evans · 2023
Closest in time.
Tell, don’t show: Declarative facts influence how llms generalize
Alexander Meinke and Owain Evans · 2023
Closest in time.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
Alexander Pan, Chan Jun Shern, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Jonathan Ng, Hanlin Zhang, Scott Emmons, and Dan Hendrycks · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein · 2023
Closest in time.
Are emergent abilities of large language models a mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2023
Closest in time.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al · 2023
Closest in time.
Model evaluation for extreme risks
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, et al · 2023
Closest in time.
Emergent deception and emergent optimization, 2023
Jacob Steinhardt · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman · 2023
Closest in time.
Uncovering mesa-optimization algorithms in transformers
Johannes von Oswald, Eyvind Niklasson, Maximilian Schlegel, Seijin Kobayashi, Nicolas Zucchet, Nino Scherrer, Nolan Miller, Mark Sandler, Max Vladymyrov, Razvan Pascanu, et al · 2023
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al · 2023
Closest in time.
Larger language models do in-context learning differently
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al · 2023
Closest in time.
Emergence of maps in the memories of blind navigation agents
Erik Wijmans, Manolis Savva, Irfan Essa, Stefan Lee, Ari S. Morcos, and Dhruv Batra · 2023
Closest in time.
What algorithms can transformers learn? a study in length generalization
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Josh Susskind, Samy Bengio, and Preetum Nakkiran · 2023
Closest in time.
Managing extreme ai risks amid rapid progress
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, et al · 2024
Closest in time.
Looking inward: Language models can learn about themselves by introspection
Felix J Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, and Owain Evans · 2024
Closest in time.
Sparse autoencoders reveal temporal difference learning in large language models
Can Demircan, Tankred Saanum, Akshay K Jagadish, Marcel Binz, and Eric Schulz · 2024
Closest in time.
Sycophancy to subterfuge: Investigating reward-tampering in large language models
Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, et al · 2024
Closest in time.
Planning behavior in a recurrent neural network that plays sokoban
Adrià Garriga-Alonso, Mohammad Taufeeque, and Adam Gleave · 2024
Closest in time.
Alignment faking in large language models
Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, et al · 2024
Closest in time.
Introduction to AI Safety, Ethics, and Society
Dan Hendrycks · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al · 2024
Closest in time.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Closest in time.
Frontier models are capable of in-context scheming
Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn · 2024
Closest in time.
Language models learn to mislead humans via rlhf
Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R Bowman, He He, and Shi Feng · 2024
Closest in time.
Monitoring reasoning models for misbehavior and the risks of promoting obfuscation
Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi · 2025
Closest in time.
Demonstrating specification gaming in reasoning models
Alexander Bondarenko, Denis Volk, Dmitrii Volkov, and Jeffrey Ladish · 2025
Closest in time.
Me, myself, and ai: The situational awareness dataset (sad) for llms
Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hariharan, Mikita Balesni, Jérémy Scheurer, Marius Hobbhahn, Alexander Meinke, and Owain Evans · 2025
Closest in time.
Utility engineering: Analyzing and controlling emergent value systems in ais
Mantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim, Bruce W Lee, Richard Ren, Long Phan, Norman Mu, Adam Khoja, Oliver Zhang, et al · 2025
Closest in time.
Connecting the dots: Llms can infer and verbalize latent structure from disparate training data
Johannes Treutlein, Dami Choi, Jan Betley, Samuel Marks, Cem Anil, Roger B Grosse, and Owain Evans · 2025
Closest in time.