Fetching the paper…
Reading the bibliography…
In reinforcement learning from human feedback, it is common to optimize against a reward model trained to predict human preferences.
On evaluating adversarial robustness, 2019
Nicholas Carlini, Anish Athalye, Nicolas Papernot, Wieland Brendel, Jonas Rauber, Dimitris Tsipras, Ian Goodfellow, Aleksander Madry, and Alexey Kurakin · 1902
Earlier work this paper cites.
Reforms as experiments
Donald T Campbell · 1969
Earlier work this paper cites.
Problems of monetary management: the uk experience in papers in monetary economics
Charles Goodhart · 1975
Earlier work this paper cites.
The "awful idea of accountability" : inscribing people into the measurement of objects
Keith Hoskin · 1996
Earlier work this paper cites.
Predictably incoherent judgments
Cass R Sunstein, Daniel Kahneman, David Schkade, and Ilana Ritov · 2001
Earlier work this paper cites.
The basic ai drives
Stephen M. Omohundro · 2008
Earlier work this paper cites.
General purpose intelligence: arguing the orthogonality thesis
Stuart Armstrong et al · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom · 2014
Earlier work this paper cites.
Towards deep neural network architectures robust to adversarial examples
Shixiang Gu and Luca Rigazio · 2014
Earlier work this paper cites.
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky · 2015
Earlier work this paper cites.
Quantilizers: A safer alternative to maximizers for limited optimization
Jessica Taylor · 2016
Earlier work this paper cites.
Improving the robustness of deep neural networks via stability training
Stephan Zheng, Yang Song, Thomas Leung, and Ian Goodfellow · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Reinforcement learning with a corrupted reward channel
Tom Everitt, Victoria Krakovna, Laurent Orseau, Marcus Hutter, and Shane Legg · 2017
Earlier work this paper cites.
Tactics of adversarial attack on deep reinforcement learning agents, 2017
Yen-Chen Lin, Zhang-Wei Hong, Yuan-Hong Liao, Meng-Li Shih, Ming-Yu Liu, and Min Sun · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Occam's razor is insufficient to infer the preferences of irrational agents
Stuart Armstrong and Sören Mindermann · 2018
Earlier work this paper cites.
Adversarial attacks and defences: A survey
Anirban Chakraborty, Manaar Alam, Vishal Dey, Anupam Chattopadhyay, and Debdeep Mukhopadhyay · 2018
Cited alongside, same era.
Adversarial attack on graph structured data
Hanjun Dai, Hui Li, Tian Tian, Xin Huang, Lin Wang, Jun Zhu, and Le Song · 2018
Cited alongside, same era.
On adversarial examples for character-level neural machine translation
Javid Ebrahimi, Daniel Lowd, and Dejing Dou · 2018
Cited alongside, same era.
Generalization and regularization in dqn
Jesse Farebrother, Marlos C Machado, and Michael Bowling · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Cited alongside, same era.
Consequences of misaligned AI
Simon Zhuang and Dylan Hadfield-Menell · 2020
Later among the works it cites.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Later among the works it cites.
Eliciting latent knowledge: How to tell if your eyes deceive you, 12 2021
Paul Christiano, Ajeya Cotra, and Mark Xu · 2021
Later among the works it cites.
Gradient-based adversarial attacks against text transformers, 2021
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela · 2021
Later among the works it cites.
Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Cited alongside, same era.
Categorizing variants of goodhart’s law
David Manheim and Scott Garrabrant · 2018
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Cited alongside, same era.
Risks from learned optimization in advanced machine learning systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Cited alongside, same era.
Classifying specification problems as variants of goodhart’s law, 8 2019
Victoria Krakovna and Ramana Kumar · 2019
Cited alongside, same era.
Observational overfitting in reinforcement learning
Xingyou Song, Yiding Jiang, Stephen Tu, Yilun Du, and Behnam Neyshabur · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Later among the works it cites.
Optimal policies tend to seek power
Alex Turner, Logan Smith, Rohin Shah, Andrew Critch, and Prasad Tadepalli · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Closest in time.
Is power-seeking AI an existential risk?
Joseph Carlsmith · 2022
Closest in time.
Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover, 2022
Ajeya Cotra · 2022
Closest in time.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trebacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Soňa Mokrá, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving · 2022
Closest in time.
Uncertainty estimation for language reward models
Adam Gleave and Geoffrey Irving · 2022
Closest in time.
Rl with kl penalties is better viewed as bayesian inference
Tomasz Korbak, Ethan Perez, and Christopher L Buckley · 2022
Closest in time.
Teaching language models to support answers with verified quotes
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, et al · 2022
Closest in time.
The alignment problem from a deep learning perspective
Richard Ngo · 2022
Closest in time.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Jan Leike, and Ryan Lowe · 2022
Closest in time.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Closest in time.
Defining and characterizing reward hacking, 2022
Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger · 2022
Closest in time.