Fetching the paper…
Reading the bibliography…
Reward hacking -- where RL agents exploit gaps in misspecified reward functions -- has been widely observed, but not yet systematically studied.
More is different
Philip W Anderson · 1972
Earlier work this paper cites.
Risk analysis of blood glucose data:a quantitative approach to optimizing the control of insulin dependent diabetes
BorIs. P. Kovatchev, Martin Straume, Daniel J. Cox, and Leon.S Farhy · 2000
Earlier work this paper cites.
Congested traffic states in empirical observations and microscopic simulations
Martin Treiber, Ansgar Hennecke, and Dirk Helbing · 2000
Earlier work this paper cites.
The relationship between precision-recall and roc curves
Jesse Davis and Mark Goadrich · 2006
Earlier work this paper cites.
Self-organizing systems: The emergence of order
F Eugene Yates · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
The UVA/PADOVA type 1 diabetes simulator: New features
Chiara Dalla Man, Francesco Micheletto, Dayu Lv, Marc Breton, Boris Kovatchev, and Claudio Cobelli · 2014
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
The potential cost implications of averting severe hypoglycemic events requiring hospitalization in high-risk adults with type 1 diabetes using real-time continuous glucose monitoring
Amy Bronstone and Claudia Graham · 2016
Earlier work this paper cites.
Quantilizers: A safer alternative to maximizers for limited optimization
Jessica Taylor · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Reinforcement learning with a corrupted reward channel
Tom Everitt, Victoria Krakovna, Laurent Orseau, and Shane Legg · 2017
Earlier work this paper cites.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Dan Hendrycks and Kevin Gimpel · 2017
Earlier work this paper cites.
AI safety gridworlds, 2017
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A. Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Robert Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in Atari
Borja Ibarz, J. Leike, Tobias Pohlen, Geoffrey Irving, S. Legg, and Dario Amodei · 2018
Cited alongside, same era.
Microscopic traffic simulation using SUMO
Pablo Alvarez Lopez, Michael Behrisch, Laura Bieker-Walz, Jakob Erdmann, Yun-Pang Flötteröd, Robert Hilbrich, Leonhard Lücken, Johannes Rummel, Peter Wagner, and Evamarie Wießner · 2018
Cited alongside, same era.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Cited alongside, same era.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher · 2018
Cited alongside, same era.
Benchmarks for reinforcement learning in mixed-autonomy traffic
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Reinforcement learning for optimization of covid-19 mitigation policies, 2020
Varun Kompella, Roberto Capobianco, Stacy Jong, Jonathan Browne, Spencer Fox, Lauren Meyers, Peter Wurman, and Peter Stone · 2020
Later among the works it cites.
How much does insulin cost? Here’s how 23 brands compare, Nov 2020
Benita Lee · 2020
Later among the works it cites.
Auditing radicalization pathways on youtube
Manoel Horta Ribeiro, Raphael Ottoni, Robert West, Virgílio A. F. Almeida, and Wagner Meira · 2020
Later among the works it cites.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eugene Vinitsky, Aboudy Kreidieh, Luc Le Flem, Nishant Kheterpal, Kathy Jang, Cathy Wu, Fangyu Wu, Richard Liaw, Eric Liang, and Alexandre M. Bayen · 2018
Cited alongside, same era.
The U.S. Insulin Crisis - Rationing a Lifesaving Medication Discovered in the 1920s
M. Fralick and A. S. Kesselheim · 2019
Cited alongside, same era.
Cost-related insulin underuse among patients with diabetes
Darby Herkert, Pavithra Vijayakumar, Jing Luo, Jeremy I. Schwartz, Tracy L. Rabin, Eunice DeFilippo, and Kasia J. Lipska · 2019
Cited alongside, same era.
Risks from learned optimization in advanced machine learning systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Cited alongside, same era.
TorchBeast: A PyTorch Platform for Distributed RL
Heinrich Küttler, Nantas Nardelli, Thibaut Lavril, Marco Selvatici, Viswanath Sivakumar, Tim Rocktäschel, and Edward Grefenstette · 2019
Cited alongside, same era.
Human Compatible: Artificial Intelligence and the Problem of Control
Stuart Russell · 2019
Cited alongside, same era.
Is deep reinforcement learning really superhuman on Atari? Leveling the playing field, 2019
Marin Toromanoff, Emilie Wirbel, and Fabien Moutarde · 2019
Cited alongside, same era.
Later among the works it cites.
Aligning ai optimization to community well-being
Jonathan Stray · 2020
Later among the works it cites.
Csi: Novelty detection via contrastive learning on distributionally shifted instances
Jihoon Tack, Sangwoo Mo, Jongheon Jeong, and Jinwoo Shin · 2020
Later among the works it cites.
Consequences of misaligned AI
Simon Zhuang and Dylan Hadfield-Menell · 2020
Later among the works it cites.
The political economy of the covid-19 pandemic
Peter Boettke and Benjamin Powell · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani et al · 2021
Later among the works it cites.
Scaling laws for acoustic models
Jasha Droppo and Oguz Elibol · 2021
Later among the works it cites.
Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish · 2021
Later among the works it cites.
Policy gradient bayesian robust optimization for imitation learning
Zaynah Javed, Daniel S Brown, Satvik Sharma, Jerry Zhu, Ashwin Balakrishna, Marek Petrik, Anca Dragan, and Ken Goldberg · 2021
Later among the works it cites.
Reward (Mis)design for Autonomous Driving
W. Bradley Knox, Alessandro Allievi, Holger Banzhaf, Felix Schmitt, and Peter Stone · 2021
Later among the works it cites.
Gathering strength, gathering storms: The one hundred year study on artificial intelligence (AI100) 2021 study panel report
Michael L. Littman, Ifeoma Ajunwa, Guy Berger, Craig Boutilier, Morgan Currie, Finale Doshi-Velez, Gillian Hadfield, Michael C. Horowitz, Charles Isbell, Hiroaki Kitano, Karen Levy, Terah Lyons, Melanie Mitchell, Julie Shah, Steven Sloman, Shannon Vallor, and Toby Walsh · 2021
Later among the works it cites.
Alexander Trott, Sunil Srinivasa, Douwe van der Wal, Sebastien Haneuse, and Stephan Zheng · 2021
Later among the works it cites.
Flow: A modular learning framework for mixed autonomy traffic
Cathy Wu, Abdul Rahman Kreidieh, Kanaad Parvate, Eugene Vinitsky, and Alexandre M. Bayen · 2021
Later among the works it cites.