Fetching the paper…
Reading the bibliography…
Because it is difficult to precisely specify complex objectives, reinforcement learning policies are often optimized using proxy reward functions that only approximate the true goal.
Lyapunov-based Safe Policy Optimization for Continuous Control, February 2019
Yinlam Chow, Ofir Nachum, Aleksandra Faust, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 1901
Earlier work this paper cites.
Efficient Exploration via State Marginal Matching, February 2020
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 1906
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library, December 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 1912
Earlier work this paper cites.
The UVA/PADOVA Type 1 Diabetes Simulator
Chiara Dalla Man, Francesco Micheletto, Dayu Lv, Marc Breton, Boris Kovatchev, and Claudio Cobelli · 1932
Earlier work this paper cites.
Algorithms for a Closed-Loop Artificial Pancreas: The Case for Proportional-Integral-Derivative Control
Garry M. Steil · 1932
Earlier work this paper cites.
Problems of Monetary Management: The UK Experience
C. A. E. Goodhart · 1984
Earlier work this paper cites.
Congested Traffic States in Empirical Observations and Microscopic Simulations
Martin Treiber, Ansgar Hennecke, and Dirk Helbing · 2000
Earlier work this paper cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning, December 2020
Hong Jun Jeon, Smitha Milli, and Anca D. Dragan · 2002
Earlier work this paper cites.
First Order Constrained Optimization in Policy Space, October 2020
Yiming Zhang, Quan Vuong, and Keith W. Ross · 2002
Earlier work this paper cites.
Leverage the Average: an Analysis of KL Regularization in RL, January 2021
Nino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin, Rémi Munos, and Matthieu Geist · 2003
Earlier work this paper cites.
Local Rademacher Complexities and Oracle Inequalities in Risk Minimization
Vladimir Koltchinskii · 2006
Earlier work this paper cites.
Avoiding Side Effects in Complex Environments, October 2020
Alexander Matt Turner, Neale Ratzlaff, and Prasad Tadepalli · 2006
Earlier work this paper cites.
From Optimizing Engagement to Measuring Value
Smitha Milli, Luca Belli, and Moritz Hardt · 2008
Earlier work this paper cites.
Deep Reinforcement Learning for Closed-Loop Blood Glucose Control, September 2020
Ian Fox, Joyce Lee, Rodica Pop-Busui, and Jenna Wiens · 2009
Earlier work this paper cites.
Learning to summarize from human feedback, September 2020
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2009
Earlier work this paper cites.
Reinforcement Learning for Optimization of COVID-19 Mitigation policies, October 2020
Varun Kompella, Roberto Capobianco, Stacy Jong, Jonathan Browne, Spencer Fox, Lauren Meyers, Peter Wurman, and Peter Stone · 2010
Earlier work this paper cites.
Artificial intelligence: a modern approach
Stuart J. Russell, Peter Norvig, and Ernest Davis · 2010
Earlier work this paper cites.
Concrete Problems in AI Safety, July 2016
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Generative Adversarial Imitation Learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
To Predict and Serve?
Kristian Lum and William Isaac · 2016
Earlier work this paper cites.
Quantilizers: A Safer Alternative to Maximizers for Limited Optimization
Jessica Taylor · 2016
Earlier work this paper cites.
Algorithmic decision making and the cost of fairness, June 2017
Sam Corbett-Davies, Emma Pierson, Avi Feller, Sharad Goel, and Aziz Huq · 2017
Earlier work this paper cites.
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart Russell, and Anca Dragan · 2017
Earlier work this paper cites.
AI Safety Gridworlds, November 2017
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A. Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Earlier work this paper cites.
Active Preference-Based Learning of Reward Functions
Dorsa Sadigh, Anca Dragan, Shankar Sastry, and Sanjit Seshia · 2017
Earlier work this paper cites.
Proximal Policy Optimization Algorithms, August 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Safe Exploration in Continuous Action Spaces, January 2018
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in Atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Earlier work this paper cites.
Policy Optimization with Demonstrations
Bingyi Kang, Zequn Jie, and Jiashi Feng · 2018
Earlier work this paper cites.
Specification gaming examples in AI, April 2018
Victoria Krakovna · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction, November 2018
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Cited alongside, same era.
RLlib: Abstractions for Distributed Reinforcement Learning, June 2018
Eric Liang, Richard Liaw, Philipp Moritz, Robert Nishihara, Roy Fox, Ken Goldberg, Joseph E. Gonzalez, Michael I. Jordan, and Ion Stoica · 2018
Cited alongside, same era.
Microscopic Traffic Simulation using SUMO
Pablo Alvarez Lopez, Michael Behrisch, Laura Bieker-Walz, Jakob Erdmann, Yun-Pang Flötteröd, Robert Hilbrich, Leonhard Lücken, Johannes Rummel, Peter Wagner, and Evamarie Wiessner · 2018
Cited alongside, same era.
Benchmarks for reinforcement learning in mixed-autonomy traffic
Eugene Vinitsky, Aboudy Kreidieh, Luc Le Flem, Nishant Kheterpal, Kathy Jang, Cathy Wu, Fangyu Wu, Richard Liaw, Eric Liang, and Alexandre M. Bayen · 2018
Cited alongside, same era.
Off-Policy Deep Reinforcement Learning without Exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Building Human Values into Recommender Systems: An Interdisciplinary Synthesis, July 2022
Jonathan Stray, Alon Halevy, Parisa Assar, Dylan Hadfield-Menell, Craig Boutilier, Amar Ashar, Lex Beattie, Michael Ekstrand, Claire Leibowicz, Connie Moon Sehat, Sara Johansen, Lianne Kerlin, David Vickrey, Spandana Singh, Sanne Vrijenhoek, Amy Zhang, McKane Andrus, Natali Helberger, Polina Proutskova, Tanushree Mitra, and Nina Vasan · 2022
Later among the works it cites.
Shentao Yang, Yihao Feng, Shujian Zhang, and Mingyuan Zhou · 2022
Later among the works it cites.
LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models, November 2023
Marwa Abdulhai, Isadora White, Charlie Snell, Charles Sun, Joey Hong, Yuexiang Zhai, Kelvin Xu, and Sergey Levine · 2023
Later among the works it cites.
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Provably Efficient Maximum Entropy Exploration, January 2019
Elad Hazan, Sham M. Kakade, Karan Singh, and Abby Van Soest · 2019
Cited alongside, same era.
Classifying specification problems as variants of Goodhart’s Law, August 2019
Victoria Krakovna · 2019
Cited alongside, same era.
Penalizing side effects using stepwise relative reachability, March 2019
Victoria Krakovna, Laurent Orseau, Ramana Kumar, Miljan Martic, and Shane Legg · 2019
Cited alongside, same era.
Categorizing Variants of Goodhart’s Law, February 2019
David Manheim and Scott Garrabrant · 2019
Cited alongside, same era.
Dissecting racial bias in an algorithm used to manage the health of populations
Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan · 2019
Cited alongside, same era.
Optimal Policies Tend To Seek Power
A. M. Turner, Logan Smith, Rohin Shah, Andrew Critch, and Prasad Tadepalli · 2019
Cited alongside, same era.
Learning Human Objectives by Evaluating Hypothetical Behavior
Siddharth Reddy, Anca Dragan, Sergey Levine, Shane Legg, and Jan Leike · 2020
Cited alongside, same era.
Later among the works it cites.
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S. Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
Detecting disparities in police deployments using dashcam data
Matt Franchi, J. D. Zamfirescu-Pereira, Wendy Ju, and Emma Pierson · 2023
Later among the works it cites.
Aligning Language Models with Preferences through f-divergence Minimization, June 2023
Dongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen, Nahyeon Ryu, and Marc Dymetman · 2023
Later among the works it cites.
trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback
Alexander Havrilla, Maksym Zhuravinskyi, Duy Phung, Aman Tiwari, Jonathan Tow, Stella Biderman, Quentin Anthony, and Louis Castricato · 2023
Later among the works it cites.
A Survey on Offline Model-Based Reinforcement Learning, May 2023
Haoyang He · 2023
Later among the works it cites.
Jon Kleinberg, Sendhil Mullainathan, and Manish Raghavan · 2023
Later among the works it cites.
Bridging RL Theory and Practice with the Effective Horizon, April 2023
Cassidy Laidlaw, Stuart Russell, and Anca Dragan · 2023
Later among the works it cites.
Performative Reinforcement Learning, February 2023
Debmalya Mandal, Stelios Triantafyllou, and Goran Radanovic · 2023
Later among the works it cites.
On The Fragility of Learned Reward Functions, January 2023
Lev McKinney, Yawen Duan, David Krueger, and Adam Gleave · 2023
Later among the works it cites.
k-Means Maximum Entropy Exploration, November 2023
Alexander Nedergaard and Matthew Cook · 2023
Later among the works it cites.
The alignment problem from a deep learning perspective, September 2023
Richard Ngo, Lawrence Chan, and Sören Mindermann · 2023
Later among the works it cites.
Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of Pessimism, July 2023
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2023
Later among the works it cites.
Causal Confusion and Reward Misidentification in Preference-Based Reward Learning, March 2023
Jeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca D. Dragan, and Daniel S. Brown · 2023
Later among the works it cites.
Jump-Start Reinforcement Learning, July 2023
Ikechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu, Mengyuan Yan, Joséphine Simon, Matthew Bennice, Chuyuan Fu, Cong Ma, Jiantao Jiao, Sergey Levine, and Karol Hausman · 2023
Later among the works it cites.
Bellman-consistent Pessimism for Offline Reinforcement Learning, October 2023
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2023
Later among the works it cites.
Reward Model Ensembles Help Mitigate Overoptimization, March 2024
Thomas Coste, Usman Anwar, Robert Kirk, and David Krueger · 2024
Closest in time.
Lukas Fluri, Leon Lang, Alessandro Abate, Patrick Forré, David Krueger, and Joar Skalse · 2024
Closest in time.
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization, February 2024
Uri Gadot, Esther Derman, Navdeep Kumar, Maxence Mohamed Elfatihi, Kfir Levy, and Shie Mannor · 2024
Closest in time.
Audrey Huang, Wenhao Zhan, Tengyang Xie, Jason D. Lee, Wen Sun, Akshay Krishnamurthy, and Dylan J. Foster · 2024
Closest in time.
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
Hamish Ivison, Yizhong Wang, Jiacheng Liu, Zeqiu Wu, Valentina Pyatkin, Nathan Lambert, Noah A. Smith, Yejin Choi, and Hannaneh Hajishirzi · 2024
Closest in time.
Thomas Kwa, Drake Thomas, and Adrià Garriga-Alonso · 2024
Closest in time.
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban, Georgios Tzannetos, Goran Radanovic, and Adish Singla · 2024
Closest in time.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model, July 2024
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2024
Closest in time.
Multi-turn Reinforcement Learning from Preference Human Feedback
Lior Shani, Aviv Rosenberg, Asaf Cassel, Oran Lang, Daniele Calandriello, Avital Zipori, Hila Noga, Orgad Keller, Bilal Piot, and Idan Szpektor · 2024
Closest in time.
STARC: A General Framework For Quantifying Differences Between Reward Functions, December 2024
Joar Skalse, Lucy Farnik, Sumeet Ramesh Motwani, Erik Jenner, Adam Gleave, and Alessandro Abate · 2024
Closest in time.