Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) requires access to a reward function that incentivizes the right behavior, but these are notoriously hard to specify for complex tasks.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, Ralph Allan and Terry, Milton E · 1952
Earlier work this paper cites.
Active learning via transductive experimental design
Yu, Kai, Bi, Jinbo, and Tresp, Volker · 2006
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, Pierre-Yves, Kaplan, Frdric, and Hafner, Verena V · 2007
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
Knox, W Bradley and Stone, Peter · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, Jürgen · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, Brian D · 2010
Earlier work this paper cites.
Preference-based policy learning
Akrour, Riad, Schoenauer, Marc, and Sebag, Michele · 2011
Earlier work this paper cites.
Policy search for motor primitives in robotics
Kober, Jens and Peters, Jan · 2011
Earlier work this paper cites.
Online human training of a myoelectric prosthesis controller via actor-critic reinforcement learning
Pilarski, Patrick M, Dawson, Michael R, Degris, Thomas, Fahimi, Farbod, Carey, Jason P, and Sutton, Richard S · 2011
Earlier work this paper cites.
An active learning algorithm for ranking from pairwise preferences with an almost optimal query complexity
Ailon, Nir · 2012
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries
Wilson, Aaron, Fern, Alan, and Tadepalli, Prasad · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, Marc G, Naddaf, Yavar, Veness, Joel, and Bowling, Michael · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, Jens, Bagnell, J Andrew, and Peters, Jan · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, Diederik P and Ba, Jimmy · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, John, Levine, Sergey, Abbeel, Pieter, Jordan, Michael, and Moritz, Philipp · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, Dario, Olah, Chris, Steinhardt, Jacob, Christiano, Paul, Schulman, John, and Mané, Dan · 2016
Earlier work this paper cites.
Beattie, Charles, Leibo, Joel Z, Teplyashin, Denis, Ward, Tom, Wainwright, Marcus, Küttler, Heinrich, Lefrancq, Andrew, Green, Simon, Valdés, Víctor, Sadik, Amir, et al · 2016
Earlier work this paper cites.
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Earlier work this paper cites.
The Oxford handbook of cognitive science
Chipman, Susan EF · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Duan, Yan, Chen, Xi, Houthooft, Rein, Schulman, John, and Abbeel, Pieter · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Yarin and Ghahramani, Zoubin · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, John, Moritz, Philipp, Levine, Sergey, Jordan, Michael, and Abbeel, Pieter · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, Christian, Vanhoucke, Vincent, Ioffe, Sergey, Shlens, Jon, and Wojna, Zbigniew · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, Paul F, Leike, Jan, Brown, Tom, Martic, Miljan, Legg, Shane, and Amodei, Dario · 2017
Earlier work this paper cites.
Inverse reward design
Hadfield-Menell, Dylan, Milli, Smitha, Abbeel, Pieter, Russell, Stuart, and Dragan, Anca · 2017
Cited alongside, same era.
Interactive learning from policy-dependent human feedback
MacGlashan, James, Ho, Mark K, Loftin, Robert, Peng, Bei, Roberts, David, Taylor, Matthew E, and Littman, Michael L · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Sadigh, Dorsa, Dragan, Anca D, Sastry, Shankar, and Seshia, Sanjit A · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, John, Wolski, Filip, Dhariwal, Prafulla, Radford, Alec, and Klimov, Oleg · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, David, Schrittwieser, Julian, Simonyan, Karen, Antonoglou, Ioannis, Huang, Aja, Guez, Arthur, Hubert, Thomas, Baker, Lucas, Lai, Matthew, Bolton, Adrian, et al · 2017
Cited alongside, same era.
Autonomous navigation of stratospheric balloons using reinforcement learning
Bellemare, Marc G, Candido, Salvatore, Castro, Pablo Samuel, Gong, Jun, Machado, Marlos C, Moitra, Subhodeep, Ponda, Sameera S, and Wang, Ziyu · 2020
Later among the works it cites.
Active preference-based gaussian process regression for reward learning
Biyik, Erdem, Huynh, Nicolas, Kochenderfer, Mykel J, and Sadigh, Dorsa · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, Tom B, Mann, Benjamin, Ryder, Nick, Subbiah, Melanie, Kaplan, Jared, Dhariwal, Prafulla, Neelakantan, Arvind, Shyam, Pranav, Sastry, Girish, Askell, Amanda, et al · 2020
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, Karl, Hesse, Chris, Hilton, Jacob, and Schulman, John · 2020
Later among the works it cites.
Derail: Diagnostic environments for reward and imitation learning
Freire, Pedro, Gleave, Adam, Toyer, Sam, and Russell, Stuart · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Biyik, Erdem and Sadigh, Dorsa · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, Tuomas, Zhou, Aurick, Abbeel, Pieter, and Levine, Sergey · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Henderson, Peter, Islam, Riashat, Bachman, Philip, Pineau, Joelle, Precup, Doina, and Meger, David · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Ibarz, Borja, Leike, Jan, Pohlen, Tobias, Irving, Geoffrey, Legg, Shane, and Amodei, Dario · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, Dmitry, Irpan, Alex, Pastor, Peter, Ibarz, Julian, Herzog, Alexander, Jang, Eric, Quillen, Deirdre, Holly, Ethan, Kalakrishnan, Mrinal, Vanhoucke, Vincent, et al · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
Leike, Jan, Krueger, David, Everitt, Tom, Martic, Miljan, Maini, Vishal, and Legg, Shane · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
Silver, David, Hubert, Thomas, Schrittwieser, Julian, Antonoglou, Ioannis, Lai, Matthew, Guez, Artfhur, Lanctot, Marc, Sifre, Laurent, Kumaran, Dharshan, Graepel, Thore, et al · 2018
Cited alongside, same era.
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, Justin, Kumar, Aviral, Nachum, Ofir, Tucker, George, and Levine, Sergey · 2020
Later among the works it cites.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Gulcehre, Caglar, Wang, Ziyu, Novikov, Alexander, Paine, Tom Le, Colmenarejo, Sergio Gomez, Zolna, Konrad, Agarwal, Rishabh, Merel, Josh, Mankowitz, Daniel, Paduraru, Cosmin, et al · 2020
Later among the works it cites.
Reinforcement learning with augmented data
Laskin, Michael, Lee, Kimin, Stooke, Adam, Pinto, Lerrel, Abbeel, Pieter, and Srinivas, Aravind · 2020
Later among the works it cites.
Learning human objectives by evaluating hypothetical behavior
Reddy, Siddharth, Dragan, Anca, Levine, Sergey, Legg, Shane, and Leike, Jan · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Srinivas, Aravind, Laskin, Michael, and Abbeel, Pieter · 2020
Later among the works it cites.
Learning to summarize from human feedback
Stiennon, Nisan, Ouyang, Long, Wu, Jeff, Ziegler, Daniel M, Lowe, Ryan, Voss, Chelsea, Radford, Alec, Amodei, Dario, and Christiano, Paul · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control
Tassa, Yuval, Tunyasuvunakool, Saran, Muldal, Alistair, Doron, Yotam, Liu, Siqi, Bohez, Steven, Merel, Josh, Erez, Tom, Lillicrap, Timothy, and Heess, Nicolas · 2020
Later among the works it cites.
Avoiding side effects in complex environments
Turner, Alexander Matt, Ratzlaff, Neale, and Tadepalli, Prasad · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online
Xu, Zhongwen, van Hasselt, Hado, Hessel, Matteo, Oh, Junhyuk, Singh, Satinder, and Silver, David · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, Tianhe, Quillen, Deirdre, He, Zhanpeng, Julian, Ryan, Hausman, Karol, Finn, Chelsea, and Levine, Sergey · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, Rishabh, Schwarzer, Max, Castro, Pablo Samuel, Courville, Aaron, and Bellemare, Marc G · 2021
Closest in time.
The impacts of known and unknown demonstrator irrationality on reward inference, 2021
Chan, Lawrence, Critch, Andrew, and Dragan, Anca · 2021
Closest in time.
URLB: Unsupervised reinforcement learning benchmark
Laskin, Michael, Yarats, Denis, Liu, Hao, Lee, Kimin, Zhan, Albert, Lu, Kevin, Cang, Catherine, Pinto, Lerrel, and Abbeel, Pieter · 2021
Closest in time.
Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
Lee, Kimin, Smith, Laura, and Abbeel, Pieter · 2021
Closest in time.
Behavior from the void: Unsupervised active pre-training
Liu, Hao and Abbeel, Pieter · 2021
Closest in time.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, Max, Anand, Ankesh, Goel, Rishab, Hjelm, R Devon, Courville, Aaron, and Bachman, Philip · 2021
Closest in time.
State entropy maximization with random encoders for efficient exploration
Seo, Younggyo, Chen, Lili, Shin, Jinwoo, Lee, Honglak, Abbeel, Pieter, and Lee, Kimin · 2021
Closest in time.
Decoupling representation learning from reinforcement learning
Stooke, Adam, Lee, Kimin, Abbeel, Pieter, and Laskin, Michael · 2021
Closest in time.
Recursively summarizing books with human feedback
Wu, Jeff, Ouyang, Long, Ziegler, Daniel M, Stiennon, Nissan, Lowe, Ryan, Leike, Jan, and Christiano, Paul · 2021
Closest in time.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, Denis, Kostrikov, Ilya, and Fergus, Rob · 2021
Closest in time.