Fetching the paper…
Reading the bibliography…
A promising approach to solving challenging long-horizon tasks has been to extract behavior priors (skills) by fitting generative models to large offline datasets of demonstrations.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. F. Christiano, and G. Irving · 1909
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 1910
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 1911
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
W. B. Knox and P. Stone · 2009
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
Online human training of a myoelectric prosthesis controller via actor-critic reinforcement learning
P. M. Pilarski, M. R. Dawson, T. Degris, F. Fahimi, J. P. Carey, and R. S. Sutton · 2011
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Concrete problems in ai safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
L. Pinto and A. Gupta · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Earlier work this paper cites.
Inverse reward design
D. Hadfield-Menell, S. Milli, P. Abbeel, S. Russell, and A. Dragan · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Cited alongside, same era.
Interactive learning from policy-dependent human feedback
J. MacGlashan, M. K. Ho, R. Loftin, B. Peng, D. Roberts, M. E. Taylor, and M. L. Littman · 2017
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2018
Cited alongside, same era.
Openai five
OpenAI · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
Solving rubik’s cube with a robot hand
I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, et al · 2019
Later among the works it cites.
Deep reinforcement learning from policy-dependent human feedback
D. Arumugam, J. K. Lee, S. Saskin, and M. L. Littman · 2019
Later among the works it cites.
Avoiding side effects in complex environments
A. M. Turner, N. Ratzlaff, and P. Tadepalli · 2020
Later among the works it cites.
Accelerating reinforcement learning with learned skill priors
K. Pertsch, Y. Lee, and J. J. Lim · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Ibarz, J. Leike, T. Pohlen, G. Irving, S. Legg, and D. Amodei · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg · 2018
Cited alongside, same era.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
G. Warnell, N. Waytowich, V. Lawhern, and P. Stone · 2018
Cited alongside, same era.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen · 2018
Cited alongside, same era.
Variational inverse control with events: A general framework for data-driven reward definition
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine · 2018
Cited alongside, same era.
Few-shot goal inference for visuomotor learning and planning
A. Xie, A. Singh, S. Levine, and C. Finn · 2018
Cited alongside, same era.
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
N. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, and M. A. Riedmiller · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Human preference scaling with demonstrations for deep reinforcement learning
Z. Cao, K. Wong, and C.-T. Lin · 2020
Later among the works it cites.
Awac: Accelerating online reinforcement learning with offline datasets, 2020
A. Nair, A. Gupta, M. Dalal, and S. Levine · 2020
Later among the works it cites.
K. Lee, L. Smith, and P. Abbeel · 2021
Closest in time.
Parrot: Data-driven behavioral priors for reinforcement learning
A. Singh, H. Liu, G. Zhou, A. Yu, N. Rhinehart, and S. Levine · 2021
Closest in time.
{OPAL}: Offline primitive discovery for accelerating offline reinforcement learning
A. Ajay, A. Kumar, P. Agrawal, S. Levine, and O. Nachum · 2021
Closest in time.