Fetching the paper…
Reading the bibliography…
This study focuses on the topic of offline preference-based reinforcement learning (PbRL), a variant of conventional reinforcement learning that dispenses with the need for online interaction or specification of reward functions.
Kumar, A., Peng, X. B., and Levine, S · 1912
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Risk-sensitive reinforcement learning
Mihatsch, O. and Neuneier, R · 2002
Earlier work this paper cites.
Policy gradient reinforcement learning for fast quadrupedal locomotion
Kohl, N. and Stone, P · 2004
Earlier work this paper cites.
Safe exploration for reinforcement learning
Hans, A., Schneegaß, D., Schäfer, A. M., and Udluft, S · 2008
Earlier work this paper cites.
Policy search for motor primitives in robotics
Kober, J. and Peters, J · 2008
Earlier work this paper cites.
Visualizing data using t-sne
van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Wainwright, M. J., Jordan, M. I., et al · 2008
Earlier work this paper cites.
Interactively shaping agents via human reinforcement
Knox, W. B. and Stone, P · 2009
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F · 2015
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Inverse reward design
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D · 2017
Earlier work this paper cites.
Deep TAMER: Interactive agent shaping in high-dimensional state spaces
Warnell, G., Waytowich, N. R., Lawhern, V. J., and Stone, P · 2017
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in Atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Cited alongside, same era.
QT-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., and Levine, S · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Cited alongside, same era.
Deep reinforcement learning from policy-dependent human feedback
Arumugam, D., Lee, J. K., Saskin, S., and Littman, M. L · 2019
Cited alongside, same era.
Accelerating reinforcement learning with learned skill priors
Pertsch, K., Lee, Y., and Lim, J. J · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Siegel, N. Y., Springenberg, J. T., Berkenkamp, F., Abdolmaleki, A., Neunert, M., Lampe, T., Hafner, R., Heess, N., and Riedmiller, M · 2020
Later among the works it cites.
Parrot: Data-driven behavioral priors for reinforcement learning
Singh, A., Liu, H., Zhou, G., Yu, A., Rhinehart, N., and Levine, S · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Learning to reach goals without reinforcement learning
Ghosh, D., Gupta, A., Fu, J., Reddy, A., Devin, C., Eysenbach, B., and Levine, S · 2019
Cited alongside, same era.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K · 2019
Cited alongside, same era.
Training agents using upside-down reinforcement learning
Srivastava, R. K., Shyam, P., Mutz, F., Jaśkowski, W., and Schmidhuber, J · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gulcehre, C., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D · 2019
Cited alongside, same era.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Ajay, A., Kumar, A., Agrawal, P., Levine, S., and Nachum, O · 2020
Cited alongside, same era.
Explore, discover and learn: Unsupervised discovery of state-covering skills
Campos, V., Trott, A., Xiong, C., Socher, R., Giró-i Nieto, X., and Torres, J · 2020
Cited alongside, same era.
Later among the works it cites.
Rvs: What is essential for offline rl via supervised learning?
Emmons, S., Eysenbach, B., Kostrikov, I., and Levine, S · 2021
Later among the works it cites.
Generalized decision transformer for offline hindsight information matching
Furuta, H., Matsuo, Y., and Gu, S. S · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Later among the works it cites.
PEBBLE: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
Lee, K., Smith, L. M., and Abbeel, P · 2021
Later among the works it cites.
Unsupervised domain adaptation with dynamics-aware rewards in reinforcement learning
Liu, J., Shen, H., Wang, D., Kang, Y., and Tian, Q · 2021
Later among the works it cites.
Offline preference-based apprenticeship learning
Shin, D. and Brown, D. S · 2021
Later among the works it cites.
Chen, X., Zhong, H., Yang, Z., Wang, Z., and Wang, L · 2022
Later among the works it cites.
Models of human preference for learning reward functions
Knox, W. B., Hatgis-Kessell, S., Booth, S., Niekum, S., Stone, P., and Allievi, A · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2022
Later among the works it cites.
What matters in learning from offline human demonstrations for robot manipulation
Mandlekar, A., Xu, D., Wong, J., Nasiriany, S., Wang, C., Kulkarni, R., Fei-Fei, L., Savarese, S., Zhu, Y., and Martín-Martín, R · 2022
Later among the works it cites.
Park, J., Seo, Y., Shin, J., Lee, H., Abbeel, P., and Lee, K · 2022
Later among the works it cites.
Scalar reward is not enough: A response to silver, singh, precup and sutton (2021)
Vamplew, P., Smith, B. J., Källström, J., Ramos, G., Rădulescu, R., Roijers, D. M., Hayes, C. F., Heintz, F., Mannion, P., Libin, P. J., et al · 2022
Later among the works it cites.
Preference transformer: Modeling human preferences using transformers for RL
Kim, C., Park, J., Shin, J., Lee, H., Abbeel, P., and Lee, K · 2023
Closest in time.
Behavior proximal policy optimization
Zhuang, Z., LEI, K., Liu, J., Wang, D., and Guo, Y · 2023
Closest in time.