Fetching the paper…
Reading the bibliography…
Sample efficiency and exploration remain major challenges in online reinforcement learning (RL).
A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Playback control of force teachable robots
Asada, H. and Hanafusa, H · 1979
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, S. and Schwartz, A · 1993
Earlier work this paper cites.
Learning from demonstration
Schaal, S · 1996
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2005
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Ross, S. and Bagnell, J. A · 2012
Earlier work this paper cites.
Guided policy search
Levine, S. and Koltun, V · 2013
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Večerík, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Rothörl, T., Lampe, T., and Riedmiller, M · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Cited alongside, same era.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., Dulac-Arnold, G., Agapiou, J., Leibo, J., and Gruslys, A · 2018
Cited alongside, same era.
Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., and Levine, S · 2018
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Nair, A., McGrew, B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Cited alongside, same era.
Aw-opt: Learning robotic skills with imitation andreinforcement at scale
Lu, Y., Hausman, K., Chebotar, Y., Yan, M., Jang, E., Herzog, A., Xiao, T., Irpan, A., Khansari, M., Kalashnikov, D., and Levine, S · 2021
Later among the works it cites.
A graph placement methodology for fast chip design
Mirhoseini, A., Goldie, A., Yazgan, M., Jiang, J. W., Songhori, E. M., Wang, S., Lee, Y.-J., Johnson, E., Pathak, O., Nazi, A., Pak, J., Tong, A., Srinivasa, K., Hang, W., Tuncer, E., Le, Q. V., Laudon, J., Ho, R., Carpenter, R., and Dean, J · 2021
Later among the works it cites.
Tactical optimism and pessimism for deep reinforcement learning
Moskovitz, T., Parker-Holder, J., Pacchiano, A., Arbel, M., and Jordan, M · 2021
Later among the works it cites.
On pathologies in KL-regularized reinforcement learning from expert demonstrations
Rudner, T. G. J., Lu, C., Osborne, M., Gal, Y., and Teh, Y. W · 2021
Later among the works it cites.
Stabilizing off-policy deep reinforcement learning from pixels
Cetin, E., Ball, P. J., Roberts, S., and Celiktutan, O · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2018
Cited alongside, same era.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Cabi, S., Colmenarejo, S. G., Novikov, A., Konyushkova, K., Reed, S. E., Jeong, R., Zolna, K., Aytar, Y., Budden, D., Vecerík, M., Sushkov, O. O., Barker, D., Scholz, J., Denil, M., de Freitas, N., and Wang, Z · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Implementation matters in deep rl: A case study on ppo and trpo
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
AWAC: Accelerating online reinforcement learning with offline datasets
Nair, A., Gupta, A., Dalal, M., and Levine, S · 2020
Cited alongside, same era.
What matters for on-policy deep actor-critic methods? a large-scale study
Andrychowicz, M., Raichuk, A., Stańczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., Gelly, S., and Bachem, O · 2021
Cited alongside, same era.
Modem: Accelerating visual model-based reinforcement learning with demonstrations
Hansen, N., Lin, Y., Su, H., Wang, X., Kumar, V., and Rajeswaran, A · 2022
Later among the works it cites.
Dropout q-functions for doubly efficient reinforcement learning
Hiraoka, T., Imagawa, T., Hashimoto, T., Onishi, T., and Tsuruoka, Y · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2022
Later among the works it cites.
Efficient deep reinforcement learning requires regulating statistical overfitting
Li, Q., Kumar, A., Kostrikov, I., and Levine, S · 2022
Later among the works it cites.
Challenges and opportunities in offline reinforcement learning from visual observations
Lu, C., Ball, P. J., Rudner, T. G. J., Parker-Holder, J., Osborne, M. A., and Teh, Y. W · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Offline reinforcement learning as anti-exploration
Rezaeifar, S., Dadashi, R., Vieillard, N., Hussenot, L., Bachem, O., Pietquin, O., and Geist, M · 2022
Later among the works it cites.
Leveraging offline data in online reinforcement learning
Wagenmaker, A. and Pacchiano, A · 2022
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2022
Later among the works it cites.
Hybrid RL: Using both offline and online data can make RL efficient
Song, Y., Zhou, Y., Sekhari, A., Bagnell, D., Krishnamurthy, A., and Sun, W · 2023
Closest in time.
Policy expansion for bridging offline-to-online reinforcement learning
Zhang, H., Xu, W., and Yu, H · 2023
Closest in time.