Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has been shown to be effective at learning control from experience.
S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in AISTATS , 2011
2011
Earlier work this paper cites.
E. Parisotto, J. L. Ba, and R. Salakhutdinov, “Actor-mimic: Deep multitask and transfer reinforcement learning,” in ICLR , 2016
2016
Earlier work this paper cites.
M. G. Bellemare, W. Dabney, and R. Munos, “A distributional perspective on reinforcement learning,” in ICML , 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, K. Hartikainen, et al. , “Soft actor-critic algorithms and applications,” in ICML , 2018
2018
Earlier work this paper cites.
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller, “Maximum a posteriori policy optimisation,” in ICLR , 2018
2018
Earlier work this paper cites.
S. Fujimoto, D. Meger, and D. Precup, “Off-policy deep reinforcement learning without exploration,” in ICML , 2019
2019
Earlier work this paper cites.
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine, “Stabilizing off-policy q-learning via bootstrapping error reduction,” in NeurIPS , 2019
2019
Cited alongside, same era.
N. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, N. Heess, and M. Riedmiller, “Keep doing what worked: Behavior modelling priors for offline reinforcement learning,” in ICLR , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative Q-learning for offline reinforcement learning,” in NeurIPS , 2020
2020
Cited alongside, same era.
T. Xie, N. Jiang, H. Wang, C. Xiong, and Y. Bai, “Policy finetuning: Bridging sample-efficient offline and online reinforcement learning,” in NeurIPS , 2021
2021
Later among the works it cites.
A. X. Lee, C. M. Devin, Y. Zhou, T. Lampe, K. Bousmalis, J. T. Springenberg, et al. , “Beyond pick-and-place: Tackling robotic stacking of diverse shapes,” in CoRL , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
S. K. S. Ghasemipour, D. Schuurmans, and S. S. Gu, “EMaQ: Expected-max Q-learning operator for simple yet effective offline and online RL,” in ICML , 2021
2021
Later among the works it cites.
S. Fujimoto and S. S. Gu, “A minimalist approach to offline reinforcement learning,” in NeurIPS , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Z. Wang et al. , “Critic regularized regression,” in NeurIPS , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
D. Tirumala et al. , “Behavior priors for efficient reinforcement learning,” arXiv:2010.14274 , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin, “Offline-to-online reinforcement learning via balanced replay and pessimistic Q-ensemble,” in CoRL , 2021
2021
Later among the works it cites.