Fetching the paper…
Reading the bibliography…
Sample efficiency is a crucial problem in deep reinforcement learning.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Double Q-learning
Hado Hasselt · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
Sergey Ioffe · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, et al · 2018
Cited alongside, same era.
How does batch normalization help optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Deepmellow: Removing the need for a target network in deep Q-learning
Seungchan Kim, Kavosh Asadi, Michael L. Littman, and George Dimitri Konidaris · 2019
Dropout q-functions for doubly efficient reinforcement learning
Takuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi, and Yoshimasa Tsuruoka · 2021
Closest in time.
Training larger networks for deep reinforcement learning
Kei Ota, Devesh K Jha, and Asako Kanezaki · 2021
Closest in time.
Stable-baselines3: Reliable reinforcement learning implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Closest in time.
Data-efficient reinforcement learning with self-predictive representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, and Philip Bachman · 2021
Closest in time.
Overcoming the spectral bias of neural value approximation
Ge Yang, Anurag Ajay, and Pulkit Agrawal · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Evalnorm: Estimating batch normalization statistics for evaluation
Saurabh Singh and Abhinav Shrivastava · 2019
Cited alongside, same era.
A theoretical analysis of deep Q-learning
Jianqing Fan, Zhaoran Wang, Yuchen Xie, and Zhuoran Yang · 2020
Cited alongside, same era.
Reinforcement learning with augmented data
Misha Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas · 2020
Cited alongside, same era.
Grokking deep reinforcement learning
M. Morales · 2020
Cited alongside, same era.
Can increasing input dimensionality improve deep reinforcement learning?
Kei Ota, Tomoaki Oiki, Devesh Jha, Toshisada Mariyama, and Daniel Nikovski · 2020
Cited alongside, same era.
Four things everyone should know to improve batch normalization
Cecilia Summers and Michael J. Dinneen · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron Courville, and Marc G Bellemare · 2021
Cited alongside, same era.
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Closest in time.
Reinforcement learning with automated auxiliary loss search
Tairan He, Yuge Zhang, Kan Ren, Minghuan Liu, Che Wang, Weinan Zhang, Yuqing Yang, and Dongsheng Li · 2022
Closest in time.
Efficient deep reinforcement learning requires regulating overfitting
Qiyang Li, Aviral Kumar, Ilya Kostrikov, and Sergey Levine · 2022
Closest in time.
The primacy bias in deep reinforcement learning
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville · 2022
Closest in time.
Continual normalization: Rethinking batch normalization for online continual learning
Quang Pham, Chenghao Liu, and Steven HOI · 2022
Closest in time.
A walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning
Laura Smith, Ilya Kostrikov, and Sergey Levine · 2022
Closest in time.
TTN: A domain-shift aware batch normalization in test-time adaptation
Hyesu Lim, Byeonggeun Kim, Jaegul Choo, and Sungha Choi · 2023
Closest in time.
Gymnasium, 2023
Mark Towers, Jordan K. Terry, Ariel Kwiatkowski, John U. Balis, Gianluca de Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Arjun KG, Markus Krimmel, Rodrigo Perez-Vicente, Andrea Pierré, Sander Schulhoff, Jun Jet Tai, Andrew Tan Jin Shen, and Omar G. Younis · 2023
Closest in time.