Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) experiments have notoriously high variance, and minor details can have disproportionately large effects on measured outcomes.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Earlier work this paper cites.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Oron Anschel, Nir Baram, and Nahum Shimkin · 2017
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Riashat Islam, Peter Henderson, Maziar Gomrokchi, and Doina Precup · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Spinning Up in Deep Reinforcement Learning
Joshua Achiam · 2018
Earlier work this paper cites.
Understanding batch normalization
Johan Bjorck, Carla Gomes, Bart Selman, and Kilian Q Weinberger · 2018
Earlier work this paper cites.
Stabilizing reinforcement learning in dynamic environment with application to online recommendation
Shi-Yong Chen, Yang Yu, Qing Da, Jun Tan, Hai-Kuan Huang, and Hai-Hong Tang · 2018
Earlier work this paper cites.
Let’s play again: Variability of deep reinforcement learning agents in atari environments
Kaleigh Clary, Emma Tosch, John Foley, and David Jensen · 2018
Earlier work this paper cites.
How many random seeds? statistical power analysis in deep reinforcement learning experiments
Cédric Colas, Olivier Sigaud, and Pierre-Yves Oudeyer · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Simple random search of static linear policies is competitive for reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Cited alongside, same era.
Improving stability in deep reinforcement learning with weight averaging
Evgenii Nikishin, Pavel Izmailov, Ben Athiwaratkun, Dmitrii Podoprikhin, Timur Garipov, Pavel Shvechikov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Deep contextualized word representations
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
Variance reduction for deep q-learning using stochastic recursive gradient
Haonan Jia, Xiao Zhang, Jun Xu, Wei Zeng, Hao Jiang, Xiaohui Yan, and Ji-Rong Wen · 2020
Later among the works it cites.
Self-supervised visual feature learning with deep neural networks: A survey
Longlong Jing and Yingli Tian · 2020
Later among the works it cites.
Evaluating the performance of reinforcement learning algorithms
Scott Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang, and Philip Thomas · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Ilya Kostrikov, Denis Yarats, and Rob Fergus · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Learning representations by maximizing mutual information across views
Philip Bachman, R Devon Hjelm, and William Buchwalter · 2019
Cited alongside, same era.
A hitchhiker’s guide to statistical comparisons of reinforcement learning algorithms
Cédric Colas, Olivier Sigaud, and Pierre-Yves Oudeyer · 2019
Cited alongside, same era.
Implementation matters in deep rl: A case study on ppo and trpo
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, G. Tucker, and Sergey Levine · 2019
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Cited alongside, same era.
Self-assembly behavior of an oligothiophene-based conjugated liquid crystal and its implication for ionic conductivity characteristics
Ziwei Liu, Ban Xuan Dong, Mayank Misra, Yangyang Sun, Joseph Strzalka, Shrayesh N Patel, Fernando A Escobedo, Paul F Nealey, and Christopher K Ober · 2019
Cited alongside, same era.
Deep reinforcement learning-based beam tracking for low-latency services in vehicular networks
Yan Liu, Zhiyuan Jiang, Shunqing Zhang, and Shugong Xu · 2020
Later among the works it cites.
A survey on reproducibility by evaluating deep reinforcement learning algorithms on real-world robots
Nicolai A Lynnerup, Laura Nolling, Rasmus Hasle, and John Hallam · 2020
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Emilio Parisotto, Francis Song, Jack Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, et al · 2020
Later among the works it cites.
Learning to be safe: Deep rl with a safety critic
Krishnan Srinivasan, Benjamin Eysenbach, Sehoon Ha, Jie Tan, and Chelsea Finn · 2020
Later among the works it cites.
dm control: Software and tasks for continuous control, 2020
Yuval Tassa, Saran Tunyasuvunakool, Alistair Muldal, Yotam Doron, Siqi Liu, Steven Bohez, Josh Merel, Tom Erez, Timothy Lillicrap, and Nicolas Heess · 2020
Later among the works it cites.
Striving for simplicity and performance in off-policy drl: Output normalization and non-uniform sampling
Che Wang, Yanqiu Wu, Quan Vuong, and Keith Ross · 2020
Later among the works it cites.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron Courville, and Marc G Bellemare · 2021
Closest in time.
Efficient and targeted covid-19 border testing via reinforcement learning, 2021
Hamsa Bastani, Kimon Drakopoulos, Vishal Gupta, Jon Vlachogiannis, Christos Hadjicristodoulou, Pagona Lagiou, Gkikas Magiorkinis, Dimitrios Paraskevis, and Sotirios Tsiodras · 2021
Closest in time.
Spectral normalisation for deep reinforcement learning: an optimisation perspective
Florin Gogianu, Tudor Berariu, Mihaela Rosca, Claudia Clopath, Lucian Busoniu, and Razvan Pascanu · 2021
Closest in time.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2021
Closest in time.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2021
Closest in time.
On the adequacy of untuned warmup for adaptive optimization
Jerry Ma and Denis Yarats · 2021
Closest in time.
Data-efficient reinforcement learning with momentum predictive representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, and Philip Bachman · 2021
Closest in time.