Fetching the paper…
Reading the bibliography…
Training large neural networks is known to be time-consuming, with the learning duration taking days or even weeks.
Algorithm as 177: Expected normal order statistics (exact and approximate)
J. P. Royston · 1982
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.
Ucb exploration via q-ensembles
Richard Y Chen, Szymon Sidor, Pieter Abbeel, and John Schulman · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Shallow updates for deep reinforcement learning
Nir Levine, Tom Zahavy, Daniel J Mankowitz, Aviv Tamar, and Shie Mannor · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
Distributed deep reinforcement learning: Learn how to play atari games in 21 minutes
Igor Adamski, Robert Adamski, Tomasz Grel, Adam Jędrych, Kamil Kaczmarek, and Henryk Michalewski · 2018
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Distributed prioritized experience replay
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado Van Hasselt, and David Silver · 2018
Earlier work this paper cites.
Model-ensemble trust-region policy optimization
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel · 2018
Earlier work this paper cites.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Earlier work this paper cites.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Cited alongside, same era.
Accelerated methods for deep reinforcement learning
Adam Stooke and Pieter Abbeel · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
Better exploration with optimistic actor critic
Kamil Ciosek, Quan Vuong, Robert Loftin, and Katja Hofmann · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Randomized ensembled double q-learning: Learning fast without a model
Xinyue Chen, Che Wang, Zijian Zhou, and Keith Ross · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Later among the works it cites.
Dropout q-functions for doubly efficient reinforcement learning
Takuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi, and Yoshimasa Tsuruoka · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2021
Later among the works it cites.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
Kimin Lee, Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Bidirectional model-based policy optimization
Hang Lai, Jian Shen, Weinan Zhang, and Yong Yu · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Takuma Seno and Michita Imai · 2021
Later among the works it cites.
Pulserl: Enabling offline reinforcement learning for digital marketing systems via conservative q-learning
Douglas W Soares, Acordo Certo, Telma Lima, and Deep Learning Brazil · 2021
Later among the works it cites.
Plas: Latent action space for offline reinforcement learning
Wenxuan Zhou, Sujay Bajracharya, and David Held · 2021
Later among the works it cites.
Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Chenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhihong Deng, Animesh Garg, Peng Liu, and Zhaoran Wang · 2022
Closest in time.
Seyed Kamyar Seyed Ghasemipour, Shixiang Shane Gu, and Ofir Nachum · 2022
Closest in time.
Showing your offline reinforcement learning work: Online evaluation budget matters
Vladislav Kurenkov and Sergey Kolesnikov · 2022
Closest in time.
Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble
Seunghyun Lee, Younggyo Seo, Kimin Lee, Pieter Abbeel, and Jinwoo Shin · 2022
Closest in time.
Reducing variance in temporal-difference value estimation via ensemble of deep networks
Litian Liang, Yaosheng Xu, Stephen McAleer, Dailin Hu, Alexander Ihler, Pieter Abbeel, and Roy Fox · 2022
Closest in time.
Can wikipedia help offline reinforcement learning?
Machel Reid, Yutaro Yamada, and Shixiang Shane Gu · 2022
Closest in time.
Offline reinforcement learning as anti-exploration
Shideh Rezaeifar, Robert Dadashi, Nino Vieillard, Léonard Hussenot, Olivier Bachem, Olivier Pietquin, and Matthieu Geist · 2022
Closest in time.
DNS: Determinantal point process based neural network sampler for ensemble reinforcement learning
Hassam Sheikh, Kizza Frisbee, and Mariano Phielipp · 2022
Closest in time.
Deepthermal: Combustion optimization for thermal power generating units using offline reinforcement learning
Xianyuan Zhan, Haoran Xu, Yue Zhang, Xiangyu Zhu, Honglei Yin, and Yu Zheng · 2022
Closest in time.