Fetching the paper…
Reading the bibliography…
Pre-training Reinforcement Learning agents in a task-agnostic manner has shown promising results.
Auto-encoding variational bayes, 2014
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Empowerment–an introduction
C. Salge, C. Glackin, and D. Polani · 2014
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
S. Mohamed and D. J. Rezende · 2015
Earlier work this paper cites.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Stochastic neural networks for hierarchical reinforcement learning
C. Florensa, Y. Duan, and P. Abbeel · 2017
Earlier work this paper cites.
Variational intrinsic control
K. Gregor, D. J. Rezende, and D. Wierstra · 2017
Earlier work this paper cites.
Control of a quadrotor with reinforcement learning
J. Hwangbo, I. Sa, R. Siegwart, and M. Hutter · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. van den Oord, O. Vinyals, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Variational option discovery algorithms, 2018
J. Achiam, H. Edwards, D. Amodei, and P. Abbeel · 2018
Earlier work this paper cites.
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2018
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Cited alongside, same era.
Learning actionable representations with goal conditioned policies
D. Ghosh, A. Gupta, and S. Levine · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Unsupervised control through non-parametric discriminative rewards
D. Warde-Farley, T. Van de Wiele, T. Kulkarni, C. Ionescu, S. Hansen, and V. Mnih · 2018
Cited alongside, same era.
Leveraging procedural generation to benchmark reinforcement learning
K. Cobbe, C. Hesse, J. Hilton, and J. Schulman · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
M. Laskin, A. Srinivas, and P. Abbeel · 2020
Later among the works it cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
A. Lee, A. Nagabandi, P. Abbeel, and S. Levine · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2020
Later among the works it cites.
Decoupling representation learning from reinforcement learning
A. Stooke, K. Lee, P. Abbeel, and M. Laskin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, et al · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, et al · 2019
Cited alongside, same era.
Keeping your distance: Solving sparse reward tasks using self-balancing shaped rewards
A. Trott, S. Zheng, C. Xiong, and R. Socher · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Explore, discover and learn: Unsupervised discovery of state-covering skills
V. Campos, A. Trott, C. Xiong, R. Socher, X. Giro-i Nieto, and J. Torres · 2020
Cited alongside, same era.
V. Campos, P. Sprechmann, S. Hansen, A. Barreto, S. Kapturowski, A. Vitvitskyi, A. P. Badia, and C. Blundell · 2021
Closest in time.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Closest in time.
An empirical study of training self-supervised visual transformers
X. Chen, S. Xie, and K. He · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Closest in time.
The minerl 2020 competition on sample efficient reinforcement learning using human priors
W. H. Guss, M. Y. Castro, S. Devlin, B. Houghton, N. S. Kuno, C. Loomis, S. Milani, S. Mohanty, K. Nakata, R. Salakhutdinov, et al · 2021
Closest in time.
Reinforcement learning with prototypical representations
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Closest in time.