Fetching the paper…
Reading the bibliography…
Building generally capable agents is a grand challenge for deep reinforcement learning (RL).
On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other
H. B. Mann and D. R. Whitney · 1947
Earlier work this paper cites.
An analysis of approximations for maximizing submodular set functions—i
G. L. Nemhauser, L. A. Wolsey, and M. L. Fisher · 1978
Earlier work this paper cites.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Making the world differentiable: On using self-supervised fully recurrent neural networks for dynamic reinforcement learning and planning in non-stationary environments
J. Schmidhuber · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
R. S. Sutton · 1991
Earlier work this paper cites.
Cost-sensitive learning of classification knowledge and its applications in robotics
M. Tan · 1993
Earlier work this paper cites.
Introduction to Reinforcement Learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
N. Houlsby, F. Huszar, Z. Ghahramani, and M. Lengyel · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Earlier work this paper cites.
Analysis of thompson sampling for the multi-armed bandit problem
S. Agrawal and N. Goyal · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
J. Kober, J. A. Bagnell, and J. Peters · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Offline policy evaluation across representations with applications to educational games
T. Mandel, Y.-E. Liu, S. Levine, E. Brunskill, and Z. Popovic · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Robots that can adapt like animals
A. Cully, J. Clune, D. Tarapore, and J.-B. Mouret · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. Bellemare, A. Graves, M. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
S. Mohamed and D. J. Rezende · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
J. Oh, X. Guo, H. Lee, R. Lewis, and S. Singh · 2015
Earlier work this paper cites.
Faulty Reward Functions in the Wild
D. Amodei and J. Clark · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Earlier work this paper cites.
Quality diversity: A new frontier for evolutionary computation
J. K. Pugh, L. B. Soros, and K. O. Stanley · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Variational intrinsic control
K. Gregor, D. J. Rezende, and D. Wierstra · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Collective robot reinforcement learning with distributed asynchronous guided policy search
A. Yahya, A. Li, M. Kalakrishnan, Y. Chebotar, and S. Levine · 2017
Earlier work this paper cites.
Variational option discovery algorithms
J. Achiam, H. Edwards, D. Amodei, and P. Abbeel · 2018
Earlier work this paper cites.
Minimalistic gridworld environment for OpenAI Gym
M. Chevalier-Boisvert, L. Willems, and S. Pal · 2018
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Earlier work this paper cites.
Improving exploration in Evolution Strategies for deep reinforcement learning via a population of novelty-seeking agents
E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. O. Stanley, and J. Clune · 2018
Earlier work this paper cites.
Scalable coordinated exploration in concurrent reinforcement learning
M. Dimakopoulou, I. Osband, and B. Van Roy · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
D. Ha and J. Schmidhuber · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Procedural level generation improves generality of deep reinforcement learning
N. Justesen, R. R. Torrado, P. Bontrager, A. Khalifa, J. Togelius, and S. Risi · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Earlier work this paper cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
S. Levine, P. Pastor, A. Krizhevsky, J. Ibarz, and D. Quillen · 2018
Earlier work this paper cites.
Learning dexterous in-hand manipulation
OpenAI, M. Andrychowicz, B. Baker, M. Chociej, R. Józefowicz, B. McGrew, J. W. Pachocki, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba · 2018
Cited alongside, same era.
Assessing generalization in deep reinforcement learning
C. Packer, K. Gao, J. Kos, P. Krähenbühl, V. Koltun, and D. Song · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. de Oliveira Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang · 2019
Cited alongside, same era.
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2019
Cited alongside, same era.
First return then explore
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2021
Later among the works it cites.
Medical dead-ends and learning to identify high-risk states and treatments
M. Fatemi, T. W. Killian, J. Subramanian, and M. Ghassemi · 2021
Later among the works it cites.
Model-value inconsistency as a signal for epistemic uncertainty
A. Filos, E. Vértes, Z. Marinho, G. Farquhar, D. Borsa, A. Friesen, F. Behbahani, T. Schaul, A. Barreto, and S. Osindero · 2021
Later among the works it cites.
Adversarially guided actor-critic
Y. Flet-Berliac, J. Ferret, O. Pietquin, P. Preux, and M. Geist · 2021
Later among the works it cites.
Collective intelligence for deep learning: A survey of recent developments
D. Ha and Y. Tang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imagined value gradients: Model-based policy optimization with transferable latent dynamics models
A. Byravan, J. T. Springenberg, A. Abdolmaleki, R. Hafner, M. Neunert, T. Lampe, N. Siegel, N. M. O. Heess, and M. A. Riedmiller · 2019
Cited alongside, same era.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
I. Clavera, A. Nagabandi, S. Liu, R. S. Fearing, P. Abbeel, S. Levine, and C. Finn · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Cited alongside, same era.
Provably efficient maximum entropy exploration
E. Hazan, S. Kakade, K. Singh, and A. Van Soest · 2019
Cited alongside, same era.
Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning
A. Kirsch, J. Van Amersfoort, and Y. Gal · 2019
Cited alongside, same era.
Efficient exploration via state marginal matching, 2019
L. Lee, B. Eysenbach, E. Parisotto, E. Xing, S. Levine, and R. Salakhutdinov · 2019
Cited alongside, same era.
Deep exploration via randomized value functions
I. Osband, B. V. Roy, D. J. Russo, and Z. Wen · 2019
Cited alongside, same era.
Mastering atari with discrete world models
D. Hafner, T. P. Lillicrap, M. Norouzi, and J. Ba · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
M. Janner, Q. Li, and S. Levine · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
R. Kirk, A. Zhang, E. Grefenstette, and T. Rocktäschel · 2021
Later among the works it cites.
URLB: Unsupervised reinforcement learning benchmark
M. Laskin, D. Yarats, H. Liu, K. Lee, A. Zhan, K. Lu, C. Cang, L. Pinto, and P. Abbeel · 2021
Later among the works it cites.
Aps: Active pretraining with successor features
H. Liu and P. Abbeel · 2021
Later among the works it cites.
Behavior from the void: Unsupervised active pre-training, 2021
H. Liu and P. Abbeel · 2021
Later among the works it cites.
Cooperative exploration for multi-agent deep reinforcement learning
I.-J. Liu, U. Jain, R. A. Yeh, and A. Schwing · 2021
Later among the works it cites.
Deployment-efficient reinforcement learning via model-based offline optimization
T. Matsushima, H. Furuta, Y. Matsuo, O. Nachum, and S. Gu · 2021
Later among the works it cites.
Discovering and achieving goals via world models, 2021
R. Mendonca, O. Rybkin, K. Daniilidis, D. Hafner, and D. Pathak · 2021
Later among the works it cites.
Model-free representation learning and exploration in low-rank mdps, 2021
A. Modi, J. Chen, A. Krishnamurthy, N. Jiang, and A. Agarwal · 2021
Later among the works it cites.
Efficient wasserstein natural gradients for reinforcement learning
T. Moskovitz, M. Arbel, F. Huszar, and A. Gretton · 2021
Later among the works it cites.
On reward-free rl with kernel and neural function approximations: Single-agent mdp and markov game
S. Qiu, J. Ye, Z. Wang, and Z. Yang · 2021
Later among the works it cites.
Offline reinforcement learning from images with latent space models
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn · 2021
Later among the works it cites.
Maximum entropy model-based reinforcement learning
O. Svidchenko and A. Shpilman · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
O. E. L. Team, A. Stooke, A. Mahajan, C. Barros, C. Deck, J. Bauer, J. Sygnowski, M. Trebacz, M. Jaderberg, M. Mathieu, N. McAleese, N. Bradley-Schmieg, N. Wong, N. Porcel, R. Raileanu, S. Hughes-Fitt, V. Dalibard, and W. M. Czarnecki · 2021
Later among the works it cites.
Learning one representation to optimize all rewards
A. Touati and Y. Ollivier · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Later among the works it cites.
Exploration by maximizing renyi entropy for reward-free RL framework
C. Zhang, Y. Cai, L. Huang, and J. Li · 2021
Later among the works it cites.
Noveld: A simple yet effective exploration criterion
T. Zhang, H. Xu, X. Wang, Y. Wu, K. Keutzer, J. E. Gonzalez, and Y. Tian · 2021
Later among the works it cites.
Do as i can and not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, and M. Yan · 2022
Closest in time.
Procedural generalization by planning with self-supervised world models
A. Anand, J. C. Walker, Y. Li, E. Vértes, J. Schrittwieser, S. Ozair, T. Weber, and J. B. Hamrick · 2022
Closest in time.
Information prioritization through empowerment in visual model-based RL
H. Bharadhwaj, M. Babaeizadeh, D. Erhan, and S. Levine · 2022
Closest in time.
Magnetic control of tokamak plasmas through deep reinforcement learning
J. Degrave, F. Felici, J. Buchli, M. Neunert, B. Tracey, F. Carpanese, T. Ewalds, R. Hafner, A. Abdolmaleki, D. de las Casas, C. Donner, L. Fritz, C. Galperti, A. Huber, J. Keeling, M. Tsimpoukelli, J. Kay, A. Merle, J.-M. Moret, S. Noury, F. Pesamosca, D. Pfau, O. Sauter, C. Sommariva, S. Coda, B. Duval, A. Fasoli, P. Kohli, K. Kavukcuoglu, D. Hassabis, and M. Riedmiller · 2022
Closest in time.
Generalized decision transformer for offline hindsight information matching
H. Furuta, Y. Matsuo, and S. S. Gu · 2022
Closest in time.
Benchmarking the spectrum of agent capabilities
D. Hafner · 2022
Closest in time.
The challenges of exploration for offline reinforcement learning
N. Lambert, M. Wulfmeier, W. F. Whitney, A. Byravan, M. Bloesch, V. Dasagi, T. Hertweck, and M. A. Riedmiller · 2022
Closest in time.
Dynamics-aware quality-diversity for efficient learning of skill repertoires
B. Lim, L. Grillotti, L. Bernasconi, and A. Cully · 2022
Closest in time.
A generalist agent, 2022
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas · 2022
Closest in time.
Can wikipedia help offline reinforcement learning?, 2022
M. Reid, Y. Yamada, and S. S. Gu · 2022
Closest in time.
The phenomenon of policy churn, 2022
T. Schaul, A. Barreto, J. Quan, and G. Ostrovski · 2022
Closest in time.
Learning more skills through optimistic exploration
D. Strouse, K. Baumli, D. Warde-Farley, V. Mnih, and S. S. Hansen · 2022
Closest in time.
MURO: Deployment constrained reinforcement learning with model-based uncertainty regularized batch optimization, 2022
D. Su, J. D. Lee, J. Mulvey, and H. V. Poor · 2022
Closest in time.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning
D. Yarats, D. Brandfonbrener, H. Liu, M. Laskin, P. Abbeel, A. Lazaric, and L. Pinto · 2022
Closest in time.
Discovering policies with domino: Diversity optimization maintaining near optimality
T. Zahavy, Y. Schroecker, F. Behbahani, K. Baumli, S. Flennerhag, S. Hou, and S. Singh · 2022
Closest in time.
Continuously discovering novel strategies via reward-switching policy optimization
Z. Zhou, W. Fu, B. Zhang, and Y. Wu · 2022
Closest in time.