Fetching the paper…
Reading the bibliography…
Efficient exploration remains a challenging problem in reinforcement learning, especially for tasks where extrinsic rewards from environments are sparse or even totally disregarded.
P. Dayan and G. E. Hinton, “Using expectation-maximization for reinforcement learning,” Neural Computation , vol. 9, no. 2, pp. 271–278, 1997
1997
Earlier work this paper cites.
R. M. Ryan and E. L. Deci, “Intrinsic and extrinsic motivations: Classic definitions and new directions,” Contemporary educational psychology , vol. 25, no. 1, pp. 54–67, 2000
2000
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Proceedings of Advances in neural information processing systems , 2000, pp. 1057–1063
2000
Earlier work this paper cites.
M. Toussaint and A. Storkey, “Probabilistic inference for solving discrete and continuous state markov decision processes,” in Proceedings of International Conference on Machine Learning , 2006, pp. 945–952
2006
Earlier work this paper cites.
J. Peters and S. Schaal, “Reinforcement learning by reward-weighted regression for operational space control,” in Proceedings of International Conference on Machine Learning , 2007, pp. 745–750
2007
Earlier work this paper cites.
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning.” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 8. Chicago, IL, USA, 2008, pp. 1433–1438
2008
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in Proceedings of International Conference on Learning Representations , 2014, pp. 1–14
2014
Earlier work this paper cites.
S. Levine, Motor skill learning with local trajectory methods . Stanford University, 2014
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, pp. 529–533, 2015
2015
Earlier work this paper cites.
K. Sohn, H. Lee, and X. Yan, “Learning structured output representation using deep conditional generative models,” in Proceedings of Advances in neural information processing systems , 2015, pp. 3483–3491
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in Proceedings of International conference on machine learning , 2015, pp. 1889–1897
2015
Earlier work this paper cites.
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, “High-dimensional continuous control using generalized advantage estimation,” in Proceedings of International Conference on Learning Representations , 2016, pp. 1–14
2016
Earlier work this paper cites.
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos, “Unifying count-based exploration and intrinsic motivation,” in Proceedings of Advances in Neural Information Processing Systems , 2016, pp. 1471–1479
2016
Earlier work this paper cites.
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel, “Vime: Variational information maximizing exploration,” in Proceedings of Advances in Neural Information Processing Systems , 2016, pp. 1109–1117
2016
Earlier work this paper cites.
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. R. Baker, M. Lai, A. Bolton, Y. Chen, T. P. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis, “Mastering the game of go without human knowledge,” Nature , vol. 550, pp. 354–359, 2017
2017
Earlier work this paper cites.
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell, “Curiosity-driven exploration by self-supervised prediction,” in Proceedings of International Conference on Machine Learning , 2017, pp. 2778–2787
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
G. Ostrovski, M. G. Bellemare, A. van den Oord, and R. Munos, “Count-based exploration with neural density models,” in Proceedings of International Conference on Machine Learning , 2017, pp. 2721–2730
2017
Earlier work this paper cites.
D. M. Blei, A. Kucukelbir, and J. D. McAuliffe, “Variational inference: A review for statisticians,” Journal of the American statistical Association , vol. 112, no. 518, pp. 859–877, 2017
2017
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Cited alongside, same era.
Z. Ren, D. Dong, H. Li, and C. Chen, “Self-paced prioritized curriculum learning with coverage penalty in deep reinforcement learning,” IEEE transactions on neural networks and learning systems , vol. 29, no. 6, pp. 2216–2226, 2018
2018
Cited alongside, same era.
Z. Yang, K. Merrick, L. Jin, and H. A. Abbass, “Hierarchical deep reinforcement learning for continuous action control,” IEEE transactions on neural networks and learning systems , vol. 29, no. 11, pp. 5174–5184, 2018
2018
Cited alongside, same era.
K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep reinforcement learning in a handful of trials using probabilistic dynamics models,” in Proceedings of Advances in Neural Information Processing Systems , 2018, pp. 4754–4765
2018
Cited alongside, same era.
P. Shyam, W. Jaśkowski, and F. Gomez, “Model-based active exploration,” in Proceedings of International Conference on Machine Learning , 2019, pp. 5779–5788
2019
Later among the works it cites.
2019
Later among the works it cites.
L. Beyer, D. Vincent, O. Teboul, S. Gelly, M. Geist, and O. Pietquin, “Mulex: Disentangling exploitation from exploration in deep rl,” in Proceedings of Exploration in RL Workshop of International Conference on Machine Learning , 2019, pp. 1–14
2019
Later among the works it cites.
H. Kim, J. Kim, Y. Jeong, S. Levine, and H. O. Song, “Emi: Exploration with mutual information,” in Proceedings of International Conference on Machine Learning , 2019, pp. 3360–3369
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Haber, D. Mrowca, S. Wang, L. F. Fei-Fei, and D. L. Yamins, “Learning to play with intrinsically-motivated, self-aware agents,” in Proceedings of Advances in Neural Information Processing Systems , 2018, pp. 8388–8399
2018
Cited alongside, same era.
D. P. Bertsekas, “Feature-based aggregation and deep reinforcement learning: A survey and some new implementations,” IEEE/CAA Journal of Automatica Sinica , vol. 6, no. 1, pp. 1–31, 2018
2018
Cited alongside, same era.
D. Corneil, W. Gerstner, and J. Brea, “Efficient model-based deep reinforcement learning with variational state tabulation,” in Proceedings of International Conference on Machine Learning . PMLR, 2018, pp. 1049–1058
2018
Cited alongside, same era.
P. Schulz, W. Aziz, and T. Cohn, “A stochastic decoder for neural machine translation,” in The 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 1243–1252
2018
Cited alongside, same era.
K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep reinforcement learning in a handful of trials using probabilistic dynamics models,” in Proceedings of Advances in Neural Information Processing Systems , 2018, pp. 4759–4770
2018
Cited alongside, same era.
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. Hausknecht, and M. Bowling, “Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents,” Journal of Artificial Intelligence Research , vol. 61, pp. 523–562, 2018
2018
Cited alongside, same era.
T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of gans for improved quality, stability, and variation,” in Proceedings of International Conference of Learning Representation , 2018, pp. 1–14
2018
Cited alongside, same era.
S. Gao, M. Zhou, Y. Wang, J. Cheng, H. Yachi, and J. Wang, “Dendritic neuron model with effective learning algorithms for classification, approximation, and prediction,” IEEE transactions on neural networks and learning systems , vol. 30, no. 2, pp. 601–614, 2018
2018
Cited alongside, same era.
Y. Kim, W. Nam, H. Kim, J.-H. Kim, and G. Kim, “Curiosity-bottleneck: Exploration by distilling task-specific novelty,” in Proceedings of International Conference on Machine Learning . PMLR, 2019, pp. 3379–3388
2019
Later among the works it cites.
J. Choi, Y. Guo, M. Moczulski, J. Oh, N. Wu, M. Norouzi, and H. Lee, “Contingency-aware exploration in reinforcement learning,” in Proceedings of International Conference of Learning Representation , 2019, pp. 1–14
2019
Later among the works it cites.
L. Chen, L. Wang, Z. Han, J. Zhao, and W. Wang, “Variational inference based kernel dynamic bayesian networks for construction of prediction intervals for industrial time series with incomplete input,” IEEE/CAA Journal of Automatica Sinica , vol. 7, no. 5, pp. 1437–1445, 2019
2019
Later among the works it cites.
M. Fellows, A. Mahajan, T. G. Rudner, and S. Whiteson, “Virel: A variational inference framework for reinforcement learning,” Advances in Neural Information Processing Systems , vol. 32, pp. 7122–7136, 2019
2019
Later among the works it cites.
X. Sun and B. Bischl, “Tutorial and survey on probabilistic graphical model and variational inference in deep reinforcement learning,” in Proceedings of 2019 IEEE Symposium Series on Computational Intelligence (SSCI) . IEEE, 2019, pp. 110–119
2019
Later among the works it cites.
O. Ivanov, M. Figurnov, and D. Vetrov, “Variational autoencoder with arbitrary conditioning,” in Proceedings of International Conference on Learning Representations , 2019, pp. 1–14
2019
Later among the works it cites.
H. Thanh-Tung, T. Tran, and S. Venkatesh, “Improving generalization and stability of generative adversarial networks,” in Proceedings of International Conference of Learning Representation , 2019, pp. 1–14
2019
Later among the works it cites.
J. Wang and T. Kumbasar, “Parameter optimization of interval type-2 fuzzy neural networks based on pso and bbbc methods,” IEEE/CAA Journal of Automatica Sinica , vol. 6, no. 1, pp. 247–257, 2019
2019
Later among the works it cites.
P. Liu, C. Bai, Y. Zhao, C. Bai, W. Zhao, and X. Tang, “Generating attentive goals for prioritized hindsight reinforcement learning,” Knowledge-Based Systems , vol. 203, p. 106140, 2020
2020
Closest in time.
A. P. Badia, P. Sprechmann, A. Vitvitskyi, D. Guo, B. Piot, S. Kapturowski, O. Tieleman, M. Arjovsky, A. Pritzel, A. Bolt, and C. Blundell, “Never give up: Learning directed exploration strategies,” in Proceedings of International Conference on Learning Representations , 2020, pp. 1–14
2020
Closest in time.
R. Sekar, O. Rybkin, K. Daniilidis, P. Abbeel, D. Hafner, and D. Pathak, “Planning to explore via self-supervised world models,” in Proceedings of International Conference on Machine Learning , 2020, pp. 8583–8592
2020
Closest in time.
C. Bai, L. Wang, Y. Wang, Z. Wang, R. Zhao, C. Bai, and P. Liu, “Addressing hindsight bias in multigoal reinforcement learning,” IEEE Transactions on Cybernetics , pp. 1–14, 2021
2021
Closest in time.
C. Bai, L. Wang, L. Han, J. Hao, A. Garg, P. Liu, and Z. Wang, “Principled exploration via optimistic bootstrapping and backward induction,” in Proceedings of International Conference on Machine Learning , vol. 139. PMLR, 2021, pp. 577–587
2021
Closest in time.
C. Bai, L. Wang, L. Han, A. Garg, J. Hao, P. Liu, and Z. Wang, “Dynamic bottleneck for robust self-supervised exploration,” in Proceedings of Advances in Neural Information Processing Systems , 2021, pp. 1–14
2021
Closest in time.