Count-based exploration with neural density models
Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Later among the works it cites.
Uncertainty-aware reinforcement learning for collision avoidance
Original
Gregory Kahn, Adam Villaflor, Vitchyr Pong, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Building machines that learn and think like people
Brenden M. Lake, Tomer D. Ullman, Joshua B. Tenenbaum, and Samuel J Gershman · 2017
Later among the works it cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Original
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Deep neuroevolution: Genetic algorithms are a competitive alternative for training deep neural networks for reinforcement learning
Original
Felipe Petroski Such, Vashisht Madhavan, Edoardo Conti, Joel Lehman, Kenneth O. Stanley, and Jeff Clune · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Original
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Later among the works it cites.
Intrinsically motivated goal exploration processes with automatic curriculum learning
Original
Sébastien Forestier, Yoan Mollard, and Pierre-Yves Oudeyer · 2017
Later among the works it cites.
Back to basics: Benchmarking canonical evolution strategies for playing atari
Patryk Chrabaszcz, Ilya Loshchilov, and Frank Hutter · 2018
Later among the works it cites.
Exploration by random network distillation
Original
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Later among the works it cites.
Deep q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, Gabriel Dulac-Arnold, John Agapiou, Joel Z. Leibo, and Audrunas Gruslys · 2018
Later among the works it cites.
Observe and look further: Achieving consistent performance on atari
Original
Tobias Pohlen, Bilal Piot, Todd Hester, Mohammad Gheshlaghi Azar, Dan Horgan, David Budden, Gabriel Barth-Maron, Hado van Hasselt, John Quan, Mel Večerík, et al · 2018
Later among the works it cites.
Learning montezuma’s revenge from a single demonstration
Original
Tim Salimans and Richard Chen · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew J. Hausknecht, and Michael H. Bowling · 2018
Later among the works it cites.
Deep curiosity search: Intra-life exploration improves performance on challenging deep reinforcement learning problems
Original
Christopher Stanton and Jeff Clune · 2018
Later among the works it cites.
The uncertainty bellman equation and exploration
Brendan O’Donoghue, Ian Osband, Rémi Munos, and Volodymyr Mnih · 2018
Later among the works it cites.
Contingency-aware exploration in reinforcement learning
Original
Jongwook Choi, Yijie Guo, Marcin Moczulski, Junhyuk Oh, Neal Wu, Mohammad Norouzi, and Honglak Lee · 2018
Later among the works it cites.
Distributed prioritized experience replay
Original
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado van Hasselt, and David Silver · 2018
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Later among the works it cites.
Playing hard exploration games by watching youtube
Original
Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, and Nando de Freitas · 2018
Later among the works it cites.
An atari model zoo for analyzing, visualizing, and comparing deep reinforcement learning agents
Original
Felipe Petroski Such, Vashisht Madhavan, Rosanne Liu, Rui Wang, Pablo Samuel Castro, Yulun Li, Ludwig Schubert, Marc Bellemare, Jeff Clune, and Joel Lehman · 2018
Later among the works it cites.
Montezuma’s revenge solved by go-explore, a new algorithm for hard-exploration problems (sets records on pitfall, too)
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune · 2018
Later among the works it cites.
Robustness to out-of-distribution inputs via task-aware generative uncertainty
Original
Rowan McAllister, Gregory Kahn, Jeff Clune, and Sergey Levine · 2018
Later among the works it cites.
Learning dexterous in-hand manipulation
Original
Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2018
Later among the works it cites.
Backplay: "man muss immer umkehren"
Original
Cinjon Resnick, Roberta Raileanu, Sanyam Kapoor, Alex Peysakhovich, Kyunghyun Cho, and Joan Bruna · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, and David Silver · 2018
Later among the works it cites.
Trial without error: Towards safe reinforcement learning via human intervention
William Saunders, Girish Sastry, Andreas Stuhlmueller, and Owain Evans · 2018
Later among the works it cites.
GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms
Cédric Colas, Olivier Sigaud, and Pierre-Yves Oudeyer · 2018
Later among the works it cites.
Self-imitation learning
Junhyuk Oh, Yijie Guo, Satinder Singh, and Honglak Lee · 2018
Later among the works it cites.
Recall traces: Backtracking models for efficient reinforcement learning
Original
Anirudh Goyal, Philemon Brakel, William Fedus, Timothy P. Lillicrap, Sergey Levine, Hugo Larochelle, and Yoshua Bengio · 2018
Later among the works it cites.
Fractal ai: A fragile theory of intelligence
Original
Sergio Hernandez Cerezo and Guillem Duran Ballester · 2018
Later among the works it cites.
Taking the scenic route: Automatic exploration for videogames
Original
Zeping Zhan, Batu Aytemiz, and Adam M Smith · 2018
Later among the works it cites.
Back to basics: Benchmarking canonical evolution strategies for playing atari
Patryk Chrabaszcz, Ilya Loshchilov, and Frank Hutter · 2018
Later among the works it cites.
Exploration by random network distillation
Original
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Later among the works it cites.
Deep q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, Gabriel Dulac-Arnold, John Agapiou, Joel Z. Leibo, and Audrunas Gruslys · 2018
Later among the works it cites.
Observe and look further: Achieving consistent performance on atari
Original
Tobias Pohlen, Bilal Piot, Todd Hester, Mohammad Gheshlaghi Azar, Dan Horgan, David Budden, Gabriel Barth-Maron, Hado van Hasselt, John Quan, Mel Večerík, et al · 2018
Later among the works it cites.
Learning montezuma’s revenge from a single demonstration
Original
Tim Salimans and Richard Chen · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew J. Hausknecht, and Michael H. Bowling · 2018
Later among the works it cites.
Deep curiosity search: Intra-life exploration improves performance on challenging deep reinforcement learning problems
Original
Christopher Stanton and Jeff Clune · 2018
Later among the works it cites.
The uncertainty bellman equation and exploration
Brendan O’Donoghue, Ian Osband, Rémi Munos, and Volodymyr Mnih · 2018
Later among the works it cites.
Contingency-aware exploration in reinforcement learning
Original
Jongwook Choi, Yijie Guo, Marcin Moczulski, Junhyuk Oh, Neal Wu, Mohammad Norouzi, and Honglak Lee · 2018
Later among the works it cites.
Distributed prioritized experience replay
Original
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado van Hasselt, and David Silver · 2018
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Later among the works it cites.
Playing hard exploration games by watching youtube
Original
Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, and Nando de Freitas · 2018
Later among the works it cites.
An atari model zoo for analyzing, visualizing, and comparing deep reinforcement learning agents
Original
Felipe Petroski Such, Vashisht Madhavan, Rosanne Liu, Rui Wang, Pablo Samuel Castro, Yulun Li, Ludwig Schubert, Marc Bellemare, Jeff Clune, and Joel Lehman · 2018
Later among the works it cites.
Montezuma’s revenge solved by go-explore, a new algorithm for hard-exploration problems (sets records on pitfall, too)
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune · 2018
Later among the works it cites.
Robustness to out-of-distribution inputs via task-aware generative uncertainty
Original
Rowan McAllister, Gregory Kahn, Jeff Clune, and Sergey Levine · 2018
Later among the works it cites.
Learning dexterous in-hand manipulation
Original
Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2018
Later among the works it cites.
Backplay: "man muss immer umkehren"
Original
Cinjon Resnick, Roberta Raileanu, Sanyam Kapoor, Alex Peysakhovich, Kyunghyun Cho, and Joan Bruna · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, and David Silver · 2018
Later among the works it cites.
Trial without error: Towards safe reinforcement learning via human intervention
William Saunders, Girish Sastry, Andreas Stuhlmueller, and Owain Evans · 2018
Later among the works it cites.
GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms
Cédric Colas, Olivier Sigaud, and Pierre-Yves Oudeyer · 2018
Later among the works it cites.
Self-imitation learning
Junhyuk Oh, Yijie Guo, Satinder Singh, and Honglak Lee · 2018
Later among the works it cites.
Recall traces: Backtracking models for efficient reinforcement learning
Original
Anirudh Goyal, Philemon Brakel, William Fedus, Timothy P. Lillicrap, Sergey Levine, Hugo Larochelle, and Yoshua Bengio · 2018
Later among the works it cites.
Fractal ai: A fragile theory of intelligence
Original
Sergio Hernandez Cerezo and Guillem Duran Ballester · 2018
Later among the works it cites.
Taking the scenic route: Automatic exploration for videogames
Original
Zeping Zhan, Batu Aytemiz, and Adam M Smith · 2018
Later among the works it cites.
Learning abstract models for long-horizon exploration, 2019
Evan Zheran Liu, Ramtin Keramati, Sudarshan Seshadri, Kelvin Guu, Panupong Pasupat, Emma Brunskill, and Percy Liang · 2019
Closest in time.
Fast exploration with simplified models and approximately optimistic planning in model based reinforcement learning, 2019
Ramtin Keramati, Jay Whang, Patrick Cho, and Emma Brunskill · 2019
Closest in time.
Explicit recall for efficient exploration, 2019
Honghua Dong, Jiayuan Mao, Xinyue Cui, and Lihong Li · 2019
Closest in time.
Learning abstract models for long-horizon exploration, 2019
Evan Zheran Liu, Ramtin Keramati, Sudarshan Seshadri, Kelvin Guu, Panupong Pasupat, Emma Brunskill, and Percy Liang · 2019
Closest in time.
Fast exploration with simplified models and approximately optimistic planning in model based reinforcement learning, 2019
Ramtin Keramati, Jay Whang, Patrick Cho, and Emma Brunskill · 2019
Closest in time.
Explicit recall for efficient exploration, 2019
Honghua Dong, Jiayuan Mao, Xinyue Cui, and Lihong Li · 2019
Closest in time.