Fetching the paper…
Reading the bibliography…
Deep reinforcement learning has proven to be successful for learning tasks in simulated environments, but applying same techniques for robots in real-world domain is more challenging, as they require hours of training.
“Policy reuse for transfer learning across tasks with different state and action spaces,”
Fernando Fernández and Manuela Veloso, · 2006
Earlier work this paper cites.
“Transfer learning for reinforcement learning domains: A survey,”
Matthew E Taylor and Peter Stone, · 2009
Earlier work this paper cites.
“The arcade learning environment: An evaluation platform for general agents,”
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, · 2013
Earlier work this paper cites.
“An empirical investigation of catastrophic forgetting in gradient-based neural networks,”
Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio, · 2013
Earlier work this paper cites.
“Human-level control through deep reinforcement learning,”
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al., · 2015
Earlier work this paper cites.
“Vizdoom: A doom-based ai research platform for visual reinforcement learning,”
Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaśkowski, · 2016
Earlier work this paper cites.
“Progressive neural networks,”
Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell, · 2016
Earlier work this paper cites.
“Torchcraft: a library for machine learning research on real-time strategy games,”
Gabriel Synnaeve, Nantas Nardelli, Alex Auvolat, Soumith Chintala, Timothée Lacroix, Zeming Lin, Florian Richoux, and Nicolas Usunier, · 2016
Cited alongside, same era.
“Dueling network architectures for deep reinforcement learning,”
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas, · 2016
Cited alongside, same era.
“Domain randomization for transferring deep neural networks from simulation to the real world,”
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel, · 2017
Cited alongside, same era.
“Starcraft ii: A new challenge for reinforcement learning,”
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, et al., · 2017
Cited alongside, same era.
“Torille: Learning environment for hand-to-hand combat,”
Anssi Kanervisto and Ville Hautamäki, · 2018
Later among the works it cites.
“Airsim: High-fidelity visual and physical simulation for autonomous vehicles,”
Shital Shah, Debadeepta Dey, Chris Lovett, and Ashish Kapoor, · 2018
Later among the works it cites.
“Learning dexterous in-hand manipulation,”
OpenAI, Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al., · 2018
Later among the works it cites.
“Holodeck: A high fidelity simulator,” 2018
Joshua Greaves, Max Robinson, Nick Walton, Mitchell Mortensen, Robert Pottorff, Connor Christopherson, Derek Hancock, and Jayden Milne David Wingate, · 2018
Later among the works it cites.
“Stable baselines,” https://github.com/hill-a/stable-baselines
Ashley Hill, Antonin Raffin, Maximilian Ernestus, Adam Gleave, Anssi Kanervisto, Rene Traore, Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rika Antonova, Silvia Cruciani, Christian Smith, and Danica Kragic, · 2017
Cited alongside, same era.
“Proximal policy optimization algorithms,”
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov, · 2017
Cited alongside, same era.
Later among the works it cites.
“Deep reinforcement learning with double q-learning.,”
Hado Van Hasselt, Arthur Guez, and David Silver, · 2094
Closest in time.