Fetching the paper…
Reading the bibliography…
Deep reinforcement learning has led to several recent breakthroughs, though the learned policies are often based on black-box neural networks.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W · 1983
Earlier work this paper cites.
Behavioural cloning: phenomena, results and problems
Bratko, Ivan, Urbančič, Tanja, and Sammut, Claude · 1995
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, Pieter and Ng, Andrew Y · 2004
Earlier work this paper cites.
Program repair as a game
Jobstmann, Barbara, Griesmayer, Andreas, and Bloem, Roderick · 2005
Earlier work this paper cites.
Automatically finding patches using genetic programming
Weimer, Westley, Nguyen, ThanhVu, Le Goues, Claire, and Forrest, Stephanie · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, Stéphane, Gordon, Geoffrey, and Bagnell, Drew · 2011
Earlier work this paper cites.
Automated feedback generation for introductory programming assignments
Singh, Rishabh, Gulwani, Sumit, and Solar-Lezama, Armando · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A., Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin A., Fidjeland, Andreas, Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Cited alongside, same era.
Rusu, Andrei A, Colmenarejo, Sergio Gomez, Gulcehre, Caglar, Desjardins, Guillaume, Kirkpatrick, James, Pascanu, Razvan, Mnih, Volodymyr, Kavukcuoglu, Koray, and Hadsell, Raia · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, John, Levine, Sergey, Abbeel, Pieter, Jordan, Michael, and Moritz, Philipp · 2015
Cited alongside, same era.
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Cited alongside, same era.
Faulty reward functions in the wild
Clark, Jack and Amodei, Dario · 2016
Mastering the game of go with deep neural networks and tree search
Silver, David, Huang, Aja, Maddison, Chris J., Guez, Arthur, Sifre, Laurent, van den Driessche, George, Schrittwieser, Julian, Antonoglou, Ioannis, Panneershelvam, Vedavyas, Lanctot, Marc, Dieleman, Sander, Grewe, Dominik, Nham, John, Kalchbrenner, Nal, Sutskever, Ilya, Lillicrap, Timothy P., Leach, Madeleine, Kavukcuoglu, Koray, Graepel, Thore, and Hassabis, Demis · 2016
Later among the works it cites.
Inverse reward design
Hadfield-Menell, Dylan, Milli, Smitha, Abbeel, Pieter, Russell, Stuart J, and Dragan, Anca · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, John, Wolski, Filip, Dhariwal, Prafulla, Radford, Alec, and Klimov, Oleg · 2017
Later among the works it cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, David, Hubert, Thomas, Schrittwieser, Julian, Antonoglou, Ioannis, Lai, Matthew, Guez, Arthur, Lanctot, Marc, Sifre, Laurent, Kumaran, Dharshan, Graepel, Thore, Lillicrap, Timothy P., Simonyan, Karen, and Hassabis, Demis · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generative adversarial imitation learning
Ho, Jonathan and Ermon, Stefano · 2016
Cited alongside, same era.
Bastani, Osbert, Pu, Yewen, and Solar-Lezama, Armando · 2018
Closest in time.
Programmatically interpretable reinforcement learning
Verma, Abhinav, Murali, Vijayaraghavan, Singh, Rishabh, Kohli, Pushmeet, and Chaudhuri, Swarat · 2018
Closest in time.