Fetching the paper…
Reading the bibliography…
Policies for complex visual tasks have been successfully learned with deep reinforcement learning, using an approach called deep Q-networks (DQN), but relatively large (task-specific) networks and extensive training are needed to achieve good performance.
How not to lie with statistics: The correct way to summarize benchmark results
Philip J. Fleming and John J. Wallace · 1986
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Handbook of learning and approximate dynamic programming
A. G. Barto and T. G. Dietterich · 2004
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L. Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
Model compression
Cristian Bucila, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Structure compilation: trading structure for features
Percy Liang, Hal Daumé III, and Dan Klein · 2008
Earlier work this paper cites.
Visualizing high-dimensional data using t-sne
L.J.P. van der Maaten and G.E. Hinton · 2008
Earlier work this paper cites.
Improving supervised learning by adapting the problem to the learner
Joshua Menke and Tony Martinez · 2009
Cited alongside, same era.
Analysis of a classification-based policy iteration algorithm
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2010
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey J Gordon, and J Andrew Bagnell · 2010
Cited alongside, same era.
Generalized classification-based approximate policy iteration
Amir-massoud Farahmand, Doina Precup, and Mohammad Ghavamzadeh · 2012
Cited alongside, same era.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Learning small-size dnn with output-distribution-based criteria
Jinyu Li, Rui Zhao, Jui-Ting Huang, and Yifan Gong · 2014
Later among the works it cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Later among the works it cites.
SelfieBoost: A Boosting Algorithm for Deep Learning
S. Shalev-Shwartz · 2014
Later among the works it cites.
Transferring knowledge from a rnn to a dnn
William Chan, Nan Rosemary Ke, and Ian Lane · 2015
Closest in time.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Cited alongside, same era.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Xiaoxiao Guo, Satinder P. Singh, Honglak Lee, Richard L. Lewis, and Xiaoshi Wang · 2014
Cited alongside, same era.
Distilling the Knowledge in a Neural Network
G. Hinton, O. Vinyals, and J. Dean · 2014
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver
Cited in the paper.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, Shane Legg, Volodymyr Mnih, Koray Kavukcuoglu, and David Silver · 2015
Closest in time.
Knowledge transfer pre-training
Zhiyuan Tang, Dong Wang, Yiqiao Pan, and Zhiyong Zhang · 2015
Closest in time.
Recurrent neural network training with dark knowledge transfer
Dong Wang, Chao Liu, Zhiyuan Tang, Zhiyong Zhang, and Mengyuan Zhao · 2015
Closest in time.