Fetching the paper…
Reading the bibliography…
The ability to act in multiple environments and transfer previous knowledge to new situations can be considered a critical aspect of any intelligent agent.
A stochastic approximation method
Robbins, Herbert and Monro, Sutton · 1951
Earlier work this paper cites.
Sensitivity analysis, ergodicity coefficients, and rank-one updates for finite markov chains
Seneta, E · 1991
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Bertsekas, Dimitri P · 1995
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, Richard S. and Barto, Andrew G · 1998
Earlier work this paper cites.
A convergent form of approximate policy iteration
Perkins, Theodore J and Precup, Doina · 2002
Earlier work this paper cites.
Autonomous shaping: Knowledge transfer in reinforcement learning
Konidaris, George and Barto, Andrew G · 2006
Earlier work this paper cites.
General game learning using knowledge transfer
Banerjee, Bikramjit and Stone, Peter · 2007
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, Matthew E and Stone, Peter · 2009
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, Stephane, Gordon, Geoffrey, and Bagnell, Andrew · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, Marc G., Naddaf, Yavar, Veness, Joel, and Bowling, Michael · 2013
Cited alongside, same era.
Guided policy search
Levine, Sergey and Koltun, Vladlen · 2013
Cited alongside, same era.
Do deep nets really need to be deep?
Ba, Jimmy and Caruana, Rich · 2014
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Distilling the knowledge in a neural network
Hinton, Geoffrey, Vinyals, Oriol, and Dean, Jeff · 2015
Closest in time.
Adam: A method for stochastic optimization
Kingma, Diederik P. and Ba, Jimmy · 2015
Closest in time.
End-to-end training of deep visuomotor policies
Levine, Sergey, Finn, Chelsea, Darrell, Trevor, and Abbeel, Pieter · 2015
Closest in time.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P., Hunt, Jonathan J., Pritzel, Alexander, Heess, Nicholas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2015
Closest in time.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A., Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K., Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guo, Xiaoxiao, Singh, Satinder, Lee, Honglak, Lewis, Richard L, and Wang, Xiaoshi · 2014
Cited alongside, same era.
Closest in time.
Fitnets: Hints for thin deep nets
Romero, Adriana, Ballas, Nicolas, Kahou, Samira Ebrahimi, Chassang, Antoine, Gatta, Carlo, and Bengio, Yoshua · 2015
Closest in time.