Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) makes it possible to train agents capable of achieving sophisticated goals in complex and uncertain environments.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1989
Earlier work this paper cites.
Unsupervised classifiers, mutual information and ’phantom targets’
J. S. Bridle, A. J. Heading, and D. J. MacKay · 1992
Earlier work this paper cites.
Signature verification using a “siamese” time delay neural network
Jane Bromley, James W Bentz, Léon Bottou, Isabelle Guyon, Yann LeCun, Cliff Moore, Eduard Säckinger, and Roopak Shah · 1993
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Stefan Schaal · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Ng, S. Russell, et al · 2000
Earlier work this paper cites.
Like me?-measures of correspondence and imitation
Chrystopher L Nehaniv and Kerstin Dautenhahn · 2001
Earlier work this paper cites.
Understanding “prior intentions” enables two–year–olds to imitatively learn a complex task
Malinda Carpenter, Josep Call, and Michael Tomasello · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Ng · 2004
Earlier work this paper cites.
Kernelized infomax clustering
D. Barber and F. V. Agakov · 2005
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
Sumit Chopra, Raia Hadsell, and Yann LeCun · 2005
Earlier work this paper cites.
Learning shared latent structure for image synthesis and robotic imitation
Aaron Shon, Keith Grochow, Aaron Hertzmann, and Rajesh P Rao · 2005
Earlier work this paper cites.
Maximum margin planning
N. Ratliff, J. A. Bagnell, and M. A. Zinkevich · 2006
Earlier work this paper cites.
On learning, representing, and generalizing a task in a humanoid robot
Sylvain Calinon, Florent Guenter, and Aude Billard · 2007
Earlier work this paper cites.
Nine billion correspondence problems
Chrystopher L Nehaniv · 2007
Earlier work this paper cites.
Bayesian inverse reinforcement learning
D. Ramachandran and E. Amir · 2007
Earlier work this paper cites.
Boosting structured prediction for imitation learning
N. Ratliff, D. Bradley, J. A. Bagnell, and J. Chestnutt · 2007
Earlier work this paper cites.
Cross-domain video concept detection using adaptive svms
Jun Yang, Rong Yan, and Alexander G Hauptmann · 2007
Cited alongside, same era.
Maximum entropy inverse reinforcement learning
B. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey · 2008
Cited alongside, same era.
A survey of robot learning from demonstration
Brenna D Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2009
Cited alongside, same era.
Robot programming by demonstration
Sylvain Calinon · 2009
Cited alongside, same era.
Domain adaptation: Learning bounds and algorithms
Yishay Mansour, Mehryar Mohri, and Afshin Rostamizadeh · 2009
Cited alongside, same era.
Autonomous helicopter aerobatics through apprenticeship learning
Pieter Abbeel, Adam Coates, and Andrew Y Ng · 2010
Decaf: A deep convolutional activation feature for generic visual recognition
Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell · 2014
Later among the works it cites.
Unsupervised domain adaptation by backpropagation
Y. Ganin and V. Lempitsky · 2014
Later among the works it cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Later among the works it cites.
Deep domain confusion: Maximizing for domain invariance
Eric Tzeng, Judy Hoffman, Ning Zhang, Kate Saenko, and Trevor Darrell · 2014
Later among the works it cites.
Direct loss minimization inverse optimal control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Discriminative clustering by regularized information maximization
A. Krause, P. Perona, and R. G. Gomes · 2010
Cited alongside, same era.
Tabula rasa: Model transfer for object category detection
Yusuf Aytar and Andrew Zisserman · 2011
Cited alongside, same era.
Relative entropy inverse reinforcement learning
A. Boularias, J. Kober, and J. Peters · 2011
Cited alongside, same era.
What you saw is not what you get: Domain adaptation using asymmetric kernel transforms
Brian Kulis, Kate Saenko, and Trevor Darrell · 2011
Cited alongside, same era.
Nonlinear inverse reinforcement learning with gaussian processes
S. Levine, Z. Popovic, and V. Koltun · 2011
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey J Gordon, and Drew Bagnell · 2011
Cited alongside, same era.
A. Doerr, N. Ratliff, J. Bohg, M. Toussaint, and S. Schaal · 2015
Later among the works it cites.
Learning transferable features with deep adaptation networks
Mingsheng Long and Jianmin Wang · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Towards adapting deep visuomotor representations from simulated to real environments
Eric Tzeng, Coline Devin, Judy Hoffman, Chelsea Finn, Xingchao Peng, Pieter Abbeel, Sergey Levine, Kate Saenko, and Trevor Darrell · 2015
Later among the works it cites.
Maximum entropy deep inverse reinforcement learning
M. Wulfmeier, P. Ondruska, and I. Posner · 2015
Later among the works it cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Later among the works it cites.
Guided cost learning: Deep inverse optimal control via policy optimization
C. Finn, S. Levine, and P. Abbeel · 2016
Later among the works it cites.
Learning dexterous manipulation for a soft robotic hand from human demonstration
Abhishek Gupta, Clemens Eppner, Sergey Levine, and Pieter Abbeel · 2016
Later among the works it cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.