Fetching the paper…
Reading the bibliography…
Robotic reinforcement learning (RL) holds the promise of enabling robots to learn complex behaviors through experience.
G. Hayes and J. Demiris, “A robot controller using learning by imitation,” 1994
1994
Earlier work this paper cites.
Y. Kuniyoshi, M. Inaba, and H. Inoue, “Learning by watching: Extracting reusable task knowledge from visual observation of human performance,” IEEE Transactions on Robotics and Automation , 1994
1994
Earlier work this paper cites.
H. Miyamoto, S. Schaal, F. Gandolfo, H. Gomi, Y. Koike, R. Osu, E. Nakano, Y. Wada, and M. Kawato, “A Kendama learning robot based on bi-directional theory,” Neural Networks , 1996
1996
Earlier work this paper cites.
A. Billard and M. Mataric, “Learning human arm movements by imitation: Evaluation of a biologically inspired connectionist architecture,” Robotics and Autonomous Systems , 2001
2001
Earlier work this paper cites.
C. Nehaniv and K. Dautenhahn, “The correspondence problem,” Imitation in Animals and Artifacts , 2002
2002
Earlier work this paper cites.
S. Schaal, A. Ijspeert, and A. Billard, “Computational approaches to motor learning by imitation,” Philosophical Transaction of the Royal Society of London, Series B , 2003
2003
Earlier work this paper cites.
N. Kohl and P. Stone, “Policy gradient reinforcement learning for fast quadrupedal locomotion,” in ICRA , 2004
2004
Earlier work this paper cites.
R. Rubinstein and D. Kroese, The Cross Entropy Method: A Unified Approach To Combinatorial Optimization, Monte-Carlo Simulation (Information Science and Statistics) . Springer-Verlag New York, Inc., 2004
2004
Earlier work this paper cites.
B. Argall, S. Chernova, M. Veloso, and B. Browning, “A survey of robot learning from demonstration,” Robotics and Autonomous Systems , 2009
2009
Earlier work this paper cites.
S. Calinon, P. Evrard, E. Gribovskaya, A. Billard, and A. Kheddar, “Learning collaborative manipulation tasks by demonstration using a haptic interface,” in ICAR , 2009
2009
Earlier work this paper cites.
J. Kober and J. Peters, “Learning motor primitives for robotics,” in ICRA , 2009
2009
Earlier work this paper cites.
P. Kormushev, S. Calinon, and D. Caldwell, “Robot motor skill coordination with EM-based reinforcement learning,” in IROS , 2010
2010
Earlier work this paper cites.
P. Pastor, L. Righetti, M. Kalakrishnan, and S. Schaal, “Online movement adaptation based on previous sensor experiences,” in IROS , 2011
2011
Earlier work this paper cites.
J. Kober and J. Peters, “Policy search for motor primitives in robotics,” Machine Learning , 2011
2011
Earlier work this paper cites.
S. Ross, G. Gordon, and J. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in AISTATS , 2011
2011
Earlier work this paper cites.
B. Akgun, M. Cakmak, J. Yoo, and A. Thomaz, “Trajectories and keyframes for kinesthetic teaching: A human-robot interaction perspective,” in HRI , 2012
2012
Earlier work this paper cites.
F. Stulp, E. Theodorou, and S. Schaal, “Reinforcement learning with sequences of motion primitives for robust manipulation,” IEEE Transactions on Robotics , 2012
2012
Earlier work this paper cites.
J. Kober, J. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” IJRR , 2013
2013
Earlier work this paper cites.
M. Kalakrishnan, P. Pastor, L. Righetti, and S. Schaal, “Learning objective functions for manipulation,” in ICRA , 2013
2013
Earlier work this paper cites.
K. Lee, Y. Su, T. Kim, and Y. Demiris, “A syntactic approach to robot imitation learning using probabilistic activity grammars,” Robotics and Autonomous Systems , 2013
2013
Earlier work this paper cites.
M. Deisenroth, D. Fox, and C. Rasmussen, “Gaussian processes for data-efficient learning in robotics and control,” PAMI , 2014
2014
Earlier work this paper cites.
S. Manschitz, J. Kober, M. Gienger, and J. Peters, “Leraning to sequence movement primitives from demonstrations,” in IROS , 2014
2014
Cited alongside, same era.
C. Daniel, M. Viering, J. Metz, O. Kroemer, and J. Peters, “Active reward learning,” in RSS , 2014
2014
Cited alongside, same era.
H. van Hoof, T. Hermans, G. Neumann, and J. Peters, “Learning robot in-hand manipulation with tactile features,” in Humanoids , 2015
2015
Cited alongside, same era.
W. Han, S. Levine, and P. Abbeel, “Learning compound multi-step controllers under unknown dynamics,” in IROS , 2015
2015
Cited alongside, same era.
Y. Yang, Y. Li, C. Fermuller, and Y. Aloimonos, “Robot learning manipulation action plans by “watching” unconstrained videos from the world wide web,” in AAAI , 2015
2015
Cited alongside, same era.
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, and S. Levine, “QT-Opt: Scalable deep reinforcement learning for vision-based robotic manipulation,” in CoRL , 2018
2018
Later among the works it cites.
T. Lesort, N. Díaz-Rodríguez, J. Goudou, and D. Filliat, “State representation learning for control: An overview,” Neural Networks , 2018
2018
Later among the works it cites.
K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep reinforcement learning in a handful of trials using probabilistic dynamics models,” in NIPS , 2018
2018
Later among the works it cites.
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine, “Variational inverse control with events: A general framework for data-driven reward definition,” in NIPS , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Zhu, T. Park, P. Isola, and A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in ICCV , 2017
2017
Cited alongside, same era.
A. Rusu, M. Vec̆erík, T. Rothörl, N. Heess, R. Pascanu, and R. Hadsell, “Sim-to-real robot learning from pixels with progressive nets,” in CoRL , 2017
2017
Cited alongside, same era.
A. Yahya, A. Li, M. Kalakrishnan, Y. Chebotar, and S. Levine, “Collective robot reinforcement learning with distributed asynchronous guided policy search,” in IROS , 2017
2017
Cited alongside, same era.
C. Schenck and D. Fox, “Visual closed-loop control for pouring liquids,” in ICRA , 2017
2017
Cited alongside, same era.
A. Nair, D. Chen, P. Agrawal, P. Isola, P. Abbeel, J. Malik, and S. Levine, “Combining self-supervised learning and imitation for vision-based rope manipulation,” in ICRA , 2017
2017
Cited alongside, same era.
P. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” in NIPS , 2017
2017
Cited alongside, same era.
K. Ramirez-Amaro, M. Beetz, and G. Cheng, “Transferring skills to humanoid robots by extracting semantic representations from observations of human activities,” Artificial Intelligence , 2017
2017
Cited alongside, same era.
2018
Later among the works it cites.
Y. Liu, A. Gupta, P. Abbeel, and S. Levine, “Imitation from observation: Learning to imitate behaviors from raw video via context translation,” in ICRA , 2018
2018
Later among the works it cites.
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, and S. Levine, “Time-contrastive networks: Self-supervised learning from video,” in ICRA , 2018
2018
Later among the works it cites.
T. Yu, C. Finn, A. Xie, S. Dasari, P. Abbeel, and S. Levine, “One-shot imitation from observing humans via domain-adaptive meta-learning,” in RSS , 2018
2018
Later among the works it cites.
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson, “Learning latent dynamics for planning from pixels,” in ICML , 2018
2018
Later among the works it cites.
D. Pathak, P. Mahmoudieh, G. Luo, P. Agrawal, D. Chen, Y. Shentu, E. Shelhamer, J. Malik, A. Efros, and T. Darrell, “Zero-shot visual imitation,” in ICLR , 2018
2018
Later among the works it cites.
F. Torabi, G. Warnell, and P. Stone, “Behavioral cloning from observation,” in IJCAI , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2019
Closest in time.
P. Sharma, D. Pathak, and A. Gupta, “Third-person visual imitation learning via decoupled hierarchical control,” in NeurIPS , 2019
2019
Closest in time.
A. Singh, L. Yang, K. Hartikainen, C. Finn, and S. Levine, “End-to-end robotic reinforcement learning without reward engineering,” in RSS , 2019
2019
Closest in time.
M. Zhang, S. Vikram, L. Smith, P. Abbeel, M. Johnson, and S. Levine, “SOLAR: Deep structured representations for model-based reinforcement learning,” in ICML , 2019
2019
Closest in time.
2019
Closest in time.
D. Jayaraman, F. Ebert, A. Efros, and S. Levine, “Time-agnostic prediction: Predicting predictable video frames,” in ICLR , 2019
2019
Closest in time.
W. Sun, A. Vemula, B. Boots, and J. Bagnell, “Provably efficient imitation learning from observation alone,” in ICML , 2019
2019
Closest in time.
A. Edwards, H. Sahni, Y. Schroecker, and C. Isbell, “Imitating latent policies from observation,” in ICML , 2019
2019
Closest in time.