Fetching the paper…
Reading the bibliography…
Rewards play a crucial role in reinforcement learning.
C. G. Atkeson and S. Schaal, “Robot learning from demonstration,” in ICML , vol. 97. Citeseer, 1997, pp. 12–20
1997
Earlier work this paper cites.
A. Y. Ng, S. J. Russell et al. , “Algorithms for inverse reinforcement learning.” in Icml , vol. 1, 2000, p. 2
2000
Earlier work this paper cites.
D. Shapiro and R. Shachter, “User-agent value alignment,” in Proc. of The 18th Nat. Conf. on Artif. Intell. AAAI , 2002
2002
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Proceedings of the twenty-first international conference on Machine learning , 2004, p. 1
2004
Earlier work this paper cites.
J. A. Bagnell, N. Ratliff, and M. Zinkevich, “Maximum margin planning,” in Proceedings of the International Conference on Machine Learning (ICML) . Citeseer, 2006
2006
Earlier work this paper cites.
D. Ramachandran and E. Amir, “Bayesian inverse reinforcement learning.” in IJCAI , vol. 7, 2007, pp. 2586–2591
2007
Earlier work this paper cites.
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning.” in AAAI , vol. 8. Chicago, IL, USA, 2008, pp. 1433–1438
2008
Earlier work this paper cites.
B. D. Argall, S. Chernova, M. Veloso, and B. Browning, “A survey of robot learning from demonstration,” Robotics and autonomous systems , vol. 57, no. 5, pp. 469–483, 2009
2009
Earlier work this paper cites.
S. Singh, R. L. Lewis, and A. G. Barto, “Where do rewards come from,” in Proceedings of the annual conference of the cognitive science society . Cognitive Science Society, 2009, pp. 2601–2606
2009
Earlier work this paper cites.
B. Settles, “Active learning literature survey,” University of Wisconsin-Madison Department of Computer Sciences, Tech. Rep., 2009
2009
Earlier work this paper cites.
M. Kalakrishnan, L. Righetti, P. Pastor, and S. Schaal, “Learning force control policies for compliant manipulation,” in 2011 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 2011, pp. 4639–4644
2011
Earlier work this paper cites.
M. Kalakrishnan, P. Pastor, L. Righetti, and S. Schaal, “Learning objective functions for manipulation,” in 2013 IEEE International Conference on Robotics and Automation . IEEE, 2013, pp. 1331–1336
2013
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in neural information processing systems , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
M. L. Puterman, Markov decision processes: discrete stochastic dynamic programming . John Wiley & Sons, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
S. Chernova and A. L. Thomaz, “Robot learning from human teachers,” Synthesis Lectures on Artificial Intelligence and Machine Learning , vol. 8, no. 3, pp. 1–121, 2014
2014
Earlier work this paper cites.
K. Muelling, A. Boularias, B. Mohler, B. Schölkopf, and J. Peters, “Learning strategies in table tennis using inverse reinforcement learning,” Biological cybernetics , vol. 108, no. 5, pp. 603–619, 2014
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Cited alongside, same era.
M. Kuderer, S. Gulati, and W. Burgard, “Learning driving styles for autonomous vehicles from demonstration,” in 2015 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2015, pp. 2641–2646
2015
Cited alongside, same era.
D. Kappler, P. Pastor, M. Kalakrishnan, M. Wüthrich, and S. Schaal, “Data-driven online decision making for autonomous manipulation.” in Robotics: Science and Systems , 2015
2015
Cited alongside, same era.
C. Finn, S. Levine, and P. Abbeel, “Guided cost learning: Deep inverse optimal control via policy optimization,” in International conference on machine learning , 2016, pp. 49–58
2016
Cited alongside, same era.
2018
Later among the works it cites.
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine, “Variational inverse control with events: A general framework for data-driven reward definition,” in Advances in Neural Information Processing Systems , 2018, pp. 8538–8547
2018
Later among the works it cites.
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain, “Time-contrastive networks: Self-supervised learning from video,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 1134–1141
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,” in Advances in neural information processing systems , 2016, pp. 2234–2242
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
J. Ho and S. Ermon, “Generative adversarial imitation learning,” in Advances in neural information processing systems , 2016, pp. 4565–4573
2016
Cited alongside, same era.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” The Journal of Machine Learning Research , vol. 17, no. 1, pp. 1334–1373, 2016
2016
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
D. Hadfield-Menell, S. Milli, P. Abbeel, S. J. Russell, and A. Dragan, “Inverse reward design,” in Advances in neural information processing systems , 2017, pp. 6765–6774
2017
Cited alongside, same era.
2018
Later among the works it cites.
2019
Later among the works it cites.
R. Martín-Martín, M. A. Lee, R. Gardner, S. Savarese, J. Bohg, and A. Garg, “Variable impedance control in end-effector space: An action space for reinforcement learning in contact-rich tasks,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2019, pp. 1010–1017
2019
Later among the works it cites.
M. A. Lee, Y. Zhu, K. Srinivasan, P. Shah, S. Savarese, L. Fei-Fei, A. Garg, and J. Bohg, “Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 8943–8950
2019
Later among the works it cites.
D. R. Scobee and S. S. Sastry, “Maximum likelihood constraint inference for inverse reinforcement learning,” in International Conference on Learning Representations , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Cabi, S. Gómez Colmenarejo, A. Novikov, K. Konyushkova, S. Reed, R. Jeong, K. Zolna, Y. Aytar, D. Budden, M. Vecerik et al. , “Scaling data-driven robotics with reward sketching and batch reinforcement learning,” arXiv , pp. arXiv–1909, 2019
2019
Later among the works it cites.
S. Tian, F. Ebert, D. Jayaraman, M. Mudigonda, C. Finn, R. Calandra, and S. Levine, “Manipulation by feel: Touch-based control with deep predictive models,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 818–824
2019
Later among the works it cites.
N. Fazeli, M. Oller, J. Wu, Z. Wu, J. B. Tenenbaum, and A. Rodriguez, “See, feel, act: Hierarchical learning for complex manipulation skills with multisensory fusion,” Science Robotics , vol. 4, no. 26, 2019
2019
Later among the works it cites.
K. Kimble, K. Van Wyk, J. Falco, E. Messina, Y. Sun, M. Shibata, W. Uemura, and Y. Yokokohji, “Benchmarking protocols for evaluating small parts robotic assembly systems,” IEEE Robotics and Automation Letters , vol. 5, no. 2, pp. 883–889, 2020
2020
Closest in time.
O. Mees, M. Merklinger, G. Kalweit, and W. Burgard, “Adversarial skill networks: Unsupervised robot skill learning from video,” 2020
2020
Closest in time.
K. Zakka, A. Zeng, J. Lee, and S. Song, “Form2fit: Learning shape priors for generalizable assembly from disassembly,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 9404–9410
2020
Closest in time.
Z. Wu, L. Sun, W. Zhan, C. Yang, and M. Tomizuka, “Efficient sampling-based maximum entropy inverse reinforcement learning with application to autonomous driving,” IEEE Robotics and Automation Letters , vol. 5, no. 4, pp. 5355–5362, 2020
2020
Closest in time.