Fetching the paper…
Reading the bibliography…
When robots learn reward functions using high capacity models that take raw state directly as input, they need to both learn a representation for what matters in the task -- the task ``features" -- as well as how to combine these features into a single objective.
Rank correlation methods
Maurice George Kendall. 1948 · 1948
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Ralph Allan Bradley and Milton E Terry. 1952 · 1952
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning. In Machine Learning (ICML), International Conference on . ACM
Pieter Abbeel and Andrew Y Ng. 2004 · 2004
Earlier work this paper cites.
Absolute identification by relative judgment
Neil Stewart, Gordon DA Brown, and Nick Chater. 2005 · 2005
Earlier work this paper cites.
Generalized non-metric multidimensional scaling. In Artificial Intelligence and Statistics . PMLR, 11–18
Sameer Agarwal, Josh Wills, Lawrence Cayton, Gert Lanckriet, David Kriegman, and Serge Belongie. 2007 · 2007
Earlier work this paper cites.
Apprenticeship learning about multiple intentions. In ICML
Monica Babes, Vukosi N Marivate, Kaushik Subramanian, and Michael L Littman. 2011 · 2011
Earlier work this paper cites.
Bayesian multitask inverse reinforcement learning. In European workshop on reinforcement learning . Springer, 273–284
Christos Dimitrakakis and Constantin A Rothkopf. 2011 · 2011
Earlier work this paper cites.
Learning Multi-modal Similarity
Brian McFee, Gert Lanckriet, and Tony Jebara. 2011 · 2011
Earlier work this paper cites.
Human-robot proxemics: physical and psychological distancing in human-robot interaction. In Proceedings of the 6th international conference on Human-robot interaction . 331–338
Jonathan Mumm and Bilge Mutlu. 2011 · 2011
Earlier work this paper cites.
Adaptively learning the crowd kernel
Omer Tamuz, Ce Liu, Serge Belongie, Ohad Shamir, and Adam Tauman Kalai. 2011 · 2011
Earlier work this paper cites.
Designing robot learners that ask good questions. In International Conference on Human-Robot Interaction, HRI’12, Boston, MA, USA - March 05 - 08, 2012 , Holly A. Yanco, Aaron Steinfeld, Vanessa Evers, and Odest Chadwicke Jenkins (Eds.). ACM, 17–24
Maya Cakmak and Andrea Lockerd Thomaz. 2012 · 2012
Earlier work this paper cites.
Nonparametric Bayesian inverse reinforcement learning for multiple reward functions
Jaedeug Choi and Kee-Eung Kim. 2012 · 2012
Earlier work this paper cites.
Learning Feature Representations with K-Means. In Neural Networks: Tricks of the Trade
Adam Coates and A. Ng. 2012 · 2012
Earlier work this paper cites.
Active Learning for Teaching a Robot Grounded Relational Symbols. In IJCAI 2013, Proceedings of the 23rd International Joint Conference on Artificial Intelligence, Beijing, China, August 3-9, 2013 , Francesca Rossi (Ed.). IJCAI/AAAI, 1451–1457
Johannes Kulick, Marc Toussaint, Tobias Lang, and Manuel Lopes. 2013 · 2013
Earlier work this paper cites.
Learning Perceptual Kernels for Visualization Design
Cagatay Demiralp, Michael Bernstein, and Jeffrey Heer. 2014 · 2014
Earlier work this paper cites.
Discovering task constraints through observation and active learning. In 2014 IEEE/RSJ International Conference on Intelligent Robots and Systems, Chicago, IL, USA, September 14-18, 2014 . IEEE, 4442–4449
Bradley Hayes and Brian Scassellati. 2014 · 2014
Earlier work this paper cites.
Diederik P. Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
A Kernel-Learning Approach to Semi-supervised Clustering with Relative Distance Comparisons, Vol. 9284
Ehsan Amid, Aristides Gionis, and Antti Ukkonen. 2015 · 2015
Earlier work this paper cites.
Unsupervised Visual Representation Learning by Context Prediction
Carl Doersch, Abhinav Kumar Gupta, and Alexei A. Efros. 2015 · 2015
Earlier work this paper cites.
Movement primitives via optimization. In 2015 IEEE International Conference on Robotics and Automation (ICRA) . 2339–2346
A. D. Dragan, K. Muelling, J. Andrew Bagnell, and S. S. Srinivasa. 2015 · 2015
Cited alongside, same era.
Learning local feature descriptors with triplets and shallow convolutional neural networks. 119.1–119.11
Vassileios Balntas, Edgar Riba, Daniel Ponsa, and Krystian Mikolajczyk. 2016 · 2016
Cited alongside, same era.
InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets. In Proceedings of the 30th International Conference on Neural Information Processing Systems (Barcelona, Spain) (NIPS’16) . Curran Associates Inc., Red Hook, NY, USA, 2180–2188
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel. 2016 · 2016
Cited alongside, same era.
Learning Robot Objectives from Physical Human Interaction. In Proceedings of the 1st Annual Conference on Robot Learning (Proceedings of Machine Learning Research, Vol. 78) , Sergey Levine, Vincent Vanhoucke, and Ken Goldberg (Eds.). PMLR, 217–226
Andrea Bajcsy, Dylan P. Losey, Marcia K. O’Malley, and Anca D. Dragan. 2017 · 2017
Contrastive Predictive Coding Based Feature for Automatic Speaker Verification
Cheng-I Lai. 2019 · 2019
Later among the works it cites.
An Incremental Feature Set Refinement in a Programming by Demonstration Scenario. In 4th IEEE International Conference on Advanced Robotics and Mechatronics, ICARM 2019, Toyonaka, Japan, July 3-5, 2019 . IEEE, 372–377
Hoai Luu-Duc and Jun Miura. 2019 · 2019
Later among the works it cites.
Smile: Scalable meta inverse reinforcement learning through context-conditional policies
Seyed Kamyar Seyed Ghasemipour, Shixiang Shane Gu, and Richard Zemel. 2019 · 2019
Later among the works it cites.
Meta-inverse reinforcement learning with probabilistic context variables
Lantao Yu, Tianhe Yu, Chelsea Finn, and Stefano Ermon. 2019 · 2019
Later among the works it cites.
Safe Imitation Learning via Fast Bayesian Reward Inference from Preferences. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119) , Hal Daumé III and Aarti Singh (Eds.). PMLR, 1165–1177
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep Reinforcement Learning from Human Preferences. In Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Cited alongside, same era.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. In Proceedings of the 34th International Conference on Machine Learning - Volume 70 (Sydney, NSW, Australia) (ICML’17) . JMLR.org, 1126–1135
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework. In ICLR
Irina Higgins, Loïc Matthey, Arka Pal, Christopher P. Burgess, Xavier Glorot, Matthew M. Botvinick, Shakir Mohamed, and Alexander Lerchner. 2017 · 2017
Cited alongside, same era.
Active preference-based learning of reward functions. In Robotics: Science and systems
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia. 2017 · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, Johannes Fürnkranz, et al · 2017
Cited alongside, same era.
Playing Hard Exploration Games by Watching YouTube. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (Montréal, Canada) (NIPS’18) . Curran Associates Inc., Red Hook, NY, USA, 2935–2945
Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, and Nando de Freitas. 2018 · 2018
Cited alongside, same era.
Batch active preference-based learning of reward functions. In Conference on robot learning . PMLR, 519–528
Erdem Biyik and Dorsa Sadigh. 2018 · 2018
Cited alongside, same era.
Human-Driven Feature Selection for a Robotic Agent Learning Classification Tasks from Demonstration. In 2018 IEEE International Conference on Robotics and Automation, ICRA 2018, Brisbane, Australia, May 21-25, 2018 . IEEE, 6923–6930
Kalesha Bullard, Sonia Chernova, and Andrea Lockerd Thomaz. 2018 · 2018
Cited alongside, same era.
Daniel Brown, Russell Coleman, Ravi Srinivasan, and Scott Niekum. 2020 · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations. In International conference on machine learning . PMLR, 1597–1607
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020 · 2020
Later among the works it cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca Dragan. 2020 · 2020
Later among the works it cites.
CURL: Contrastive Unsupervised Representations for Reinforcement Learning. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119) , Hal Daumé III and Aarti Singh (Eds.). PMLR, 5639–5650
Michael Laskin, Aravind Srinivas, and Pieter Abbeel. 2020 · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. 2020 · 2020
Later among the works it cites.
Fine-grained driving behavior prediction via context-aware multi-task inverse reinforcement learning. In 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2281–2287
Kentaro Nishi and Masamichi Shimosaka. 2020 · 2020
Later among the works it cites.
Feature Expansive Reward Learning: Rethinking Human Input. In Proceedings of the 2021 ACM/IEEE International Conference on Human-Robot Interaction (Boulder, CO, USA) (HRI ’21) . Association for Computing Machinery, New York, NY, USA, 216–224
Andreea Bobu, Marius Wiggert, Claire Tomlin, and Anca D. Dragan. 2021 · 2021
Later among the works it cites.
Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos
Annie S. Chen, Suraj Nair, and Chelsea Finn. 2021 · 2021
Later among the works it cites.
Meta Preference Learning for Fast User Adaptation in Human-Supervisory Multi-Robot Deployments. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 5851–5856
Chao Huang, Wenhao Luo, and Rui Liu. 2021 · 2021
Later among the works it cites.
Roial: Region of interest active learning for characterizing exoskeleton gait preference landscapes. In 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 3212–3218
Kejun Li, Maegan Tucker, Erdem Bıyık, Ellen Novoseller, Joel W Burdick, Yanan Sui, Dorsa Sadigh, Yisong Yue, and Aaron D Ames. 2021 · 2021
Later among the works it cites.
Learning Perceptual Concepts by Bootstrapping From Human Queries
Andreea Bobu, Chris Paxton, Wei Yang, Balakumar Sundaralingam, Yu-Wei Chao, Maya Cakmak, and Dieter Fox. 2022 · 2022
Later among the works it cites.
On the Effectiveness of Fine-tuning Versus Meta-reinforcement Learning
Zhao Mandi, Pieter Abbeel, and Stephen James. 2022 · 2022
Later among the works it cites.
Self-Supervised Pretraining Improves Self-Supervised Pretraining. In 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) . IEEE Computer Society, Los Alamitos, CA, USA, 1050–1060
C. J. Reed, X. Yue, A. Nrusimha, S. Ebrahimi, V. Vijaykumar, R. Mao, B. Li, S. Zhang, D. Guillory, S. Metzger, K. Keutzer, and T. Darrell. 2022 · 2022
Later among the works it cites.
Teaching Robots to Span the Space of Functional Expressive Motion
Arjun Sripathy, Andreea Bobu, Zhongyu Li, Koushil Sreenath, Daniel S. Brown, and Anca D. Dragan. 2022 · 2022
Later among the works it cites.