Fetching the paper…
Reading the bibliography…
Data generation and labeling are often expensive in robot learning.
Individual comparisons by ranking methods
Frank Wilcoxon. 1945 · 1945
Earlier work this paper cites.
Submodularity in data subset selection and active learning. In International Conference on Machine Learning . 1954–1963
Kai Wei, Rishabh Iyer, and Jeff Bilmes. 2015 · 1963
Earlier work this paper cites.
Least squares quantization in PCM
Stuart Lloyd. 1982 · 1982
Earlier work this paper cites.
Clustering by means of medoids
Leonard Kaufman and Peter Rousseeuw. 1987 · 1987
Earlier work this paper cites.
Simulated annealing
Dimitris Bertsimas and John Tsitsiklis. 1993 · 1993
Earlier work this paper cites.
An exact algorithm for maximum entropy sampling
Chun-Wa Ko, Jon Lee, and Maurice Queyranne. 1995 · 1995
Earlier work this paper cites.
Elements of information theory
Thomas M Cover. 1999 · 1999
Earlier work this paper cites.
An adaptive Metropolis algorithm
Heikki Haario, Eero Saksman, Johanna Tamminen, et al · 2001
Earlier work this paper cites.
Eynard–Mehta theorem, Schur process, and their Pfaffian analogs
Alexei Borodin and Eric M Rains. 2005 · 2005
Earlier work this paper cites.
Scalable training of L 1-regularized log-linear models. In Proceedings of the 24th international conference on Machine learning . ACM, 33–40
Galen Andrew and Jianfeng Gao. 2007 · 2007
Earlier work this paper cites.
Discriminative batch mode active learning. In Advances in neural information processing systems . 593–600
Yuhong Guo and Dale Schuurmans. 2008 · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning.. In Aaai , Vol. 8. Chicago, IL, USA, 1433–1438
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey. 2008 · 2008
Earlier work this paper cites.
On selecting a maximum volume sub-matrix of a matrix and related problems
Ali Çivril and Malik Magdon-Ismail. 2009 · 2009
Earlier work this paper cites.
Preference learning in recommender systems
Marco De Gemmis, Leo Iaquinta, Pasquale Lops, Cataldo Musto, Fedelucio Narducci, and Giovanni Semeraro. 2009 · 2009
Earlier work this paper cites.
An axiomatic approach for result diversification. In Proceedings of the 18th international conference on World wide web . ACM, 381–390
Sreenivas Gollapudi and Aneesh Sharma. 2009 · 2009
Earlier work this paper cites.
Maximizing global entropy reduction for active learning in speech recognition. In Acoustics, Speech and Signal Processing, 2009. ICASSP 2009. IEEE International Conference on . IEEE, 4721–4724
Balakrishnan Varadarajan, Dong Yu, Li Deng, and Alex Acero. 2009 · 2009
Earlier work this paper cites.
Computational complexity between K-means and K-medoids clustering algorithms for normal and uniform distributions of data points
T Velmurugan and T Santhanam. 2010 · 2010
Earlier work this paper cites.
Preference-based policy learning. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 12–27
Riad Akrour, Marc Schoenauer, and Michele Sebag. 2011 · 2011
Earlier work this paper cites.
k-DPPs: Fixed-size determinantal point processes. In Proceedings of the 28th International Conference on Machine Learning (ICML-11) . 1193–1200
Alex Kulesza and Ben Taskar. 2011 · 2011
Earlier work this paper cites.
An active learning algorithm for ranking from pairwise preferences with an almost optimal query complexity
Nir Ailon. 2012 · 2012
Earlier work this paper cites.
Keyframe-based learning from demonstration
Baris Akgun, Maya Cakmak, Karl Jiang, and Andrea L Thomaz. 2012 · 2012
Earlier work this paper cites.
April: Active preference learning-based reinforcement learning. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 116–131
Riad Akrour, Marc Schoenauer, and Michèle Sebag. 2012 · 2012
Earlier work this paper cites.
Max-sum diversification, monotone submodular functions and dynamic updates. In Proceedings of the 31st ACM SIGMOD-SIGACT-SIGAI symposium on Principles of Database Systems . ACM, 155–166
Allan Borodin, Hyun Chul Lee, and Yuli Ye. 2012 · 2012
Earlier work this paper cites.
Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
Johannes Fürnkranz, Eyke Hüllermeier, Weiwei Cheng, and Sang-Hyeun Park. 2012 · 2012
Earlier work this paper cites.
Determinantal Point Processes for Machine Learning
Alex Kulesza and Ben Taskar. 2012 · 2012
Earlier work this paper cites.
Continuous inverse optimal control with locally optimal examples. In Proceedings of the 29th International Conference on Machine Learning . 475–482
Sergey Levine and Vladlen Koltun. 2012 · 2012
Earlier work this paper cites.
Preference-learning based inverse reinforcement learning for dialog control. In Thirteenth Annual Conference of the International Speech Communication Association
Hiroaki Sugiyama, Toyomi Meguro, and Yasuhiro Minami. 2012 · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control. In Intelligent Robots and Systems (IROS), 2012 IEEE/RSJ International Conference on . IEEE, 5026–5033
Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012 · 2012
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries. In Advances in neural information processing systems . 1133–1141
Aaron Wilson, Alan Fern, and Prasad Tadepalli. 2012 · 2012
Earlier work this paper cites.
Near-optimal Batch Mode Active Learning and Adaptive Submodular Optimization
Yuxin Chen and Andreas Krause. 2013 · 2013
Earlier work this paper cites.
Exponential inapproximability of selecting a maximum volume sub-matrix
Ali Civril and Malik Magdon-Ismail. 2013 · 2013
Cited alongside, same era.
Active learning for probabilistic hypotheses using the maximum Gibbs error criterion. In Advances in Neural Information Processing Systems . 1457–1465
Nguyen Viet Cuong, Wee Sun Lee, Nan Ye, Kian Ming A Chai, and Hai Leong Chieu. 2013 · 2013
Cited alongside, same era.
Numpy/scipy Recipes for Data Science: k-Medoids Clustering
Christian Bauckhage. 2015 · 2015
Cited alongside, same era.
Learning preferences for manipulation tasks from online coactive feedback
Ashesh Jain, Shikhar Sharma, Thorsten Joachims, and Ashutosh Saxena. 2015 · 2015
Cited alongside, same era.
Randomized rounding for the largest simplex problem. In Proceedings of the forty-seventh annual ACM symposium on Theory of computing . ACM, 861–870
Aleksandar Nikolov. 2015 · 2015
Cited alongside, same era.
Modified log-Sobolev inequalities for strong-Rayleigh measures
Jonathan Hermon and Justin Salez. 2019 · 2019
Later among the works it cites.
Algorithms for optimization
Mykel J Kochenderfer and Tim A Wheeler. 2019 · 2019
Later among the works it cites.
Learning Reward Functions by Integrating Human Demonstrations and Preferences. In Proceedings of Robotics: Science and Systems (RSS)
Malayandi Palan, Nicholas C. Landolfi, Gleb Shevchuk, and Dorsa Sadigh. 2019 · 2019
Later among the works it cites.
Single shot active learning using pseudo annotators
Yazhou Yang and Marco Loog. 2019 · 2019
Later among the works it cites.
Active mini-batch sampling using repulsive point processes. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 5741–5748
Cheng Zhang, Cengiz Öztireli, Stephan Mandt, and Giampiero Salvi. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-class active learning by uncertainty sampling with diversity maximization
Yi Yang, Zhigang Ma, Feiping Nie, Xiaojun Chang, and Alexander G Hauptmann. 2015 · 2015
Cited alongside, same era.
Monte Carlo Markov chain algorithms for sampling strongly Rayleigh distributions and determinantal point processes. In Conference on Learning Theory . 103–115
Nima Anari, Shayan Oveis Gharan, and Alireza Rezaei. 2016 · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
Dissimilarity-based sparse subset selection
Ehsan Elhamifar, Guillermo Sapiro, and S Shankar Sastry. 2016 · 2016
Cited alongside, same era.
Active comparison based learning incorporating user uncertainty and noise. In RSS Workshop on Model Learning for Human-Robot Communication
Rachel Holladay, Shervin Javdani, Anca Dragan, and Siddhartha Srinivasa. 2016 · 2016
Cited alongside, same era.
Active reinforcement learning: Observing rewards at a cost. In Future of Interactive Learning Machines, NIPS Workshop
David Krueger, Jan Leike, Owain Evans, and John Salvatier. 2016 · 2016
Cited alongside, same era.
Fast mixing markov chains for strongly Rayleigh measures, DPPs, and constrained sampling. In Advances in Neural Information Processing Systems . 4188–4196
Chengtao Li, Suvrit Sra, and Stefanie Jegelka. 2016 · 2016
Cited alongside, same era.
Active Preference-Based Gaussian Process Regression for Reward Learning. In Proceedings of Robotics: Science and Systems (RSS)
Erdem Biyik, Nicolas Huynh, Mykel J. Kochenderfer, and Dorsa Sadigh. 2020 · 2020
Later among the works it cites.
Better-than-demonstrator imitation learning via automatically-ranked demonstrations. In Conference on Robot Learning . PMLR, 330–359
Daniel S Brown, Wonjoon Goo, and Scott Niekum. 2020 · 2020
Later among the works it cites.
When Humans Aren’t Optimal: Robots that Collaborate with Risk-Aware Humans. In ACM/IEEE International Conference on Human-Robot Interaction (HRI)
Minae Kwon, Erdem Biyik, Aditi Talati, Karan Bhasin, Dylan P. Losey, and Dorsa Sadigh. 2020 · 2020
Later among the works it cites.
ROIAL: Region of Interest Active Learning for Characterizing Exoskeleton Gait Preference Landscapes
Kejun Li, Maegan Tucker, Erdem Biyik, Ellen Novoseller, Joel W. Burdick, Yanan Sui, Dorsa Sadigh, Yisong Yue, and Aaron D. Ames. 2020 · 2020
Later among the works it cites.
Controlling Assistive Robots with Learned Latent Actions. In International Conference on Robotics and Automation (ICRA)
Dylan P. Losey, Krishnan Srinivasan, Ajay Mandlekar, Animesh Garg, and Dorsa Sadigh. 2020 · 2020
Later among the works it cites.
Interactive robot training for non-markov tasks
Ankit Shah, Samir Wadhwania, and Julie Shah. 2020 · 2020
Later among the works it cites.
Preference-based learning for exoskeleton gait optimization. In 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2351–2357
Maegan Tucker, Ellen Novoseller, Claudia Kann, Yanan Sui, Yisong Yue, Joel W Burdick, and Aaron D Ames. 2020 · 2020
Later among the works it cites.
Active Preference Learning using Maximum Regret. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
Nils Wilde, Dana Kulic, and Stephen L Smith. 2020 · 2020
Later among the works it cites.
Preference-based learning of reward function features
Sydney M Katz, Amir Maleki, Erdem Bıyık, and Mykel J Kochenderfer. 2021 · 2021
Later among the works it cites.
B-Pref: Benchmarking Preference-Based Reinforcement Learning. In Neural Information Processing Systems (NeurIPS)
Kimin Lee, Laura Smith, Anca Dragan, and Pieter Abbeel. 2021 · 2021
Later among the works it cites.
Learning Reward Functions from Scale Feedback. In Proceedings of the 5th Conference on Robot Learning (CoRL)
Nils Wilde and Erdem Biyik. 2021 · 2021
Later among the works it cites.
Learning reward functions from diverse sources of human feedback: Optimally integrating demonstrations and preferences
Erdem Bıyık, Dylan P Losey, Malayandi Palan, Nicholas C Landolfi, Gleb Shevchuk, and Dorsa Sadigh. 2022a · 2022
Later among the works it cites.
APReL: A Library for Active Preference-based Reward Learning Algorithms. In Proceedings of the 2022 ACM/IEEE International Conference on Human-Robot Interaction . 613–617
Erdem Bıyık, Aditi Talati, and Dorsa Sadigh. 2022b · 2022
Later among the works it cites.
Learning multimodal rewards from rankings. In Conference on Robot Learning . PMLR, 342–352
Vivek Myers, Erdem Biyik, Nima Anari, and Dorsa Sadigh. 2022 · 2022
Later among the works it cites.
Do we use the Right Measure? Challenges in Evaluating Reward Learning Algorithms. In 6th Annual Conference on Robot Learning
Nils Wilde and Javier Alonso-Mora. 2022 · 2022
Later among the works it cites.
Active Preference-Based Gaussian Process Regression for Reward Learning and Optimization
Erdem Biyik, Nicolas Huynh, Mykel J. Kochenderfer, and Dorsa Sadigh. 2023a · 2023
Later among the works it cites.
Preference Elicitation with Soft Attributes in Interactive Recommendation
Erdem Biyik, Fan Yao, Yinlam Chow, Alex Haig, Chih-wei Hsu, Mohammad Ghavamzadeh, and Craig Boutilier. 2023b · 2023
Later among the works it cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al · 2023
Later among the works it cites.
Efficient preference-based reinforcement learning using learned dynamics models
Yi Liu, Gaurav Datta, Ellen Novoseller, and Daniel S Brown. 2023 · 2023
Later among the works it cites.
Active Reward Learning from Online Preferences. In International Conference on Robotics and Automation (ICRA)
Vivek Myers, Erdem Biyik, and Dorsa Sadigh. 2023 · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023 · 2023
Later among the works it cites.
A Generalized Acquisition Function for Preference-based Reward Learning. In International Conference on Robotics and Automation (ICRA)
Evan Ellis, Gaurav R. Ghosal, Stuart J. Russell, Anca Dragan, and Erdem Biyik. 2024 · 2024
Closest in time.
Contrastive Preference Learning: Learning From Human Feedback without RL. In International Conference on Learning Representations (ICLR)
Joey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn, Scott Niekum, W. Bradley Knox, and Dorsa Sadigh. 2024 · 2024
Closest in time.
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
Yufei Wang, Zhanyi Sun, Jesse Zhang, Zhou Xian, Erdem Biyik, David Held, and Zackory Erickson. 2024 · 2024
Closest in time.