Fetching the paper…
Reading the bibliography…
Reward functions are a common way to specify the objective of a robot.
arXiv preprint arXiv:1906.07975
Bıyık E, Wang K, Anari N and Sadigh D (2019) Batch active learning using determinantal point processes · 1906
Earlier work this paper cites.
Management science 23(11): 1224–1233
Krishnan K (1977) Incorporating thresholds of indifference in probabilistic choice models · 1977
Earlier work this paper cites.
MIT press
Ben-Akiva ME, Lerman SR and Lerman SR (1985) Discrete choice analysis: theory and application to travel demand , volume 9 · 1985
Earlier work this paper cites.
In: 2011 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, pp. 1986–1993
Cakmak M, Srinivasa SS, Lee MK, Forlizzi J and Kiesler S (2011) Human preferences for robot-human hand-over configurations · 1993
Earlier work this paper cites.
In: Icml , volume 1. p. 2
Ng AY, Russell SJ et al. (2000) Algorithms for inverse reinforcement learning · 2000
Earlier work this paper cites.
arXiv preprint arXiv:2003.02232
Shah A and Shah J (2020) Interactive robot training for non-markov tasks · 2003
Earlier work this paper cites.
In: Proceedings of the twenty-first international conference on Machine learning . ACM, p. 1
Abbeel P and Ng AY (2004) Apprenticeship learning via inverse reinforcement learning · 2004
Earlier work this paper cites.
In: Proceedings of the 22nd international conference on Machine learning . ACM, pp. 1–8
Abbeel P and Ng AY (2005) Exploration and apprenticeship learning in reinforcement learning · 2005
Earlier work this paper cites.
Journal of machine learning research 6(Jul): 1019–1041
Chu W and Ghahramani Z (2005) Gaussian processes for ordinal regression · 2005
Earlier work this paper cites.
Nature 441(7095): 876
Daw ND, O’doherty JP, Dayan P, Seymour B and Dolan RJ (2006) Cortical substrates for exploratory decisions in humans · 2006
Earlier work this paper cites.
In: IJCAI , volume 7. pp. 2586–2591
Ramachandran D and Amir E (2007) Bayesian inverse reinforcement learning · 2007
Earlier work this paper cites.
In: Aaai , volume 8. Chicago, IL, USA, pp. 1433–1438
Ziebart BD, Maas AL, Bagnell JA and Dey AK (2008) Maximum entropy inverse reinforcement learning · 2008
Earlier work this paper cites.
In: Advances in neural information processing systems . pp. 985–992
Lucas CG, Griffiths TL, Xu F and Fawcett C (2009) A rational model of preference learning and choice prediction by children · 2009
Earlier work this paper cites.
In: Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics . pp. 289–296
Guo S and Sanner S (2010) Real-time multiattribute bayesian preference elicitation with pairwise comparison queries · 2010
Earlier work this paper cites.
In: Advances in neural information processing systems . pp. 2352–2360
Viappiani P and Boutilier C (2010) Optimal bayesian recommendation sets and myopically optimal choice query sets · 2010
Earlier work this paper cites.
Journal of Machine Learning Research 13(Jan): 137–164
Ailon N (2012) An active learning algorithm for ranking from pairwise preferences with an almost optimal query complexity · 2012
Earlier work this paper cites.
International Journal of Social Robotics 4(4): 343–355
Akgun B, Cakmak M, Jiang K and Thomaz AL (2012) Keyframe-based learning from demonstration · 2012
Earlier work this paper cites.
In: Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, pp. 116–131
Akrour R, Schoenauer M and Sebag M (2012) April: Active preference learning-based reinforcement learning · 2012
Earlier work this paper cites.
John Wiley & Sons
Cover TM and Thomas JA (2012) Elements of information theory · 2012
Earlier work this paper cites.
MIT Press, July
Dragan AD and Srinivasa SS (2012) Formalizing assistive teleoperation · 2012
Cited alongside, same era.
Courier Corporation
Luce RD (2012) Individual choice behavior: A theoretical analysis · 2012
Cited alongside, same era.
In: Joint European conference on machine learning and knowledge discovery in databases . Springer, pp. 148–163
Michini B and How JP (2012) Bayesian nonparametric inverse reinforcement learning · 2012
Cited alongside, same era.
In: 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, pp. 5026–5033
Todorov E, Erez T and Tassa Y (2012) Mujoco: A physics engine for model-based control · 2012
Cited alongside, same era.
In: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . ACM, pp. 3075–3084
Kulesza T, Amershi S, Caruana R, Fisher D and Charles D (2014) Structured labeling for facilitating concept evolution in machine learning · 2014
Cited alongside, same era.
In: Conference on Robot Learning (CoRL)
Biyik E and Sadigh D (2018) Batch active preference-based learning of reward functions · 2018
Later among the works it cites.
In: Conference on Robot Learning . pp. 796–805
Bobu A, Bajcsy A, Fisac JF and Dragan AD (2018) Learning under misspecified objective spaces · 2018
Later among the works it cites.
In: Advances in neural information processing systems . pp. 8011–8023
Ibarz B, Leike J, Pohlen T, Irving G, Legg S and Amodei D (2018) Reward learning from human preferences and demonstrations in atari · 2018
Later among the works it cites.
In: Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
Basu C, Biyik E, He Z, Singhal M and Sadigh D (2019) Active learning of reward dynamics from hierarchical queries · 2019
Later among the works it cites.
In: International Conference on Machine Learning . pp. 783–792
Brown D, Goo W, Nagarajan P and Niekum S (2019) Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Javdani S, Srinivasa SS and Bagnell JA (2015) Shared autonomy via hindsight optimization · 2015
Cited alongside, same era.
Presence: Teleoperators and Virtual Environments 24(2): 132–154
Khurshid RP and Kuchenbecker KJ (2015) Data-driven motion mappings improve transparency in teleoperation · 2015
Cited alongside, same era.
Journal of Aerospace Information Systems 12(10): 634–645
Lepird JR, Owen MP and Kochenderfer MJ (2015) Bayesian preference elicitation for multiobjective engineering design optimization · 2015
Cited alongside, same era.
In: 2015 10th ACM/IEEE International Conference on Human-Robot Interaction (HRI) . IEEE, pp. 189–196
Nikolaidis S, Ramakrishnan R, Gu K and Shah J (2015) Efficient model learning from joint-action demonstrations for human-robot collaborative tasks · 2015
Cited alongside, same era.
arXiv preprint arXiv:1606.01540
Brockman G, Cheung V, Pettersson L, Schneider J, Schulman J, Tang J and Zaremba W (2016) Openai gym · 2016
Cited alongside, same era.
In: RSS Workshop on Model Learning for Human-Robot Communication
Holladay R, Javdani S, Dragan A and Srinivasa S (2016) Active comparison based learning incorporating user uncertainty and noise · 2016
Cited alongside, same era.
In: Robotics: Science and Systems , volume 2. Ann Arbor, MI, USA
Sadigh D, Sastry S, Seshia SA and Dragan AD (2016) Planning for autonomous cars that leverage effects on human actions · 2016
Cited alongside, same era.
In: Workshop on Safety and Robustness in Decision Making at the 33rd Conference on Neural Information Processing Systems (NeurIPS) 2019
Brown DS and Niekum S (2019) Deep bayesian reward learning from preferences · 2019
Later among the works it cites.
In: 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI) . IEEE, pp. 317–325
Choudhury R, Swamy G, Hadfield-Menell D and Dragan AD (2019) On the utility of model learning in hri · 2019
Later among the works it cites.
In: 2019 IEEE/AIAA 38th Digital Avionics Systems Conference (DASC)
Katz SM, Bihan ACL and Kochenderfer MJ (2019) Learning an urban air mobility encounter model from expert preferences · 2019
Later among the works it cites.
In: Proceedings of Robotics: Science and Systems (RSS)
Palan M, Shevchuk G, Landolfi NC and Sadigh D (2019) Learning reward functions by integrating human demonstrations and preferences · 2019
Later among the works it cites.
IEEE Robotics and Automation Letters 4(2): 1691–1698
Wilde N, Kulić D and Smith SL (2019) Bayesian active learning for collaborative task specification using equivalence regions · 2019
Later among the works it cites.
In: Proceedings of Robotics: Science and Systems (RSS)
Biyik E, Huynh N, Kochenderfer MJ and Sadigh D (2020) Active preference-based gaussian process regression for reward learning · 2020
Closest in time.
In: Conference on robot learning . PMLR, pp. 330–359
Brown DS, Goo W and Niekum S (2020) Better-than-demonstrator imitation learning via automatically-ranked demonstrations · 2020
Closest in time.
In: Conference on robot learning . PMLR
Chen L, Paleja R and Gombolay M (2020) Learning from suboptimal demonstration via self-supervised reward regression · 2020
Closest in time.
In: ACM/IEEE International Conference on Human-Robot Interaction (HRI)
Kwon M, Biyik E, Talati A, Bhasin K, Losey DP and Sadigh D (2020) When humans aren’t optimal: Robots that collaborate with risk-aware humans · 2020
Closest in time.
In: Proceedings of the Conference on Robot Learning , Proceedings of Machine Learning Research , volume 100. PMLR, pp. 1005–1014
Park D, Noseworthy M, Paul R, Roy S and Roy N (2020) Inferring task goals and constraints using bayesian nonparametric inverse reinforcement learning · 2020
Closest in time.
In: 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE
Tucker M, Novoseller E, Kann C, Sui Y, Yue Y, Burdick J and Ames AD (2020) Preference-based learning for exoskeleton gait optimization · 2020
Closest in time.
arXiv preprint arXiv:2107.01995
Habibian S, Jonnavittula A and Losey DP (2021) Here’s what I’ve learned: Asking questions that reveal reward learning · 2021
Closest in time.
arXiv preprint arXiv:2103.02727
Katz S, Maleki A, Biyik E and Kochenderfer MJ (2021) Preference-based learning of reward function features · 2021
Closest in time.