Fetching the paper…
Reading the bibliography…
Preference-based Reinforcement Learning (PbRL) methods utilize binary feedback from the human in the loop (HiL) over queried trajectory pairs to learn a reward model in an attempt to approximate the human's underlying reward function capturing their preferences.
Making Smart Homes Smarter: Optimizing Energy Consumption with Human in the Loop
Verma, M.; Bhambri, S.; Gupta, S.; and Buduru, A. B. 2019 · 1912
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Bradley, R. A.; and Terry, M. E. 1952 · 1952
Earlier work this paper cites.
Learning from demonstration
Schaal, S. 1996 · 1996
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y.; Russell, S.; et al. 2000 · 2000
Earlier work this paper cites.
Provable benefit of orthogonal initialization in optimizing deep linear networks
Hu, W.; Xiao, L.; and Pennington, J. 2020 · 2001
Earlier work this paper cites.
Sreedharan, S.; Soni, U.; Verma, M.; Srivastava, S.; and Kambhampati, S. 2020 · 2002
Earlier work this paper cites.
Pattern recognition and machine learning , volume 4
Bishop, C. M.; and Nasrabadi, N. M. 2006 · 2006
Earlier work this paper cites.
Explanation augmented feedback in human-in-the-loop reinforcement learning
Guan, L.; Verma, M.; and Kambhampati, S. 2020 · 2006
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Earlier work this paper cites.
Deep reinforcement learning: A brief survey
Arulkumaran, K.; Deisenroth, M. P.; Brundage, M.; and Bharath, A. A. 2017 · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017 · 2017
Cited alongside, same era.
Inverse reward design
Hadfield-Menell, D.; Milli, S.; Abbeel, P.; Russell, S. J.; and Dragan, A. 2017 · 2017
Cited alongside, same era.
Imitation learning: A survey of learning methods
Hussein, A.; Gaber, M. M.; Elyan, E.; and Jayne, C. 2017 · 2017
Cited alongside, same era.
On weight initialization in deep neural networks
Kumar, S. K. 2017 · 2017
Cited alongside, same era.
Lee, K.; Smith, L.; and Abbeel, P. 2021 · 2021
Later among the works it cites.
Perfect Observability is a Myth: Restraining Bolts in the Real World
Verma, M.; Shah, N.; Nayyar, R. K.; and Hanni, A. 2021 · 2021
Later among the works it cites.
Trust-aware planning: Modeling trust evolution in longitudinal human-robot interaction
Zahedi, Z.; Verma, M.; Sreedharan, S.; and Kambhampati, S. 2021 · 2021
Later among the works it cites.
Symbols as a lingua franca for bridging human-ai chasm for explainable and advisable ai systems
Kambhampati, S.; Sreedharan, S.; Verma, M.; Zha, Y.; and Guan, L. 2022 · 2022
Later among the works it cites.
A review on weight initialization strategies for neural networks
Narkhede, M. V.; Bartakke, P. P.; and Sutaone, M. S. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haarnoja, T.; Zhou, A.; Hartikainen, K.; Tucker, G.; Ha, S.; Tan, J.; Kumar, V.; Zhu, H.; Gupta, A.; Abbeel, P.; et al. 2018 · 2018
Cited alongside, same era.
Human-aligned artificial intelligence is a multiobjective problem
Vamplew, P.; Dazeley, R.; Foale, C.; Firmin, S.; and Mummery, J. 2018 · 2018
Cited alongside, same era.
Specification gaming: The flip side of AI ingenuity— DeepMind
Krakovna, V.; Uesato, J.; Mikulik, V.; et al. 2020 · 2020
Cited alongside, same era.
Benefits of assistance over reward learning
Shah, R.; Freire, P.; Alex, N.; Freedman, R.; Krasheninnikov, D.; Chan, L.; Dennis, M. D.; Abbeel, P.; Dragan, A.; and Russell, S. 2020 · 2020
Cited alongside, same era.
Synthesizing Policies That Account For Human Execution Errors Caused By State Aliasing In Markov Decision Processes
Gopalakrishnan, S.; Verma, M.; and Kambhampati, S. 2021b · 2021
Cited alongside, same era.
Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation
Guan, L.; Verma, M.; Guo, S. S.; Zhang, R.; and Kambhampati, S. 2021 · 2021
Cited alongside, same era.
Gopalakrishnan, S.; Verma, M.; and Kambhampati, S. 2021a
Cited in the paper.
Later among the works it cites.
Park, J.; Seo, Y.; Shin, J.; Lee, H.; Abbeel, P.; and Lee, K. 2022 · 2022
Later among the works it cites.
Soni, U.; Sreedharan, S.; Verma, M.; Guan, L.; Marquez, M.; and Kambhampati, S. 2022 · 2022
Later among the works it cites.
Advice Conformance Verification by Reinforcement Learning agents for Human-in-the-Loop
Verma, M.; Kharkwal, A.; and Kambhampati, S. 2022 · 2022
Later among the works it cites.
Symbol Guided Hindsight Priors for Reward Learning from Human Preferences
Verma, M.; and Metcalf, K. 2022 · 2022
Later among the works it cites.
Modeling the Interplay between Human Trust and Monitoring
Zahedi, Z.; Sreedharan, S.; Verma, M.; and Kambhampati, S. 2022 · 2022
Later among the works it cites.