Fetching the paper…
Reading the bibliography…
The advent of large language models (LLMs) has sparked significant interest in using natural language for preference learning.
Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
Bradley, R. A.; and Terry, M. E. 1952 · 1952
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y.; and Russell, S. J. 2000 · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P.; and Ng, A. Y. 2004 · 2004
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Ramachandran, D.; and Amir, E. 2007 · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D.; Maas, A. L.; Bagnell, J. A.; and Dey, A. K. 2008 · 2008
Earlier work this paper cites.
Learning behavior styles with inverse reinforcement learning
Lee, S. J.; and Popović, Z. 2010 · 2010
Earlier work this paper cites.
Building strong semi-autonomous systems
Zilberstein, S. 2015 · 2015
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Gal, Y.; and Ghahramani, Z. 2016 · 2016
Earlier work this paper cites.
Steps toward robust artificial intelligence
Dietterich, T. G. 2017 · 2017
Earlier work this paper cites.
Planet dump retrieved from https://planet.osm.org
OpenStreetMap Contributors. 2017 · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Sadigh, D.; Dragan, A. D.; Sastry, S. S.; and Seshia, S. A. 2017 · 2017
Earlier work this paper cites.
Robust imitation of diverse behaviors
Wang, Z.; Merel, J. S.; Reed, S. E.; de Freitas, N.; Wayne, G.; and Heess, N. 2017 · 2017
Earlier work this paper cites.
Learning from richer human guidance: Augmenting comparison-based learning with feature queries
Basu, C.; Singhal, M.; and Dragan, A. D. 2018 · 2018
Earlier work this paper cites.
SFV: Reinforcement learning of physical skills from videos
Peng, X. B.; Kanazawa, A.; Malik, J.; Abbeel, P.; and Levine, S. 2018 · 2018
Earlier work this paper cites.
Cost functions for robot motion style
Zhou, A.; and Dragan, A. D. 2018 · 2018
Earlier work this paper cites.
Asking easy questions: A user-friendly approach to active reward learning
Biyik, E.; Palan, M.; Landolfi, N. C.; Losey, D. P.; and Sadigh, D. 2019 · 2019
Earlier work this paper cites.
Better-than-demonstrator imitation learning via automatically-ranked demonstrations
Brown, D. S.; Goo, W.; and Niekum, S. 2019 · 2019
Earlier work this paper cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Brown, D. S.; Goo, W.; Prabhat, N.; and Niekum, S. 2019 · 2019
Cited alongside, same era.
SDRL: interpretable and data-efficient deep reinforcement learning leveraging symbolic planning
Lyu, D.; Yang, F.; Liu, B.; and Gustafson, S. 2019 · 2019
Cited alongside, same era.
Safe imitation learning via fast Bayesian reward inference from preferences
Brown, D. S.; Niekum, S.; Coleman, R.; and Srinivasan, R. 2020 · 2020
Cited alongside, same era.
Symbolic plans as high-level instructions for reinforcement learning
Illanes, L.; Yan, X.; Icarte, R. T.; and McIlraith, S. A. 2020 · 2020
Cited alongside, same era.
CARL: Controllable agent with reinforcement learning for quadruped locomotion
Luo, Y.-S.; Soeseno, J. H.; Chen, T. P.-C.; and Chen, W.-C. 2020 · 2020
Cited alongside, same era.
Soni, U.; Thakur, N.; Sreedharan, S.; Guan, L.; Verma, M.; Marquez, M.; and Kambhampati, S. 2022 · 2022
Later among the works it cites.
Tevet, G.; Raab, S.; Gordon, B.; Shafir, Y.; Cohen-Or, D.; and Bermano, A. H. 2022 · 2022
Later among the works it cites.
A dual representation framework for robot learning with human guidance
Zhang, R.; Bansal, D.; Hao, Y.; Hiranaka, A.; Gao, J.; Wang, C.; Martín-Martín, R.; Fei-Fei, L.; and Wu, J. 2022 · 2022
Later among the works it cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Later among the works it cites.
LATTE: LAnguage Trajectory TransformEr
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bobu, A.; Paxton, C.; Yang, W.; Sundaralingam, B.; Chao, Y.-W.; Cakmak, M.; and Fox, D. 2021 · 2021
Cited alongside, same era.
Value alignment verification
Brown, D. S.; Schneider, J. J.; and Niekum, S. 2021 · 2021
Cited alongside, same era.
Actionable models: Unsupervised offline reinforcement learning of robotic skills
Chebotar, Y.; Hausman, K.; Lu, Y.; Xiao, T.; Kalashnikov, D.; Varley, J.; Irpan, A.; Eysenbach, B.; Julian, R.; Finn, C.; et al. 2021 · 2021
Cited alongside, same era.
Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation
Guan, L.; Verma, M.; Guo, S. S.; Zhang, R.; and Kambhampati, S. 2021 · 2021
Cited alongside, same era.
AMP: Adversarial motion priors for stylized physics-based character control
Peng, X. B.; Ma, Z.; Abbeel, P.; Levine, S.; and Kanazawa, A. 2021 · 2021
Cited alongside, same era.
Learning preferences for interactive autonomy
Biyik, E. 2022 · 2022
Cited alongside, same era.
Leveraging approximate symbolic models for reinforcement learning via skill diversity
Guan, L.; Sreedharan, S.; and Kambhampati, S. 2022 · 2022
Cited alongside, same era.
Bucker, A.; Figueredo, L. F. C.; Haddadin, S.; Kapoor, A.; Ma, S.; Vemprala, S.; and Bonatti, R. 2023 · 2023
Later among the works it cites.
Minigrid & Miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks
Chevalier-Boisvert, M.; Dai, B.; Towers, M.; Perez-Vicente, R.; Willems, L.; Lahlou, S.; Pal, S.; Castro, P. S.; and Terry, J. 2023 · 2023
Later among the works it cites.
“No, to the Right” – Online language corrections for robotic manipulation via shared autonomy
Cui, Y.; Karamcheti, S.; Palleti, R.; Shivakumar, N.; Liang, P.; and Sadigh, D. 2023 · 2023
Later among the works it cites.
Learning to model the world with language
Lin, J.; Du, Y.; Watkins, O.; Hafner, D.; Abbeel, P.; Klein, D.; and Dragan, A. 2023 · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Ma, Y. J.; Liang, W.; Wang, G.; Huang, D.-A.; Bastani, O.; Jayaraman, D.; Zhu, Y.; Fan, L.; and Anandkumar, A. 2023 · 2023
Later among the works it cites.
Explanation-guided reward alignment
Mahmud, S.; Saisubramanian, S.; and Zilberstein, S. 2023 · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
Yu, W.; Gileadi, N.; Fu, C.; Kirmani, S.; Lee, K.-H.; Arenas, M. G.; Chiang, H.-T. L.; Erez, T.; Hasenclever, L.; Humplik, J.; et al. 2023 · 2023
Later among the works it cites.
Lou, X.; Zhang, J.; Wang, Z.; Huang, K.; and Du, Y. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2024 · 2024
Closest in time.
RoboCLIP: One demonstration is enough to learn robot policies
Sontakke, S.; Zhang, J.; Arnold, S.; Pertsch, K.; Biyik, E.; Sadigh, D.; Finn, C.; and Itti, L. 2024 · 2024
Closest in time.
A survey on few-shot class-incremental learning
Tian, S.; Li, L.; Li, W.; Ran, H.; Ning, X.; and Tiwari, P. 2024 · 2024
Closest in time.
Optimizing robot behavior via comparative language feedback
Tien, J.; Yang, Z.; Jun, M.; Russell, S. J.; Dragan, A.; and Biyik, E. 2024 · 2024
Closest in time.
RL-VLM-F: Reinforcement learning from vision language foundation model feedback
Wang, Y.; Sun, Z.; Zhang, J.; Xian, Z.; Biyik, E.; Held, D.; and Erickson, Z. 2024 · 2024
Closest in time.