Fetching the paper…
Reading the bibliography…
While reinforcement learning (RL) has become a more popular approach for robotics, designing sufficiently informative reward functions for complex tasks has proven to be extremely difficult due their inability to capture human intent and policy exploitation.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 1910
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
The magical number seven, plus or minus two: Some limits on our capacity for processing information
G. A. Miller · 1956
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Bayesian inverse reinforcement learning
D. Ramachandran and E. Amir · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey, et al · 2008
Earlier work this paper cites.
Tamer: Training an agent manually via evaluative reinforcement
W. B. Knox and P. Stone · 2008
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
W. B. Knox and P. Stone · 2009
Earlier work this paper cites.
Robot motor skill coordination with em-based reinforcement learning
P. Kormushev, S. Calinon, and D. G. Caldwell · 2010
Earlier work this paper cites.
Human preferences for robot-human hand-over configurations
M. Cakmak, S. S. Srinivasa, M. K. Lee, J. Forlizzi, and S. Kiesler · 2011
Earlier work this paper cites.
Keyframe-based learning from demonstration
B. Akgun, M. Cakmak, K. Jiang, and A. L. Thomaz · 2012
Earlier work this paper cites.
Formalizing assistive teleoperation
A. D. Dragan and S. S. Srinivasa · 2012
Earlier work this paper cites.
A bayesian approach for policy learning from trajectory preference queries
A. Wilson, A. Fern, and P. Tadepalli · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Reinforcement learning from demonstration through shaping
T. Brys, A. Harutyunyan, H. B. Suay, S. Chernova, M. E. Taylor, and A. Nowé · 2015
Earlier work this paper cites.
Data-driven motion mappings improve transparency in teleoperation
R. P. Khurshid and K. J. Kuchenbecker · 2015
Earlier work this paper cites.
Bayesian preference elicitation for multiobjective engineering design optimization
J. R. Lepird, M. P. Owen, and M. J. Kochenderfer · 2015
Earlier work this paper cites.
Active reward learning with a novel acquisition function
C. Daniel, O. Kroemer, M. Viering, J. Metz, and J. Peters · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Rl2: Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
Perspectives of approximate dynamic programming
W. B. Powell · 2016
Earlier work this paper cites.
Inverse reward design
D. Hadfield-Menell, S. Milli, P. Abbeel, S. J. Russell, and A. Dragan · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. S. Sastry, and S. A. Seshia · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Cited alongside, same era.
Do you want your autonomous car to drive like you?
C. Basu, Q. Yang, D. Hungerman, M. Sinahal, and A. D. Draqan · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Learning to reinforcement learn
J. X. Wang, Z. Kurth-Nelson, H. Soyer, J. Z. Leibo, D. Tirumala, R. Munos, C. Blundell, D. Kumaran, and M. M. Botvinick · 2017
Cited alongside, same era.
Controlling assistive robots with learned latent actions
D. P. Losey, K. Srinivasan, A. Mandlekar, A. Garg, and D. Sadigh · 2020
Later among the works it cites.
Avid: Learning multi-stage tasks via pixel-level translation of human videos
L. Smith, N. Dhawan, M. Zhang, P. Abbeel, and S. Levine · 2020
Later among the works it cites.
Active preference-based gaussian process regression for reward learning
E. Biyik, N. Huynh, M. J. Kochenderfer, and D. Sadigh · 2020
Later among the works it cites.
Avoiding side effects in complex environments
A. Turner, N. Ratzlaff, and P. Tadepalli · 2020
Later among the works it cites.
Better-than-demonstrator imitation learning via automatically-ranked demonstrations
D. S. Brown, W. Goo, and S. Niekum · 2020
Later among the works it cites.
When humans aren’t optimal: Robots that collaborate with risk-aware humans
M. Kwon, E. Biyik, A. Talati, K. Bhasin, D. P. Losey, and D. Sadigh · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Few-shot goal inference for visuomotor learning and planning
A. Xie, A. Singh, S. Levine, and C. Finn · 2018
Cited alongside, same era.
Practical obstacles to deploying active learning
D. Lowell, Z. C. Lipton, and B. C. Wallace · 2018
Cited alongside, same era.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
G. Warnell, N. Waytowich, V. Lawhern, and P. Stone · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Reptile: a scalable metalearning algorithm
A. Nichol and J. Schulman · 2018
Cited alongside, same era.
Guiding policies with language via meta-learning
J. D. Co-Reyes, A. Gupta, S. Sanjeev, N. Altieri, J. Andreas, J. DeNero, P. Abbeel, and S. Levine · 2019
Cited alongside, same era.
Later among the works it cites.
Learning from suboptimal demonstration via self-supervised reward regression
L. Chen, R. Paleja, and M. Gombolay · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control
S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. Lillicrap, N. Heess, and Y. Tassa · 2020
Later among the works it cites.
Shaping rewards for reinforcement learning with imperfect demonstrations using generative models
Y. Wu, M. Mozifian, and F. Shkurti · 2021
Later among the works it cites.
S. Karamcheti, R. Krishna, L. Fei-Fei, and C. D. Manning · 2021
Later among the works it cites.
Learning human objectives from sequences of physical corrections
M. Li, A. Canberk, D. P. Losey, and D. Sadigh · 2021
Later among the works it cites.
Replacing rewards with examples: Example-based policy search via recursive classification
B. Eysenbach, S. Levine, and R. R. Salakhutdinov · 2021
Later among the works it cites.
Polymetis
Y. Lin, A. S. Wang, G. Sutanto, A. Rai, and F. Meier · 2021
Later among the works it cites.
The effects of reward misspecification: Mapping and mitigating misaligned models
A. Pan, K. Bhatia, and J. Steinhardt · 2022
Closest in time.
High-Dimensional Data Analysis with Low-Dimensional Models: Principles, Computation, and Applications
J. Wright and Y. Ma · 2022
Closest in time.
Learning latent actions to control assistive robots
D. P. Losey, H. J. Jeon, M. Li, K. Srinivasan, A. Mandlekar, A. Garg, J. Bohg, and D. Sadigh · 2022
Closest in time.
Learning multimodal rewards from rankings
V. Myers, E. Biyik, N. Anari, and D. Sadigh · 2022
Closest in time.
Skill preferences: Learning to extract and execute robotic skills from human feedback
X. Wang, K. Lee, K. Hakhamaneshi, P. Abbeel, and M. Laskin · 2022
Closest in time.
SURF: Semi-supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning
J. Park, Y. Seo, J. Shin, H. Lee, P. Abbeel, and K. Lee · 2022
Closest in time.
Reward uncertainty for exploration in preference-based reinforcement learning
X. Liang, K. Shu, K. Lee, and P. Abbeel · 2022
Closest in time.
Personalized meta-learning for domain agnostic learning from demonstration
M. L. Schrum, E. Hedlund-Botti, and M. C. Gombolay · 2022
Closest in time.