2022

Reward Uncertainty for Exploration in Preference-based Reinforcement Learning

Liang, Xinran, Shu, Katherine, Lee, Kimin et al.

Understand

Conveying complex objectives to reinforcement learning (RL) agents often requires meticulous reward engineering.

  • Preference-based RL methods are able to learn a more flexible reward model based on human preferences by actively incorporating human feedback, i.e.
  • teacher's preferences between two clips of behaviors.
  • However, poor feedback-efficiency still remains a problem in current preference-based RL algorithms, as tailored human feedback is very expensive.

Reading the bibliography…