2022

Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions

Escontrela, Alejandro, Peng, Xue Bin, Yu, Wenhao et al.

Understand

Training a high-dimensional simulated agent with an under-specified reward function often leads the agent to learn physically infeasible strategies that are ineffective when deployed in the real world.

  • To mitigate these unnatural behaviors, reinforcement learning practitioners often utilize complex reward functions that encourage physically plausible behaviors.
  • However, a tedious labor-intensive tuning process is often required to create hand-designed rewards which might not easily generalize across platforms and tasks.
  • We propose substituting complex reward functions with "style rewards" learned from a dataset of motion capture demonstrations.

Reading the bibliography…