2022

VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training

Ma, Yecheng Jason, Sodhani, Shagun, Jayaraman, Dinesh et al.

Understand

Reward and representation learning are two long-standing challenges for learning an expanding set of robot manipulation skills from sensory observations.

  • Given the inherent cost and scarcity of in-domain, task-specific robot data, learning from large, diverse, offline human videos has emerged as a promising path towards acquiring a generally useful visual representation for control; however, how these human videos can be used for general-purpose reward learning remains an open question.
  • We introduce $\textbf{V}$alue-$\textbf{I}$mplicit $\textbf{P}$re-training (VIP), a self-supervised pre-trained visual representation capable of generating dense and smooth reward functions for unseen robotic tasks.
  • VIP casts representation learning from human videos as an offline goal-conditioned reinforcement learning problem and derives a self-supervised dual goal-conditioned value-function objective that does not depend on actions, enabling pre-training on unlabeled human videos.

Reading the bibliography…