2022

Large Language Models can Implement Policy Iteration

Brooks, Ethan, Walls, Logan, Lewis, Richard L. et al.

Understand

This work presents In-Context Policy Iteration, an algorithm for performing Reinforcement Learning (RL), in-context, using foundation models.

  • While the application of foundation models to RL has received considerable attention, most approaches rely on either (1) the curation of expert demonstrations (either through manual design or task-specific pretraining) or (2) adaptation to the task of interest using gradient methods (either fine-tuning or training of adapter layers).
  • Both of these techniques have drawbacks.
  • Collecting demonstrations is labor-intensive, and algorithms that rely on them do not outperform the experts from which the demonstrations were derived.

Reading the bibliography…