2022

When is Realizability Sufficient for Off-Policy Reinforcement Learning?

Zanette, Andrea

Understand

Model-free algorithms for reinforcement learning typically require a condition called Bellman completeness in order to successfully operate off-policy with function approximation, unless additional conditions are met.

  • However, Bellman completeness is a requirement that is much stronger than realizability and that is deemed to be too strong to hold in practice.
  • In this work, we relax this structural assumption and analyze the statistical complexity of off-policy reinforcement learning when only realizability holds for the prescribed function class.
  • We establish finite-sample guarantees for off-policy reinforcement learning that are free of the approximation error term known as inherent Bellman error, and that depend on the interplay of three factors.

Reading the bibliography…