2025

XQC: Well-conditioned Optimization Accelerates Deep Reinforcement Learning

Palenicek, Daniel, Vogt, Florian, Watson, Joe et al.

Understand

Sample efficiency is a central property of effective deep reinforcement learning algorithms.

  • Recent work has improved this through added complexity, such as larger models, exotic network architectures, and more complex algorithms, which are typically motivated purely by empirical performance.
  • We take a more principled approach by focusing on the optimization landscape of the critic network.
  • Using the eigenspectrum and condition number of the critic's Hessian, we systematically investigate the impact of common architectural design decisions on training dynamics.

Reading the bibliography…