2020

Value Function Approximations via Kernel Embeddings for No-Regret Reinforcement Learning

Chowdhury, Sayak Ray, Oliveira, Rafael

Understand

We consider the regret minimization problem in reinforcement learning (RL) in the episodic setting.

  • In many real-world RL environments, the state and action spaces are continuous or very large.
  • Existing approaches establish regret guarantees by either a low-dimensional representation of the stochastic transition model or an approximation of the $Q$-functions.
  • However, the understanding of function approximation schemes for state-value functions largely remains missing.

Reading the bibliography…