Understand
In reinforcement learning an agent interacts with the environment by taking actions and observing the next state and reward.
- When sampled probabilistically, these state transitions, rewards, and actions can all induce randomness in the observed long-term return.
- Traditionally, reinforcement learning algorithms average over this randomness to estimate the value function.
- In this paper, we build on recent work advocating a distributional approach to reinforcement learning in which the distribution over returns is modeled explicitly instead of only estimating the mean.
Built on
Nothing clear enough to list yet.
Similar
Nothing clear enough to list yet.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…