Understand
We consider stochastic multi-armed bandit problems with complex actions over a set of basic arms, where the decision maker plays a complex action rather than a basic arm in each round.
- The reward of the complex action is some function of the basic arms' rewards, and the feedback observed may not necessarily be the reward per-arm.
- For instance, when the complex actions are subsets of the arms, we may only observe the maximum reward over the chosen subset.
- Thus, feedback across complex actions may be coupled due to the nature of the reward function.
Built on
Nothing clear enough to list yet.
Similar
Nothing clear enough to list yet.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…