Fetching the paper…

Learning from an Exploring Demonstrator: Optimal Reward Estimation for Bandits · Around