Fetching the paper…
Reading the bibliography…
We introduce the "inverse bandit" problem of estimating the rewards of a multi-armed bandit instance from observing the learning process of a low-regret demonstrator.
Asymptotically efficient adaptive allocation rules
T. L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Learning agents for uncertain environments
S. Russell · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng, S. J. Russell, et al · 2000
Earlier work this paper cites.
Behavioral models of strategies in multi-armed bandit problems
C. M. Anderson · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Cortical substrates for exploratory decisions in humans
N. D. Daw, J. P. O’doherty, P. Dayan, B. Seymour, and R. J. Dolan · 2006
Earlier work this paper cites.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
E. Even-Dar, S. Mannor, and Y. Mansour · 2006
Earlier work this paper cites.
Bayesian inverse reinforcement learning
D. Ramachandran and E. Amir · 2007
Earlier work this paper cites.
Drosophila rnai screen identifies host genes important for influenza virus replication
L. Hao, A. Sakurai, T. Watanabe, E. Sorensen, C. A. Nidom, M. A. Newton, P. Ahlquist, and Y. Kawaoka · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Adapting to the shifting intent of search queries
U. Syed, A. Slivkins, and N. Mishra · 2010
Earlier work this paper cites.
Novelty and inductive generalization in human reinforcement learning
S. J. Gershman and Y. Niv · 2015
Earlier work this paper cites.
Between imitation and intention learning
J. MacGlashan and M. L. Littman · 2015
Earlier work this paper cites.
Learning and decisions in contextual multi-armed bandit tasks
E. Schulz, E. Konstantinidis, and M. Speekenbrink · 2015
Earlier work this paper cites.
Uncertainty and exploration in a restless bandit problem
M. Speekenbrink and E. Konstantinidis · 2015
Cited alongside, same era.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Cited alongside, same era.
Empirical priors for reinforcement learning models
S. J. Gershman · 2016
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
Top arm identification in multi-armed bandits with batch arm pulls
K.-S. Jun, K. Jamieson, R. Nowak, and X. Zhu · 2016
Cited alongside, same era.
On the complexity of best-arm identification in multi-armed bandit models
E. Kaufmann, O. Cappé, and A. Garivier · 2016
Why adaptively collected data have negative bias and how to correct for it
X. Nie, X. Tian, J. Taylor, and J. Zou · 2018
Later among the works it cites.
Joint modeling of reaction times and choice improves parameter identifiability in reinforcement learning models
I. C. Ballard and S. M. McClure · 2019
Later among the works it cites.
The assistive multi-armed bandit
L. Chan, D. Hadfield-Menell, S. Srinivasa, and A. Dragan · 2019
Later among the works it cites.
Learning from a learner
A. Jacq, M. Geist, A. Paiva, and O. Pietquin · 2019
Later among the works it cites.
Are sample means in multi-armed bandits positively or negatively biased?
J. Shin, A. Ramdas, and A. Rinaldo · 2019
Later among the works it cites.
Introduction to multi-armed bandits
A. Slivkins · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning from demonstration for shaping through inverse reinforcement learning
H. B. Suay, T. Brys, M. E. Taylor, and S. Chernova · 2016
Cited alongside, same era.
Repeated inverse reinforcement learning
K. Amin, N. Jiang, and S. Singh · 2017
Cited alongside, same era.
Bandit models of human behavior: Reward processing in mental disorders
D. Bouneffouf, I. Rish, and G. A. Cecchi · 2017
Cited alongside, same era.
One-shot visual imitation learning via meta-learning
C. Finn, T. Yu, T. Zhang, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2017
Cited alongside, same era.
Reinforcement learning from imperfect demonstrations
Y. Gao, H. Xu, J. Lin, F. Yu, S. Levine, and T. Darrell · 2018
Cited alongside, same era.
Later among the works it cites.
Imitation learning from imperfect demonstration
Y.-H. Wu, N. Charoenphakdee, H. Bao, V. Tangkaratt, and M. Sugiyama · 2019
Later among the works it cites.
Closed-loop optimization of fast-charging protocols for batteries with machine learning
P. M. Attia, A. Grover, N. Jin, K. A. Severson, T. M. Markov, Y.-H. Liao, M. H. Chen, B. Cheong, N. Perkins, Z. Yang, et al · 2020
Later among the works it cites.
On-policy robot imitation learning from a converging supervisor
A. Balakrishna, B. Thananjeyan, J. Lee, F. Li, A. Zahed, J. E. Gonzalez, and K. Goldberg · 2020
Later among the works it cites.
Identifying reward functions using anchor actions
S. Geng, H. Nassif, C. A. Manzanares, A. M. Reppen, and R. Sircar · 2020
Later among the works it cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
H. J. Jeon, S. Milli, and A. D. Dragan · 2020
Later among the works it cites.
Bandit algorithms
T. Lattimore and C. Szepesvári · 2020
Later among the works it cites.
Inverse reinforcement learning from a gradient-based learner
G. Ramponi, G. Drappo, and M. Restelli · 2020
Later among the works it cites.
Inverse reinforcement learning from like-minded teachers
R. Noothigattu, T. Yan, and A. D. Procaccia · 2021
Closest in time.