Understand
We consider the problem of learning to behave optimally in a Markov Decision Process when a reward function is not specified, but instead we have access to a set of demonstrators of varying performance.
- We assume the demonstrators are classified into one of k ranks, and use ideas from ordinal regression to find a reward function that maximizes the margin between the different ranks.
- This approach is based on the idea that agents should not only learn how to behave from experts, but also how not to behave from non-experts.
- We show there are MDPs where important differences in the reward function would be hidden from existing algorithms by the behaviour of the expert.
Built on
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Brown, D. S., Goo, W., Nagarajan, P., and Niekum, S. (2019) · 1904
Earlier work this paper cites.
ALVINN: An Autonomous Land Vehicle in a Neural Network
Pomerleau, D. (1989) · 1989
Earlier work this paper cites.
A robot controller using learning by imitation
Hayes, G. and Demiris, J. (1994) · 1994
Earlier work this paper cites.
Uncovering cabdrivers’ behavior patterns from their digital traces
Liu, L., Andris, C., and Ratti, C. (1994) · 1994
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Similar
Algorithms for inverse reinforcement learning
Ng, A. Y. and Russell, S. (2000) · 2000
Cited alongside, same era.
Learning movement sequences from demonstration
Amit, R. and Mataric, M. (2002) · 2002
Cited alongside, same era.
Ranking with Large Margin Principle: Two Approaches
Shashua, A. and Levin, A. (2002) · 2002
Cited alongside, same era.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y. (2004) · 2004
Cited alongside, same era.
Maximum margin planning
Ratliff, N. D., Bagnell, J. A., and Zinkevich, M. A. (2006) · 2006
Cited alongside, same era.
Then
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A., Bagnell, J. A., and Dey, A. K. (2008) · 2008
Later among the works it cites.
Taxi-Aware Map: Identifying and predicting vacant taxis in the city
Phithakkitnukoon, S., Veloso, M., Biderman, A., Bento, C., and Ratti, C. (2010) · 2010
Later among the works it cites.
Hunting or waiting? discovering passenger-finding strategies from a large-scale real-world taxi dataset
Li, B., Zhang, D., Sun, L., Chen, C., Li, S., Qi, G., and Yang, Q. (2011) · 2011
Later among the works it cites.
Waiting/cruising location recommendation for efficient taxi business
Takayama, T., Matsumoto, K., and Kumagai, A. (2011) · 2011
Later among the works it cites.
Efficient probabilistic performance bounds for inverse reinforcement learning
Brown, D. S. and Niekum, S. (2018) · 2018
Later among the works it cites.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…