Fetching the paper…
Reading the bibliography…
We consider a setting for Inverse Reinforcement Learning (IRL) where the learner is extended with the ability to actively select multiple environments, observing an agent's behavior on each environment.
Hit-and-run mixes fast
L. Lovász · 1999
Earlier work this paper cites.
Making rational decisions using adaptive utility elicitation
U. Chajewska, D. Koller, and R. Parr · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng and S. J. Russell · 2000
Earlier work this paper cites.
Effective reinforcement learning for mobile robots
W. D. Smart and L. P. Kaelbling · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Maximum margin planning
N. D. Ratliff, J. A. Bagnell, and M. A. Zinkevich · 2006
Earlier work this paper cites.
An application of reinforcement learning to aerobatic helicopter flight
P. Abbeel, A. Coates, M. Quigley, and A. Y. Ng · 2007
Earlier work this paper cites.
Bayesian inverse reinforcement learning
D. Ramachandran and E. Amir · 2007
Cited alongside, same era.
A game-theoretic approach to apprenticeship learning
U. Syed and R. E. Schapire · 2007
Cited alongside, same era.
Theory of games and economic behavior (60th Anniversary Commemorative Edition)
J. Von Neumann and O. Morgenstern · 2007
Cited alongside, same era.
Learning for control from multiple demonstrations
A. Coates, P. Abbeel, and A. Y. Ng · 2008
Cited alongside, same era.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Cited alongside, same era.
Apprenticeship learning for helicopter control
A. Coates, P. Abbeel, and A. Y. Ng · 2009
Cited alongside, same era.
Regret-based reward elicitation for markov decision processes
K. Regan and C. Boutilier · 2009
Later among the works it cites.
Adaptive submodularity: A new approach to active learning and stochastic optimization
D. Golovin and A. Krause · 2010
Later among the works it cites.
Interactive submodular set cover
A. Guillory and J. Bilmes · 2010
Later among the works it cites.
Robust policy computation in reward-uncertain mdps using nondominated policies
K. Regan and C. Boutilier · 2010
Later among the works it cites.
Eliciting additive reward functions for markov decision processes
K. Regan and C. Boutilier · 2011
Later among the works it cites.
Preference elicitation and inverse reinforcement learning
C. A. Rothkopf and C. Dimitrakakis · 2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Active learning for reward estimation in inverse reinforcement learning
M. Lopes, F. Melo, and L. Montesano · 2009
Cited alongside, same era.