Fetching the paper…
Reading the bibliography…
This paper studies offline Imitation Learning (IL) where an agent learns to imitate an expert demonstrator without additional online environment interactions.
ALVINN: An autonomous land vehicle in a neural network
D. A. Pomerlau · 1989
Earlier work this paper cites.
Upper and lower bounds on the learning curve for gaussian processes
C. K. Williams and F. Vivarelli · 2000
Earlier work this paper cites.
Learning curves for gaussian process regression: Approximations and bounds
P. Sollich and A. Halees · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Local rademacher complexities
P. L. Bartlett, O. Bousquet, and S. Mendelson · 2005
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
D. Ernst, P. Geurts, and L. Wehenkel · 2005
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
C. E. Rasmussen and C. K. I. Williams · 2005
Earlier work this paper cites.
Learning bounds for kernel regression using effective data dimensionality
T. Zhang · 2005
Earlier work this paper cites.
Gaussian processes and reinforcement learning for identification and control of an autonomous blimp
J. Ko, D. J. Klein, D. Fox, and D. Haehnel · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
A. Antos, C. Szepesvári, and R. Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
R. Munos and C. Szepesvári · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Earlier work this paper cites.
Information consistency of nonparametric gaussian process methods
M. W. Seeger, S. M. Kakade, and D. P. Foster · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Efficient reductions for imitation learning
S. Ross and J. A. Bagnell · 2010
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
N. Srinivas, A. Krause, S. Kakade, and M. Seeger · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and J. Bagnell · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
A cascaded supervised learning approach to inverse reinforcement learning
E. Klein, B. Piot, M. Geist, and O. Pietquin · 2013
Earlier work this paper cites.
Finite-time analysis of kernelised contextual bandits
M. Valko, N. Korda, R. Munos, I. Flaounas, and N. Cristianini · 2013
Earlier work this paper cites.
Reinforcement and imitation learning via interactive no-regret learning
S. Ross and J. A. Bagnell · 2014
Earlier work this paper cites.
Learning to search for dependencies
K.-W. Chang, H. He, H. Daumé III, and J. Langford · 2015
Earlier work this paper cites.
Asymptotic analysis of the learning curve for gaussian process regression
L. Le Gratiet, L. Le Gratiet, J. Garnier, and J. Garnier · 2015
Earlier work this paper cites.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
On the equivalence between kernel quadrature rules and random feature expansions
F. Bach · 2017
Earlier work this paper cites.
Goal-driven dynamics learning via bayesian optimization
S. Bansal, R. Calandra, T. Xiao, S. Levine, and C. J. Tomiin · 2017
Earlier work this paper cites.
Iterative noise injection for scalable imitation learning
M. Laskey, J. Lee, W. Y. Hsieh, R. Liaw, J. Mahler, R. Fox, and K. Goldberg · 2017
Cited alongside, same era.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
W. Sun, A. Venkatraman, G. J. Gordon, B. Boots, and J. A. Bagnell · 2017
Cited alongside, same era.
Efficient exploration through bayesian deep q-networks
K. Azizzadenesheli, E. Brunskill, and A. Anandkumar · 2018
Cited alongside, same era.
A general safety framework for learning-based control in uncertain robotic systems
J. F. Fisac, A. K. Akametalu, M. N. Zeilinger, S. Kaynama, J. Gillula, and C. J. Tomlin · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
I. Osband, J. Aslanides, and A. Cassirer · 2018
Cited alongside, same era.
Information theoretic regret bounds for online nonlinear control
S. Kakade, A. Krishnamurthy, K. Lowrey, M. Ohnishi, and W. Sun · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Batch policy learning in average reward markov decision processes
P. Liao, Z. Qi, and S. Murphy · 2020
Later among the works it cites.
Provably good batch off-policy reinforcement learning without great exploration
Y. Liu, A. Swaminathan, A. Agarwal, and E. Brunskill · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2018
Cited alongside, same era.
An uncertainty-based control lyapunov approach for control-affine systems modeled by gaussian process
J. Umlauft, L. Pöhler, and S. Hirche · 2018
Cited alongside, same era.
Reinforcement learning: Theory and algorithms
A. Agarwal, N. Jiang, S. M. Kakade, and W. Sun · 2019
Cited alongside, same era.
Disagreement-regularized imitation learning
K. Brantley, W. Sun, and M. Henaff · 2019
Cited alongside, same era.
Gaussian process optimization with adaptive sketching: Scalable and no regret
D. Calandriello, L. Carratino, A. Lazaric, M. Valko, and L. Rosasco · 2019
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
J. Chen and N. Jiang · 2019
Cited alongside, same era.
Online learning in kernelized markov decision processes
S. R. Chowdhury and A. Gopalan · 2019
Cited alongside, same era.
Deployment-efficient reinforcement learning via model-based offline optimization
T. Matsushima, H. Furuta, Y. Matsuo, O. Nachum, and S. Gu · 2020
Later among the works it cites.
Toward the fundamental limits of imitation learning
N. Rajaraman, L. F. Yang, J. Jiao, and K. Ramachandran · 2020
Later among the works it cites.
Sqil: Imitation learning via reinforcement learning with sparse rewards
S. Reddy, A. D. Dragan, and S. Levine · 2020
Later among the works it cites.
Stable policy optimization via off-policy divergence regularization
A. Touati, A. Zhang, J. Pineau, and P. Vincent · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
M. Uehara, J. Huang, and N. Jiang · 2020
Later among the works it cites.
What are the statistical limits of offline rl with linear function approximation?
R. Wang, D. P. Foster, and S. M. Kakade · 2020
Later among the works it cites.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
T. Xie and N. Jiang · 2020
Later among the works it cites.
Provably efficient reinforcement learning with kernel and neural function approximations
Z. Yang, C. Jin, Z. Wang, M. Wang, and M. Jordan · 2020
Later among the works it cites.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
M. Yin and Y.-X. Wang · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
A. Zanette · 2020
Later among the works it cites.
Gendice: Generalized offline estimation of stationary values
R. Zhang, B. Dai, L. Li, and D. Schuurmans · 2020
Later among the works it cites.
Scalable bayesian inverse reinforcement learning
A. J. Chan and M. van der Schaar · 2021
Closest in time.
Bilinear classes: A structural framework for provable generalization in rl
S. S. Du, S. M. Kakade, J. D. Lee, S. Lovett, G. Mahajan, W. Sun, and R. Wang · 2021
Closest in time.
Risk bounds and rademacher complexity in batch reinforcement learning
Y. Duan, C. Jin, and Z. Li · 2021
Closest in time.
Continuous doubly constrained batch reinforcement learning
R. Fakoor, J. Mueller, P. Chaudhari, and A. J. Smola · 2021
Closest in time.
Optimism is all you need: Model-based imitation learning from observation alone
R. Kidambi, J. Chang, and W. Sun · 2021
Closest in time.
Visual adversarial imitation learning using variational models
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn · 2021
Closest in time.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
P. Rashidinejad, B. Zhu, C. Ma, J. Jiao, and S. Russell · 2021
Closest in time.
Feedback in imitation learning: The three regimes of covariate shift
J. Spencer, S. Choudhury, A. Venkatraman, B. Ziebart, and J. A. Bagnell · 2021
Closest in time.
M. Uehara, M. Imaizumi, N. Jiang, N. Kallus, W. Sun, and T. Xie · 2021
Closest in time.
Near-optimal offline reinforcement learning via double variance reduction
M. Yin, Y. Bai, and Y.-X. Wang · 2021
Closest in time.
Combo: Conservative offline model-based policy optimization
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Closest in time.