Fetching the paper…
Reading the bibliography…
When learning policies for real-world domains, two important questions arise: (i) how to efficiently use pre-collected off-policy, non-optimal behavior data; and (ii) how to mediate among different competing objectives and constraints.
Problem complexity and method efficiency in optimization
Nemirovsky, A. S. and Yudin, D. B · 1983
Earlier work this paper cites.
Strict stationarity of generalized autoregressive processes
Bougerol, P. and Picard, N · 1992
Earlier work this paper cites.
Sphere packing numbers for subsets of the boolean n-cube with bounded vapnik-chervonenkis dimension
Haussler, D · 1995
Earlier work this paper cites.
The effect of representation and knowledge on goal-directed exploration with reinforcement-learning algorithms
Koenig, S. and Simmons, R. G · 1996
Earlier work this paper cites.
Efficient agnostic learning of neural networks with bounded fan-in
Lee, W. S., Bartlett, P. L., and Williamson, R. C · 1996
Earlier work this paper cites.
Exponentiated gradient versus gradient descent for linear predictors
Kivinen, J. and Warmuth, M. K · 1997
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Freund, Y. and Schapire, R. E · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R. S., and Singh, S. P · 2000
Earlier work this paper cites.
The elements of statistical learning
Friedman, J., Hastie, T., and Tibshirani, R · 2001
Earlier work this paper cites.
Off-policy temporal difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Kernel-based reinforcement learning
Ormoneit, D. and Sen, Ś · 2002
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Convex optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
A distribution-free theory of nonparametric regression
Györfi, L., Kohler, M., Krzyzak, A., and Walk, H · 2006
Earlier work this paper cites.
Performance bounds in l_p-norm for approximate value iteration
Munos, R · 2007
Earlier work this paper cites.
Theory of games and economic behavior (commemorative edition)
Von Neumann, J. and Morgenstern, O · 2007
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Cited alongside, same era.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Cited alongside, same era.
Regularized policy iteration
Farahmand, A. M., Ghavamzadeh, M., Mannor, S., and Szepesvári, C · 2009
Cited alongside, same era.
Reinforcement learning for robot soccer
Riedmiller, M., Gabel, T., Hafner, R., and Lange, S · 2009
Cited alongside, same era.
Finite-sample analysis of lstd
Lazaric, A., Ghavamzadeh, M., and Munos, R · 2010
Cited alongside, same era.
Finite-sample analysis of bellman residual minimization
Maillard, O.-A., Munos, R., Lazaric, A., and Ghavamzadeh, M · 2010
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Later among the works it cites.
Batch learning from logged bandit feedback through counterfactual risk minimization
Swaminathan, A. and Joachims, T · 2015
Later among the works it cites.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Later among the works it cites.
Smooth imitation learning for online sequence prediction
Le, H. M., Kang, A., Yue, Y., and Carr, P · 2016
Later among the works it cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Later among the works it cites.
Guided policy search via approximate mirror descent
Montgomery, W. H. and Levine, S · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Cited alongside, same era.
Approximate policy iteration: A survey and some new methods
Bertsekas, D. P · 2011
Cited alongside, same era.
Chance-constrained optimal path planning with obstacles
Blackmore, L., Ono, M., and Williams, B. C · 2011
Cited alongside, same era.
Doubly robust policy evaluation and learning
Dudík, M., Langford, J., and Li, L · 2011
Cited alongside, same era.
Transfer from multiple mdps
Lazaric, A. and Restelli, M · 2011
Cited alongside, same era.
Sample-efficient batch reinforcement learning for dialogue management optimization
Pietquin, O., Geist, M., Chandramohan, S., and Frezza-Buet, H · 2011
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Later among the works it cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Later among the works it cites.
Nearly-tight vc-dimension bounds for piecewise linear neural networks
Bartlett, P. L., Harvey, N., Liaw, C., and Mehrabian, A · 2017
Later among the works it cites.
Using options and covariance testing for long horizon off-policy policy evaluation
Guo, Z., Thomas, P. S., and Brunskill, E · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Wang, Y.-X., Agarwal, A., and Dudík, M · 2017
Later among the works it cites.
A reductions approach to fair classification
Agarwal, A., Beygelzimer, A., Dudík, M., Langford, J., and Wallach, H · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Farajtabar, M., Chow, Y., and Ghavamzadeh, M · 2018
Later among the works it cites.
Ha, D. and Schmidhuber, J · 2018
Later among the works it cites.
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Osband, I., et al · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D · 2018
Later among the works it cites.
Self-imitation learning
Oh, J., Guo, Y., Singh, S., and Lee, H · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Accelerating imitation learning with predictive models
Cheng, C.-A., Yan, X., Theodorou, E., and Boots, B · 2019
Closest in time.
Model-predictive policy learning with uncertainty regularization for driving in dense traffic
Henaff, M., Canziani, A., and LeCun, Y · 2019
Closest in time.