Fetching the paper…
Reading the bibliography…
In this work, we consider the problem of model selection for deep reinforcement learning (RL) in real-world environments.
The proof and measurement of association between two things
S. Spearman, C · 1904
Earlier work this paper cites.
A generalization of sampling without replacement from a finite universe
Horvitz, D. G. and Thompson, D. J · 1952
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R. S., and Singh, S · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
A generalization error for Q-Learning
Murphy, S. A · 2005
Earlier work this paper cites.
Bias and variance approximation in value function estimates
Mannor, S., Simester, D., Sun, P., and Tsitsiklis, J. N · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Crossing the reality gap in evolutionary robotics by promoting transferable controllers
Koos, S., Mouret, J.-B., and Doncieux, S · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Dudik, M., Langford, J., and Li, L · 2011
Earlier work this paper cites.
Model selection in reinforcement learning
Farahmand, A.-M. and Szepesvári, C · 2011
Earlier work this paper cites.
The transferability approach: Crossing the reality gap in evolutionary robotics
Koos, S., Mouret, J.-B., and Doncieux, S · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Doubly robust policy evaluation and optimization
Dudík, M., Erhan, D., Langford, J., Li, L., et al · 2014
Cited alongside, same era.
Weighted importance sampling for off-policy learning with linear function approximation
Mahmood, A. R., van Hasselt, H. P., and Sutton, R. S · 2014
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Rudder: Return decomposition for delayed rewards
Arjona-Medina, J. A., Gillhofer, M., Widrich, M., Unterthiner, T., Brandstetter, J., and Hochreiter, S · 2018
Later among the works it cites.
Stochastic variational video prediction
Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S · 2018
Later among the works it cites.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2018
Later among the works it cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Personalized ad recommendation systems for life-time value optimization with guarantees
Theocharous, G., Thomas, P. S., and Ghavamzadeh, M · 2015
Cited alongside, same era.
High-Confidence Off-Policy evaluation
Thomas, P. S., Theocharous, G., and Ghavamzadeh, M · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Data-Efficient Off-Policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Bootstrapping with models: Confidence intervals for Off-Policy evaluation
Hanna, J. P., Stone, P., and Niekum, S · 2017
Cited alongside, same era.
Lee, A. X., Zhang, R., Ebert, F., Abbeel, P., Finn, C., and Levine, S · 2018
Later among the works it cites.
Representation balancing mdps for off-policy policy evaluation
Liu, Y., Gottesman, O., Raghu, A., Komorowski, M., Faisal, A., Doshi-Velez, F., and Brunskill, E · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Later among the works it cites.
Gotta learn fast: A new benchmark for generalization in rl
Nichol, A., Pfau, V., Hesse, C., Klimov, O., and Schulman, J · 2018
Later among the works it cites.
Deep reinforcement learning for Vision-Based robotic grasping: A simulated comparative evaluation of Off-Policy methods
Quillen, D., Jang, E., Nachum, O., Finn, C., Ibarz, J., and Levine, S · 2018
Later among the works it cites.
Can deep reinforcement learning solve erdos-selfridge-spencer games?
Raghu, M., Irpan, A., Andreas, J., Kleinberg, R., Le, Q., and Kleinberg, J · 2018
Later among the works it cites.
Learning by playing-solving sparse reward tasks from scratch
Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., Van de Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T · 2018
Later among the works it cites.
Sim-to-real via sim-to-sim: Data-efficient robotic grasping via randomized-to-canonical adaptation networks
James, S., Wohlhart, P., Kalakrishnan, M., Kalashnikov, D., Irpan, A., Ibarz, J., Levine, S., Hadsell, R., and Bousmalis, K · 2019
Closest in time.