Fetching the paper…
Reading the bibliography…
In this tutorial article, we aim to provide the reader with the conceptual tools needed to get started on research on offline reinforcement learning algorithms: reinforcement learning algorithms that utilize previously collected data, without additional online data collection.
Diagnosing bottlenecks in deep Q-learning algorithms
Fu, J., Kumar, A., Soh, M., and Levine, S. (2019) · 1902
Earlier work this paper cites.
Pipps: Flexible model-based policy search robust to the curse of chaos
Parmas, P., Rasmussen, C. E., Peters, J., and Doya, K. (2019) · 1902
Earlier work this paper cites.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al. (2019a) · 1903
Earlier work this paper cites.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al. (2019b) · 1903
Earlier work this paper cites.
Domain adaptation with asymmetrically-relaxed distribution alignment
Wu, Y., Winston, E., Kaushik, D., and Lipton, Z. (2019b) · 1903
Earlier work this paper cites.
Off-policy policy gradient with state distribution correction
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E. (2019) · 1904
Earlier work this paper cites.
Learning when-to-treat policies
Nie, X., Brunskill, E., and Wager, S. (2019) · 1905
Earlier work this paper cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M. (2019) · 1907
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. (2019) · 1907
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. (2019) · 1907
Earlier work this paper cites.
Safe policy improvement with soft baseline bootstrapping
Nadjahi, K., Laroche, R., and Combes, R. T. d. (2019) · 1907
Earlier work this paper cites.
Trajectory-wise control variates for variance reduction in policy gradient methods
Cheng, C.-A., Yan, X., and Boots, B. (2019) · 1908
Earlier work this paper cites.
A framework for data-driven robotics
Cabi, S., Colmenarejo, S. G., Novikov, A., Konyushkova, K., Reed, S., Jeong, R., Żołna, K., Aytar, Y., Budden, D., Vecerik, M., et al. (2019) · 1909
Earlier work this paper cites.
Kallus, N. and Uehara, M. (2019a) · 1909
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
Dasari, S., Ebert, F., Tian, S., Nair, S., Bucher, B., Schmeckpeper, K., Singh, S., Levine, S., and Finn, C. (2019) · 1910
Earlier work this paper cites.
From importance sampling to doubly robust policy gradient
Huang, J. and Jiang, N. (2019) · 1910
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S. (2019) · 1910
Earlier work this paper cites.
Doubly robust bias reduction in infinite horizon off-policy estimation
Tang, Z., Feng, Y., Li, L., Zhou, D., and Liu, Q. (2019) · 1910
Earlier work this paper cites.
Minimax weight and q-function learning for off-policy evaluation
Uehara, M. and Jiang, N. (2019) · 1910
Earlier work this paper cites.
Causality for machine learning
Schölkopf, B. (2019) · 1911
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O. (2019a) · 1911
Earlier work this paper cites.
Algaedice: Policy gradient from arbitrary experience
Nachum, O., Dai, B., Kostrikov, I., Chow, Y., Li, L., and Schuurmans, D. (2019b) · 1912
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S. (1991) · 1991
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J. (1992) · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
Tesauro, G. (1994) · 1994
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Autonomous helicopter control using reinforcement learning policy search methods
Bagnell, J. A. and Schneider, J. G. (2001) · 2001
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Murphy, S. A., van der Laan, M. J., Robins, J. M., and Group, C. P. P. R. (2001) · 2001
Earlier work this paper cites.
Reinforcement learning via fenchel-rockafellar duality
Nachum, O. and Dai, B. (2020) · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S. (2001) · 2001
Earlier work this paper cites.
Gradientdice: Rethinking generalized offline estimation of stationary values
Zhang, S., Liu, B., and Whiteson, S. (2020b) · 2001
Earlier work this paper cites.
Badgr: An autonomous self-supervised learning-based navigation system
Kahn, G., Abbeel, P., and Levine, S. (2020) · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2002) · 2002
Earlier work this paper cites.
Learning from scarce experience
Peshkin, L. and Shelton, C. R. (2002) · 2002
Earlier work this paper cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Siegel, N. Y., Springenberg, J. T., Berkenkamp, F., Abdolmaleki, A., Neunert, M., Lampe, T., Hafner, R., and Riedmiller, M. (2020) · 2002
Earlier work this paper cites.
Discor: Corrective feedback in reinforcement learning via distribution correction
Kumar, A., Gupta, A., and Levine, S. (2020a) · 2003
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R. (2003) · 2003
Earlier work this paper cites.
Learning to generalize across long-horizon tasks from human demonstrations
Mandlekar, A., Xu, D., Martín-Martín, R., Savarese, S., and Fei-Fei, L. (2020) · 2003
Earlier work this paper cites.
Black-box off-policy estimation for infinite-horizon reinforcement learning
Mousavi, A., Li, L., Liu, Q., and Zhou, D. (2020) · 2003
Earlier work this paper cites.
Batch stationary distribution estimation
Wen, J., Dai, B., Li, L., and Schuurmans, D. (2020) · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L. (2005) · 2005
Earlier work this paper cites.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T. (2020) · 2005
Earlier work this paper cites.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M. (2005) · 2005
Earlier work this paper cites.
Accelerating online reinforcement learning with offline datasets
Nair, A., Dalal, M., Gupta, A., and Levine, S. (2020) · 2006
Earlier work this paper cites.
Linearly-solvable markov decision problems
Todorov, E. (2006) · 2006
Earlier work this paper cites.
Gaussian processes and reinforcement learning for identification and control of an autonomous blimp
Ko, J., Klein, D. J., Fox, D., and Haehnel, D. (2007) · 2007
Earlier work this paper cites.
Adaptive treatment of epilepsy via batch-mode reinforcement learning
Guez, A., Vincent, R. D., Avoli, M., and Pineau, J. (2008) · 2008
Earlier work this paper cites.
Hybrid reinforcement/supervised learning of dialogue policies from fixed data sets
Henderson, J., Lemon, O., and Georgila, K. (2008) · 2008
Earlier work this paper cites.
Exploration scavenging
Langford, J., Strehl, A., and Wortman, J. (2008) · 2008
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Parr, R., Li, L., Taylor, G., Painter-Wakefield, C., and Littman, M. L. (2008) · 2008
Earlier work this paper cites.
Neuroevolutionary reinforcement learning for generalized helicopter control
Koppejan, R. and Whiteson, S. (2009) · 2009
Earlier work this paper cites.
On integral probability metrics, \ \backslash phi-divergences and binary classification
Sriperumbudur, B. K., Fukumizu, K., Gretton, A., Schölkopf, B., and Lanckriet, G. R. (2009) · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E. (2009) · 2009
Earlier work this paper cites.
Generalized off-policy actor-critic
Zhang, S., Boehmer, W., and Whiteson, S. (2019) · 2009
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Farahmand, A.-m., Szepesvári, C., and Munos, R. (2010) · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010) · 2010
Cited alongside, same era.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D. (2010) · 2010
Cited alongside, same era.
Learning from logged implicit exploration data
Strehl, A., Langford, J., Li, L., and Kakade, S. M. (2010) · 2010
Cited alongside, same era.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E. (2011) · 2011
Cited alongside, same era.
Reinforcement learning in feedback control
Hafner, R. and Riedmiller, M. (2011) · 2011
Cited alongside, same era.
Sample-efficient batch reinforcement learning for dialogue management optimization
Pietquin, O., Geist, M., Chandramohan, S., and Frezza-Buet, H. (2011) · 2011
Cited alongside, same era.
What uncertainties do we need in bayesian deep learning for computer vision?
Kendall, A. and Gal, Y. (2017) · 2017
Later among the works it cites.
Safe policy improvement with baseline bootstrapping
Laroche, R., Trichelair, P., and Combes, R. T. d. (2017) · 2017
Later among the works it cites.
1 year, 1000 km: The oxford robotcar dataset
Maddern, W., Pascoe, G., Linegar, C., and Newman, P. (2017) · 2017
Later among the works it cites.
Value prediction network
Oh, J., Singh, S., and Lee, H. (2017) · 2017
Later among the works it cites.
Why is posterior sampling better than optimism for reinforcement learning?
Osband, I. and Van Roy, B. (2017) · 2017
Later among the works it cites.
Agile autonomous driving using end-to-end deep imitation learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D. (2011) · 2011
Cited alongside, same era.
Informing sequential clinical decision-making through reinforcement learning: an empirical study
Shortreed, S. M., Laber, E., Lizotte, D. J., Stroup, T. S., Pineau, J., and Murphy, S. A. (2011) · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D. (2011) · 2011
Cited alongside, same era.
Degris, T., White, M., and Sutton, R. S. (2012) · 2012
Cited alongside, same era.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M. (2012) · 2012
Cited alongside, same era.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Tassa, Y., Erez, T., and Todorov, E. (2012) · 2012
Cited alongside, same era.
Pan, Y., Cheng, C.-A., Saigol, K., Lee, K., Yan, X., Theodorou, E., and Boots, B. (2017) · 2017
Later among the works it cites.
A reinforcement learning approach to weaning of mechanical ventilation in intensive care units
Prasad, N., Cheng, L.-F., Chivers, C., Draugelis, M., and Engelhardt, B. E. (2017) · 2017
Later among the works it cites.
Deep reinforcement learning for sepsis treatment
Raghu, A., Komorowski, M., Ahmed, I., Celi, L., Szolovits, P., and Ghassemi, M. (2017) · 2017
Later among the works it cites.
CAD2RL: Real single-image flight without a single real image
Sadeghi, F. and Levine, S. (2017) · 2017
Later among the works it cites.
Deep reinforcement learning framework for autonomous driving
Sallab, A. E., Abdou, M., Perot, E., and Yogamani, S. (2017) · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017) · 2017
Later among the works it cites.
Certifying some distributional robustness with principled adversarial training
Sinha, A., Namkoong, H., and Duchi, J. (2017) · 2017
Later among the works it cites.
Off-policy evaluation for slate recommendation
Swaminathan, A., Krishnamurthy, A., Agarwal, A., Dudik, M., Langford, J., Jose, D., and Zitouni, I. (2017) · 2017
Later among the works it cites.
Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing
Thomas, P. S., Theocharous, G., Ghavamzadeh, M., Durugkar, I., and Brunskill, E. (2017) · 2017
Later among the works it cites.
Deep reinforcement learning for automated radiation adaptation in lung cancer
Tseng, H.-H., Luo, Y., Cui, S., Chien, J.-T., Ten Haken, R. K., and El Naqa, I. (2017) · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Wang, Y.-X., Agarwal, A., and Dudik, M. (2017) · 2017
Later among the works it cites.
End-to-end offline goal-oriented dialog policy learning via policy gradient
Zhou, L., Small, K., Rokhlenko, O., and Elkan, C. (2017) · 2017
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S. (2018) · 2018
Later among the works it cites.
End-to-end driving via conditional imitation learning
Codevilla, F., Miiller, M., López, A., Koltun, V., and Dosovitskiy, A. (2018) · 2018
Later among the works it cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A., and Levine, S. (2018) · 2018
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al. (2018) · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Farajtabar, M., Chow, Y., and Ghavamzadeh, M. (2018) · 2018
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D. (2018) · 2018
Later among the works it cites.
Offline a/b testing for recommender systems
Gilotte, A., Calauzènes, C., Nedelec, T., Abraham, A., and Dollé, S. (2018) · 2018
Later among the works it cites.
Evaluating reinforcement learning algorithms in observational health settings
Gottesman, O., Johansson, F., Meier, J., Dent, J., Lee, D., Srinivasan, S., Zhang, L., Ding, Y., Wihl, D., Peng, X., et al. (2018) · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J. (2018) · 2018
Later among the works it cites.
An off-policy policy gradient theorem using emphatic weightings
Imani, E., Graves, E., and White, M. (2018) · 2018
Later among the works it cites.
Composable action-conditioned predictors: Flexible off-policy learning for robot navigation
Kahn, G., Villaflor, A., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al. (2018) · 2018
Later among the works it cites.
Stochastic primal-dual q-learning
Lee, D. and He, N. (2018) · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S. (2018) · 2018
Later among the works it cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Levine, S., Pastor, P., Krizhevsky, A., Ibarz, J., and Quillen, D. (2018) · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D. (2018) · 2018
Later among the works it cites.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Luo, Y., Xu, H., Li, Y., Tian, Y., Darrell, T., and Ma, T. (2018) · 2018
Later among the works it cites.
Mo, K., Li, H., Lin, Z., and Lee, J.-Y. (2018) · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S. (2018) · 2018
Later among the works it cites.
The uncertainty bellman equation and exploration
O’Donoghue, B., Osband, I., Munos, R., and Mnih, V. (2018) · 2018
Later among the works it cites.
Reward-estimation variance elimination in sequential decision processes
Pankov, S. (2018) · 2018
Later among the works it cites.
Deep imitative models for flexible inference, planning, and control
Rhinehart, N., McAllister, R., and Levine, S. (2018) · 2018
Later among the works it cites.
Sim-to-real: Learning agile locomotion for quadruped robots
Tan, J., Zhang, T., Coumans, E., Iscen, A., Bai, Y., Hafner, D., Bohez, S., and Vanhoucke, V. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning and the deadly triad
Van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J. (2018) · 2018
Later among the works it cites.
Supervised reinforcement learning with recurrent neural network for dynamic treatment recommendation
Wang, L., Zhang, W., He, X., and Zha, H. (2018) · 2018
Later among the works it cites.
Bdd100k: A diverse driving video database with scalable annotation tooling
Yu, F., Xian, W., Chen, Y., Liu, F., Liao, M., Madhavan, V., and Darrell, T. (2018) · 2018
Later among the works it cites.
Learning synergies between pushing and grasping with self-supervised deep reinforcement learning
Zeng, A., Song, S., Welker, S., Lee, J., Rodriguez, A., and Funkhouser, T. (2018) · 2018
Later among the works it cites.
Solar: deep structured representations for model-based reinforcement learning
Zhang, M., Vikram, S., Smith, L., Abbeel, P., Johnson, M. J., and Levine, S. (2018) · 2018
Later among the works it cites.
Closing the sim-to-real loop: Adapting simulation randomization with real world experience
Chebotar, Y., Handa, A., Makoviychuk, V., Macklin, M., Issac, J., Ratliff, N., and Fox, D. (2019) · 2019
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Gelada, C. and Bellemare, M. G. (2019) · 2019
Later among the works it cites.
Guidelines for reinforcement learning in healthcare
Gottesman, O., Johansson, F., Komorowski, M., Faisal, A., Sontag, D., Doshi-Velez, F., and Celi, L. A. (2019) · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S. (2019) · 2019
Later among the works it cites.
Learning to drive in a day
Kendall, A., Hawke, J., Janz, D., Mazur, P., Reda, D., Allen, J.-M., Lam, V.-D., Bewley, A., and Shah, A. (2019) · 2019
Later among the works it cites.
Data-driven deep reinforcement learning
Kumar, A. (2019) · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S. (2019) · 2019
Later among the works it cites.
Distributionally robust neural networks
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P. (2019) · 2019
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. (2020) · 2020
Closest in time.
Does on-policy data collection fix errors in reinforcement learning?
Kumar, A. and Gupta, A. (2020) · 2020
Closest in time.
Off-policy bandits with deficient support
Sachdeva, N., Su, Y., and Joachims, T. (2020) · 2020
Closest in time.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T. (2020) · 2020
Closest in time.
A survey of autonomous driving: Common practices and emerging technologies
Yurtsever, E., Lambert, J., Carballo, A., and Takeda, K. (2020) · 2020
Closest in time.