Fetching the paper…
Reading the bibliography…
Reinforcement learning from large-scale offline datasets provides us with the ability to learn policies without potentially unsafe or impractical exploration.
Challenges of real-world reinforcement learning
Dulac-Arnold, G., Mankowitz, D. J., and Hester, T. (2019) · 1904
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
OpenAI, Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., Schneider, J., Tezak, N., Tworek, J., Welinder, P., Weng, L., Yuan, Q., Zaremba, W., and Zhang, L. (2019) · 1910
Earlier work this paper cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T. P., and Silver, D. (2019) · 1911
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., Tucker, G., and Nachum, O. (2019) · 1911
Earlier work this paper cites.
Estimating the mean and variance of the target probability distribution
Nix, D. A. and Weigend, A. S. (1994) · 1994
Earlier work this paper cites.
Automatic data augmentation for generalization in deep reinforcement learning
Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., and Fergus, R. (2020) · 2006
Earlier work this paper cites.
Contextual markov decision processes
Hallak, A., Castro, D. D., and Mannor, S. (2015) · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Earlier work this paper cites.
Hidden parameter markov decision processes: A semiparametric regression approach for discovering latent task parametrizations
Doshi-Velez, F. and Konidaris, G. (2016) · 2016
Earlier work this paper cites.
Continuous deep q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S. (2016) · 2016
Earlier work this paper cites.
Reinforcement learning for pivoting task
Antonova, R., Cruciani, S., Smith, C., and Kragic, D. (2017) · 2017
Earlier work this paper cites.
Benchmark environments for multitask learning in continuous domains
Henderson, P., Chang, W.-D., Shkurti, F., Hansen, J., Meger, D., and Dudek, G. (2017) · 2017
Earlier work this paper cites.
Transferring end-to-end visuomotor control from simulation to real world for a multi-stage task
James, S., Davison, A. J., and Johns, E. (2017) · 2017
Earlier work this paper cites.
Robust and efficient transfer learning with hidden parameter markov decision processes
Killian, T., Konidaris, G., and Doshi-Velez, F. (2017) · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C. (2017) · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P. (2017) · 2017
Earlier work this paper cites.
Preparing for the unknown: Learning a universal policy with online system identification
Yu, W., Tan, J., Liu, C. K., and Turk, G. (2017) · 2017
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S. (2018) · 2018
Earlier work this paper cites.
Model-based reinforcement learning via meta-policy optimization
Clavera, I., Rothfuss, J., Schulman, J., Fujita, Y., Asfour, T., and Abbeel, P. (2018) · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J. (2018) · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., and Levine, S. (2018) · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P. (2018) · 2018
Cited alongside, same era.
Markov decision processes with continuous side information
Modi, A., Jiang, N., Singh, S., and Tewari, A. (2018) · 2018
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
Peng, X. B., Andrychowicz, M., Zaremba, W., and Abbeel, P. (2018) · 2018
Cited alongside, same era.
Generalizing to unseen domains via adversarial data augmentation
Counterfactual data augmentation using locally factored dynamics
Pitis, S., Creager, E., and Garg, A. (2020) · 2020
Later among the works it cites.
Meta-learning requires meta-augmentation
Rajendran, J., Irpan, A., and Jang, E. (2020) · 2020
Later among the works it cites.
Trajectory-wise multiple choice learning for dynamics generalization in reinforcement learning
Seo, Y., Lee, K., Clavera, I., Kurutach, T., Shin, J., and Abbeel, P. (2020) · 2020
Later among the works it cites.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B. (2020) · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T. (2020) · 2020
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning
Zintgraf, L., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Volpi, R., Namkoong, H., Sener, O., Duchi, J. C., Murino, V., and Savarese, S. (2018) · 2018
Cited alongside, same era.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Clavera, I., Nagabandi, A., Liu, S., Fearing, R. S., Abbeel, P., Levine, S., and Finn, C. (2019) · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J. (2019) · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S. (2019) · 2019
Cited alongside, same era.
Deep online learning via meta-learning: Continual adaptation for model-based RL
Nagabandi, A., Finn, C., and Levine, S. (2019) · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D. (2019) · 2019
Cited alongside, same era.
Environment probing interaction policies
Zhou, W., Pinto, L., and Gupta, A. (2019) · 2019
Cited alongside, same era.
Later among the works it cites.
OPAL: Offline primitive discovery for accelerating offline reinforcement learning
Ajay, A., Kumar, A., Agrawal, P., Levine, S., and Nachum, O. (2021) · 2021
Closest in time.
Model-based offline planning
Argenson, A. and Dulac-Arnold, G. (2021) · 2021
Closest in time.
D4{rl}: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. (2021) · 2021
Closest in time.
Self-supervised policy adaptation during deployment
Hansen, N., Jangir, R., Sun, Y., Alenyà, G., Abbeel, P., Efros, A. A., Pinto, L., and Wang, X. (2021) · 2021
Closest in time.
Transient non-stationarity and generalisation in deep reinforcement learning
Igl, M., Farquhar, G., Luketina, J., Boehmer, W., and Whiteson, S. (2021) · 2021
Closest in time.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R. (2021) · 2021
Closest in time.
Linear representation meta-reinforcement learning for instant adaptation
Peng, M., Zhu, B., and Jiao, J. (2021) · 2021
Closest in time.
On pathologies in kl-regularized reinforcement learning from expert demonstrations
Rudner, T. G. J., Lu, C., Osborne, M. A., Gal, Y., and Teh, Y. W. (2021) · 2021
Closest in time.
Dropout’s dream land: Generalization from learned simulators to reality
Wellmer, Z. and Kwok, J. T. (2021) · 2021
Closest in time.
Learning robust state abstractions for hidden-parameter block mdps
Zhang, A., Sodhani, S., Khetarpal, K., and Pineau, J. (2021) · 2021
Closest in time.
Exploration in approximate hyper-state space for meta reinforcement learning
Zintgraf, L. M., Feng, L., Lu, C., Igl, M., Hartikainen, K., Hofmann, K., and Whiteson, S. (2021) · 2021
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D. (2019) · 2062
Closest in time.