Fetching the paper…
Reading the bibliography…
We consider the problem of learning useful robotic skills from previously collected offline data without access to manually specified rewards or additional online exploration, a setting that is becoming increasingly important for scaling robot learning by reusing past robotic data.
Videoflow: A flow-based generative model for video
Kumar, M., Babaeizadeh, M., Erhan, D., Finn, C., Levine, S., Dinh, L., and Kingma, D · 1903
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P · 1993
Earlier work this paper cites.
The Cross Entropy Method: A Unified Approach To Combinatorial Optimization, Monte-Carlo Simulation (Information Science and Statistics)
Rubinstein, R. Y. and Kroese, D. P · 2004
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Mathieu, M., Couprie, C., and Lecun, Y · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J. T., Boedecker, J., and Riedmiller, M · 2015
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
Williams, G., Aldrich, A., and Theodorou, E · 2015
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., Van Hasselt, H., and Silver, D · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Crow, D., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Earlier work this paper cites.
A probabilistic data-driven model for planar pushing
Bauza, M. and Rodriguez, A · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
Finn, C. and Levine, S · 2017
Earlier work this paper cites.
Safe policy improvement with baseline bootstrapping
Laroche, R., Trichelair, P., and Combes, R. T. d · 2017
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Cited alongside, same era.
Stochastic video generation with a learned prior
Denton, E. and Fergus, R · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2018
Reinforcement learning without ground-truth state
Lin, X., Baweja, H. S., and Held, D · 2019
Later among the works it cites.
Learning latent plans from play
Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P · 2019
Later among the works it cites.
Policy continuation with hindsight inverse dynamics
Sun, H., Li, Z., Liu, X., Lin, D., and Zhou, B · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
Wang, T. and Ba, J · 2019
Later among the works it cites.
Positive-unlabeled reward learning
Xu, D. and Denil, M · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Successor uncertainties: exploration and uncertainty in temporal difference learning
Janz, D., Hron, J., Mazur, P., Hofmann, K., Hernández-Lobato, J. M., and Tschiatschek, S · 2018
Cited alongside, same era.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., and Levine, S · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
Nair, A., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S · 2018
Cited alongside, same era.
Self-imitation learning
Oh, J., Guo, Y., Singh, S., and Lee, H · 2018
Cited alongside, same era.
Temporal difference models: Model-free deep rl for model-based control
Pong, V., Gu, S., Dalal, M., and Levine, S · 2018
Cited alongside, same era.
Goal-conditioned imitation learning
Ding, Y., Florensa, C., Phielipp, M., and Abbeel, P · 2019
Cited alongside, same era.
Search on the replay buffer: Bridging planning and reinforcement learning
Eysenbach, B., Salakhutdinov, R., and Levine, S · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Later among the works it cites.
Rewriting history with inverse RL: hindsight inference for policy improvement
Eysenbach, B., Geng, X., Levine, S., and Salakhutdinov, R. R · 2020
Later among the works it cites.
Distributed reinforcement learning of targeted grasping with active vision for mobile manipulators
Fujita, Y., Uenishi, K., Ummadisingu, A., Nagarajan, P., Masuda, S., and Castro, M. Y · 2020
Later among the works it cites.
Broadly-exploring, local-policy trees for long-horizon task planning
Ichter, B., Sermanet, P., and Lynch, C · 2020
Later among the works it cites.
Sim2real predictivity: Does evaluation in simulation predict real-world performance?
Kadian, A., Truong, J., Gokaslan, A., Clegg, A., Wijmans, E., Lee, S., Savva, M., Chernova, S., and Batra, D · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Deep dynamics models for learning dexterous manipulation
Nagabandi, A., Konolige, K., Levine, S., and Kumar, V · 2020
Later among the works it cites.
Schroecker, Y. and Isbell, C · 2020
Later among the works it cites.
Cog: Connecting new skills to past experience with offline reinforcement learning
Singh, A., Yu, A., Yang, J., Zhang, J., Kumar, A., and Levine, S · 2020
Later among the works it cites.
Offline learning from demonstrations and unlabeled experience
Zolna, K., Novikov, A., Konyushkova, K., Gulcehre, C., Wang, Z., Aytar, Y., Denil, M., de Freitas, N., and Reed, S · 2020
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
Kalashnikov, D., Varley, J., Chebotar, Y., Swanson, B., Jonschkowski, R., Finn, C., Levine, S., and Hausman, K · 2021
Closest in time.