Fetching the paper…
Reading the bibliography…
Generalization has been a long-standing challenge for reinforcement learning (RL).
Mixmatch: A holistic approach to semi-supervised learning
Berthelot, D., Carlini, N., Goodfellow, I., Papernot, N., Oliver, A., and Raffel, C · 1905
Earlier work this paper cites.
Remixmatch: Semi-supervised learning with distribution alignment and augmentation anchoring
Berthelot, D., Carlini, N., Cubuk, E. D., Kurakin, A., Sohn, K., Zhang, H., and Raffel, C · 1911
Earlier work this paper cites.
Learning from demonstration
Schaal, S. et al · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S · 2006
Earlier work this paper cites.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S · 2006
Earlier work this paper cites.
Visualizing data using t-sne
Maaten, L. v. d. and Hinton, G · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A · 2013
Earlier work this paper cites.
Bisimulation metrics are optimal value functions
Ferns, N. and Precup, D · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Rusu, A. A., Colmenarejo, S. G., Gulcehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R · 2015
Earlier work this paper cites.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., Schrittwieser, J., Anderson, K., York, S., Cant, M., Cain, A., Bolton, A., Gaffney, S., King, H., Hassabis, D., Legg, S., and Petersen, S · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Earlier work this paper cites.
Carla: An open urban driving simulator
Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Asymmetric actor critic for image-based robot learning
Pinto, L., Andrychowicz, M., Welinder, P., Zaremba, W., and Abbeel, P · 2017
Earlier work this paper cites.
Towards generalization and simplicity in continuous control
Rajeswaran, A., Lowrey, K., Todorov, E., and Kakade, S. M · 2017
Earlier work this paper cites.
Distral: Robust multitask reinforcement learning
Teh, Y., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R · 2017
Cited alongside, same era.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R. H., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2017
Cited alongside, same era.
Multi-task learning for continuous control
Arora, H., Kumar, R., Krone, J., and Li, C · 2018
Cited alongside, same era.
Playing hard exploration games by watching youtube
Aytar, Y., Pfaff, T., Budden, D., Paine, T., Wang, Z., and de Freitas, N · 2018
Cited alongside, same era.
Single episode policy transfer in reinforcement learning, 2019
Yang, J., Petersen, B., Zha, H., and Faissol, D · 2019
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., and Fergus, R · 2019
Later among the works it cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y · 2019
Later among the works it cites.
Learning dexterous in-hand manipulation
Andrychowicz, O. M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al · 2020
Later among the works it cites.
Learning by cheating
Chen, D., Zhou, B., Koltun, V., and Krähenbühl, P · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2018
Cited alongside, same era.
Generalization and regularization in dqn
Farebrother, J., Machado, M. C., and Bowling, M. H · 2018
Cited alongside, same era.
Visualizing and understanding atari agents
Greydanus, S., Koul, A., Dodge, J., and Fern, A · 2018
Cited alongside, same era.
Robot learning in homes: Improving generalization and reducing dataset bias
Gupta, A., Murali, A., Gandhi, D. P., and Pinto, L · 2018
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and surface variations
Hendrycks, D. and Dietterich, T. G · 2018
Cited alongside, same era.
Illuminating generalization in deep reinforcement learning through procedural level generation
Justesen, N., Torrado, R. R., Bontrager, P., Khalifa, A., Togelius, J., and Risi, S · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M. J., and Bowling, M. H · 2018
Cited alongside, same era.
Hansen, N. and Wang, X · 2020
Later among the works it cites.
Self-supervised policy adaptation during deployment
Hansen, N., Sun, Y., Abbeel, P., Efros, A. A., Pinto, L., and Wang, X · 2020
Later among the works it cites.
The impact of non-stationarity on generalisation in deep reinforcement learning
Igl, M., Farquhar, G., Luketina, J., Böhmer, W., and Whiteson, S · 2020
Later among the works it cites.
Jiang, M., Grefenstette, E., and Rocktäschel, T · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2020
Later among the works it cites.
Reinforcement learning with augmented data
Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A · 2020
Later among the works it cites.
Ride: Rewarding impact-driven exploration for procedurally-generated environments
Raileanu, R. and Rocktäschel, T · 2020
Later among the works it cites.
Automatic data augmentation for generalization in deep reinforcement learning
Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., and Fergus, R · 2020
Later among the works it cites.
Rl-cyclegan: Reinforcement learning aware simulation-to-real
Rao, K., Harris, C., Irpan, A., Levine, S., Ibarz, J., and Khansari, M · 2020
Later among the works it cites.
Kornia: an open source differentiable computer vision library for pytorch
Riba, E., Mishkin, D., Ponsa, D., Rublee, E., and Bradski, G · 2020
Later among the works it cites.
igibson, a simulation environment for interactive tasks in large realistic scenes
Shen, B., Xia, F., Li, C., Martın-Martın, R., Fan, L., Wang, G., Buch, S., D’Arpino, C., Srivastava, S., Tchapmi, L. P., Vainio, K., Fei-Fei, L., and Savarese, S · 2020
Later among the works it cites.
Fixmatch: Simplifying semi-supervised learning with consistency and confidence
Sohn, K., Berthelot, D., Li, C.-L., Zhang, Z., Carlini, N., Cubuk, E. D., Kurakin, A., Zhang, H., and Raffel, C · 2020
Later among the works it cites.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Srinivas, A., Laskin, M., and Abbeel, P · 2020
Later among the works it cites.
Decoupling representation learning from reinforcement learning
Stooke, A., Lee, K., Abbeel, P., and Laskin, M · 2020
Later among the works it cites.
Improving generalization in reinforcement learning with mixture regularization
Wang, K., Kang, B., Shao, J., and Feng, J · 2020
Later among the works it cites.
Interactive gibson benchmark: A benchmark for interactive navigation in cluttered environments
Xia, F., Shen, W. B., Li, C., Kasimbeg, P., Tchapmi, M. E., Toshev, A., Martín-Martín, R., and Savarese, S · 2020
Later among the works it cites.
Domain adaptation through task distillation
Zhou, B., Kalra, N., and Krähenbühl, P · 2020
Later among the works it cites.
robosuite: A modular simulation framework and benchmark for robot learning
Zhu, Y., Wong, J., Mandlekar, A., and Martín-Martín, R · 2020
Later among the works it cites.