Fetching the paper…
Reading the bibliography…
Autonomous agents trained using deep reinforcement learning (RL) often lack the ability to successfully generalise to new environments, even when these environments share characteristics with the ones they have encountered during training.
Evolutionary robotics and the radical envelope-of-noise hypothesis
Jakobi, N · 1997
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents (extended abstract)
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2015
Earlier work this paper cites.
Wave function collapse algorithm
Gumin, M · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
White, T · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Crow, D., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Earlier work this paper cites.
Information-theoretic analysis of generalization capability of learning algorithms
Xu, A. and Raginsky, M · 2017
Cited alongside, same era.
Minimalistic gridworld environment for openai gym
Chevalier-Boisvert, M., Willems, L., and Pal, S · 2018
Cited alongside, same era.
Assessing generalization in deep reinforcement learning
Packer, C., Gao, K., Kos, J., Krähenbühl, P., Koltun, V., and Song, D. X · 2018
Cited alongside, same era.
A study on overfitting in deep reinforcement learning
Zhang, C., Vinyals, O., Munos, R., and Bengio, S · 2018
Cited alongside, same era.
Solving the rubik’s cube with deep reinforcement learning and search
Agostinelli, F., McAleer, S., Shmakov, A., and Baldi, P · 2019
Cited alongside, same era.
Why generalization in RL is difficult: Epistemic pomdps and implicit partial observability
Ghosh, D., Rahme, J., Kumar, A., Zhang, A., Adams, R. P., and Levine, S · 2021
Later among the works it cites.
Replay-guided adversarial environment design
Jiang, M., Dennis, M., Parker-Holder, J., Foerster, J. N., Grefenstette, E., and Rocktäschel, T · 2021
Later among the works it cites.
Prioritized level replay
Jiang, M., Grefenstette, E., and Rocktäschel, T · 2021
Later among the works it cites.
Openrooms: An open framework for photorealistic indoor scene datasets
Li, Z., Yu, T., Sang, S., Wang, S., Song, M., Liu, Y., Yeh, Y., Zhu, R., Gundavarapu, N. B., Shi, J., Bi, S., Yu, H., Xu, Z., Sunkavalli, K., Hasan, M., Ramamoorthi, R., and Chandraker, M · 2021
Later among the works it cites.
Automatic data augmentation for generalization in reinforcement learning
Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., and Fergus, R · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to optimize in swarms
Cao, Y., Chen, T., Wang, Z., and Shen, Y · 2019
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2019
Cited alongside, same era.
Generalization in reinforcement learning with selective noise injection and information bottleneck
Igl, M., Ciosek, K., Li, Y., Tschiatschek, S., Zhang, C., Devlin, S., and Hofmann, K · 2019
Cited alongside, same era.
How powerful are graph neural networks?
Xu, K., Hu, W., Leskovec, J., and Jegelka, S · 2019
Cited alongside, same era.
Instance-based generalization in reinforcement learning
Bertrán, M., Martínez, N., Phielipp, M., and Sapiro, G · 2020
Cited alongside, same era.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Cited alongside, same era.
Emergent complexity and zero-shot transfer via unsupervised environment design
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A. M., Russell, S., Critch, A., and Levine, S · 2020
Cited alongside, same era.
Wilson, B., Qi, W., Agarwal, T., Lambert, J., Singh, J., Khandelwal, S., Pan, B., Kumar, R., Hartnett, A., Kaesemodel Pontes, J., Ramanan, D., Carr, P., and Hays, J · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, D., Kostrikov, I., and Fergus, R · 2021
Later among the works it cites.
Grounding aleatoric uncertainty for unsupervised environment design
Jiang, M., Dennis, M., Parker-Holder, J., Lupu, A., Küttler, H., Grefenstette, E., Rocktäschel, T., and Foerster, J · 2022
Later among the works it cites.
Evolving curricula with regret-based environment design
Parker-Holder, J., Jiang, M., Dennis, M., Samvelyan, M., Foerster, J. N., Grefenstette, E., and Rocktäschel, T · 2022
Later among the works it cites.
Learning to walk in minutes using massively parallel deep reinforcement learning
Rudin, N., Hoeller, D., Reist, P., and Hutter, M · 2022
Later among the works it cites.
Clutr: Curriculum learning via unsupervised task representation learning
Azad, A. S., Gur, I., Emhoff, J., Alexis, N., Faust, A., Abbeel, P., and Stoica, I · 2023
Later among the works it cites.
A survey of zero-shot generalisation in deep reinforcement learning
Kirk, R., Zhang, A., Grefenstette, E., and Rocktäschel, T · 2023
Later among the works it cites.
Metabox: A benchmark platform for meta-black-box optimization with reinforcement learning
Ma, Z., Guo, H., Chen, J., Li, Z., Peng, G., Gong, Y.-J., Ma, Y., and Cao, Z · 2023
Later among the works it cites.
Multi-view disentanglement for reinforcement learning with multiple cameras
Dunion, M. and Albrecht, S. V · 2024
Closest in time.
Conditional mutual information for disentangled representations in reinforcement learning
Dunion, M., McInroe, T., Luck, K. S., Hanna, J., and Albrecht, S · 2024
Closest in time.