Fetching the paper…
Reading the bibliography…
A broad challenge of research on generalization for sequential decision-making tasks in interactive environments is designing benchmarks that clearly landmark progress.
Table for estimating the goodness of fit of empirical distributions
Smirnov, N · 1948
Earlier work this paper cites.
The wasserstein distance and approximation theorems
Rüschendorf, L · 1985
Earlier work this paper cites.
PDDL–the planning domain definition language, 1998
McDermott, D., Ghallab, M., Howe, A., Knoblock, C., Ram, A., Veloso, M., Weld, D., and Wilkins, D · 1998
Earlier work this paper cites.
The ingredients of real-world robotic reinforcement learning
Zhu, H., Yu, J., Gupta, A., Shah, D., Hartikainen, K., Singh, A., Kumar, V., and Levine, S · 2004
Earlier work this paper cites.
Classical electricity and magnetism
Panofsky, W. K. and Phillips, M · 2005
Earlier work this paper cites.
Box2D, 2006
Catto, E · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
robosuite: A modular simulation framework and benchmark for robot learning
Zhu, Y., Wong, J., Mandlekar, A., and Martín-Martín, R · 2009
Earlier work this paper cites.
Relational dynamic influence diagram language (rddl): Language description, 2010
Sanner, S · 2010
Earlier work this paper cites.
Displacement interpolation using lagrangian mass transport
Bonneel, N., Van De Panne, M., Paris, S., and Heidrich, W · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: Applications and real-world challenges
Kormushev, P., Calinon, S., and Caldwell, D. G · 2013
Earlier work this paper cites.
Sliced and radon wasserstein barycenters of measures
Bonneel, N., Rabin, J., Peyré, G., and Pfister, H · 2015
Earlier work this paper cites.
Simulation tools for model-based robotics: Comparison of bullet, havok, mujoco, ode and physx
Erez, T., Tassa, Y., and Todorov, E · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Generalizing skills with semi-supervised reinforcement learning
Finn, C., Yu, T., Fu, J., Abbeel, P., and Levine, S · 2016
Earlier work this paper cites.
The Malmo platform for artificial intelligence experimentation
Johnson, M., Hofmann, K., Hutton, T., and Bignell, D · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Earlier work this paper cites.
Imitation learning: A survey of learning methods
Hussein, A., Gaber, M. M., Elyan, E., and Jayne, C · 2017
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Gordon, D., Zhu, Y., Gupta, A., and Farhadi, A · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. A · 2017
Cited alongside, same era.
Continual learning through synaptic intelligence
Zenke, F., Poole, B., and Ganguli, S · 2017
Cited alongside, same era.
Mutual information neural estimation
Belghazi, M. I., Baratin, A., Rajeshwar, S., Ozair, S., Bengio, Y., Courville, A., and Hjelm, D · 2018
Cited alongside, same era.
Textworld: A learning environment for text-based games
Côté, M.-A., Kádár, A., Yuan, X., Kybartas, B., Barnes, T., Fine, E., Moore, J., Hausknecht, M., Asri, L. E., Adada, M., et al · 2018
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q · 2019
Later among the works it cites.
Causalworld: A robotic manipulation benchmark for causal structure and transfer learning
Ahmed, O., Träuble, F., Goyal, A., Neitz, A., Bengio, Y., Schölkopf, B., Wüthrich, M., and Bauer, S · 2020
Later among the works it cites.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2020
Later among the works it cites.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dietterich, T., Trimponias, G., and Chen, Z · 2018
Cited alongside, same era.
Meta-reinforcement learning of structured exploration strategies
Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Learning deep representations by mutual information estimation and maximization
Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., and Bengio, Y · 2018
Cited alongside, same era.
Self-supervised deep reinforcement learning with generalized computation graphs for robot navigation
Kahn, G., Villaflor, A., Ding, B., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Assessing generalization in deep reinforcement learning
Packer, C., Gao, K., Kos, J., Krähenbühl, P., Koltun, V., and Song, D · 2018
Cited alongside, same era.
An atari model zoo for analyzing, visualizing, and comparing deep reinforcement learning agents
Such, F. P., Madhavan, V., Liu, R., Wang, R., Castro, P. S., Li, Y., Zhi, J., Schubert, L., Bellemare, M. G., Clune, J., et al · 2018
Cited alongside, same era.
Küttler, H., Nardelli, N., Miller, A. H., Raileanu, R., Selvatici, M., Grefenstette, E., and Rocktäschel, T · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Laskin, M., Srinivas, A., and Abbeel, P · 2020
Later among the works it cites.
Deep reinforcement and infomax learning
Mazoure, B., Tachet des Combes, R., Doan, T. L., Bachman, P., and Hjelm, R. D · 2020
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2020
Later among the works it cites.
Pddlgym: Gym environments from pddl problems
Silver, T. and Chitnis, R · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S · 2020
Later among the works it cites.
Contrastive behavioral similarity embeddings for generalization in reinforcement learning
Agarwal, R., Machado, M. C., Castro, P. S., and Bellemare, M. G · 2021
Later among the works it cites.
Minimalistic gridworld environment (MiniGrid), 2021
Chevalier-Boisvert, M · 2021
Later among the works it cites.
Secant: Self-expert cloning for zero-shot generalization of visual policies
Fan, L., Wang, G., Huang, D.-A., Yu, Z., Fei-Fei, L., Zhu, Y., and Anandkumar, A · 2021
Later among the works it cites.
Pot: Python optimal transport
Flamary, R., Courty, N., Gramfort, A., Alaya, M. Z., Boisbunon, A., Chambon, S., Chapel, L., Corenflos, A., Fatras, K., Fournier, N., Gautheron, L., Gayraud, N. T., Janati, H., Rakotomamonjy, A., Redko, I., Rolet, A., Schutz, A., Seguy, V., Sutherland, D. J., Tavenard, R., Tong, A., and Vayer, T · 2021
Later among the works it cites.
Why generalization in rl is difficult: Epistemic pomdps and implicit partial observability
Ghosh, D., Rahme, J., Kumar, A., Zhang, A., Adams, R. P., and Levine, S · 2021
Later among the works it cites.
Benchmarking the spectrum of agent capabilities
Hafner, D · 2021
Later among the works it cites.
How to train your robot with deep reinforcement learning: lessons we have learned
Ibarz, J., Tan, J., Finn, C., Kalakrishnan, M., Pastor, P., and Levine, S · 2021
Later among the works it cites.
Systematic evaluation of causal discovery in visual model based reinforcement learning
Ke, N. R., Didolkar, A., Mittal, S., Goyal, A., Lajoie, G., Bauer, S., Rezende, D., Bengio, Y., Mozer, M., and Pal, C · 2021
Later among the works it cites.
On the effect of auxiliary tasks on representation dynamics
Lyle, C., Rowland, M., Ostrovski, G., and Dabney, W · 2021
Later among the works it cites.
Cross-trajectory representation learning for zero-shot generalization in rl
Mazoure, B., Ahmed, A. M., MacAlpine, P., Hjelm, R. D., and Kolobov, A · 2021
Later among the works it cites.
Pretraining representations for data-efficient reinforcement learning
Schwarzer, M., Rajkumar, N., Noukhovitch, M., Anand, A., Charlin, L., Hjelm, R. D., Bachman, P., and Courville, A. C · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
Team, O. E. L., Stooke, A., Mahajan, A., Barros, C., Deck, C., Bauer, J., Sygnowski, J., Trebacz, M., Jaderberg, M., Mathieu, M., et al · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Later among the works it cites.