Fetching the paper…
Reading the bibliography…
Inspired by the diversity and depth of XLand and the simplicity and minimalism of MiniGrid, we present XLand-MiniGrid, a suite of tools and grid-world environments for meta-reinforcement learning research.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Towards generalization and simplicity in continuous control
Rajeswaran, A., Lowrey, K., Todorov, E. V., and Kakade, S. M · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Distributed deep reinforcement learning: Learn how to play atari games in 21 minutes
Adamski, I., Adamski, R., Grel, T., Jędrych, A., Kaczmarek, K., and Michalewski, H · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Earlier work this paper cites.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., Van Hasselt, H., and Silver, D · 2018
Earlier work this paper cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2018
Earlier work this paper cites.
Some considerations on learning to explore via meta-reinforcement learning
Stadie, B. C., Yang, G., Houthooft, R., Chen, X., Duan, Y., Wu, Y., Abbeel, P., and Sutskever, I · 2018
Earlier work this paper cites.
Accelerated methods for deep reinforcement learning
Stooke, A. and Abbeel, P · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
A dissection of overfitting and generalization in continuous reinforcement learning
Zhang, A., Ballas, N., and Pineau, J · 2018
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Dębiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Earlier work this paper cites.
dm_env: A python interface for reinforcement learning environments, 2019
Muldal, A., Doron, Y., Aslanides, J., Harley, T., Ward, T., and Liu, S · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Triton: An intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H. T., and Cox, D · 2019
Earlier work this paper cites.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2019
Earlier work this paper cites.
Varibad: A very good method for bayes-adaptive deep rl via meta-learning
Zintgraf, L., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S · 2019
Earlier work this paper cites.
A brief look at generalization in visual meta-reinforcement learning
Alver, S. and Precup, D · 2020
Cited alongside, same era.
What matters in on-policy reinforcement learning? a large-scale empirical study
Andrychowicz, M., Raichuk, A., Stańczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., et al · 2020
Cited alongside, same era.
Griddly: A platform for ai research in games, 2020
Bamford, C., Huang, S., and Lucas, S · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Sample factory: Egocentric 3d control from pixels at 100000 fps with asynchronous reinforcement learning
Petrenko, A., Huang, Z., Kumar, T., Sukhatme, G., and Koltun, V · 2020
Envpool: A highly parallel reinforcement learning environment execution engine
Weng, J., Lin, M., Huang, S., Liu, B., Makoviichuk, D., Makoviychuk, V., Liu, Z., Song, Y., Luo, T., Jiang, Y., et al · 2022
Later among the works it cites.
Iglu gridworld: Simple and fast environment for embodied dialog agents
Zholus, A., Skrynnik, A., Mohanty, S., Volovikova, Z., Kiseleva, J., Szlam, A., Coté, M.-A., and Panov, A. I · 2022
Later among the works it cites.
Jumanji: a diverse suite of scalable reinforcement learning environments in jax
Bonnet, C., Luo, D., Byrne, D., Surana, S., Coyette, V., Duckworth, P., Midgley, L. I., Kalloniatis, T., Abramowitz, S., Waters, C. N., et al · 2023
Closest in time.
Emergence of collective open-ended exploration from decentralized meta-reinforcement learning
Bornemann, R., Hamon, G., Nisioti, E., and Moulin-Frier, C · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
Bebold: Exploration beyond the boundary of explored regions
Zhang, T., Xu, H., Wang, X., Wu, Y., Keutzer, K., Gonzalez, J. E., and Tian, Y · 2020
Cited alongside, same era.
Brax–a differentiable physics engine for large scale rigid body simulation
Freeman, C. D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O · 2021
Cited alongside, same era.
The minerl 2020 competition on sample efficient reinforcement learning using human priors
Guss, W. H., Castro, M. Y., Devlin, S., Houghton, B., Kuno, N. S., Loomis, C., Milani, S., Mohanty, S., Nakata, K., Salakhutdinov, R., et al · 2021
Cited alongside, same era.
Benchmarking the spectrum of agent capabilities
Hafner, D · 2021
Cited alongside, same era.
Grounding language to entities and dynamics for generalization in reinforcement learning
Hanjie, A. W., Zhong, V. Y., and Narasimhan, K · 2021
Cited alongside, same era.
Podracer architectures for scalable reinforcement learning
Hessel, M., Kroiss, M., Clark, A., Kemaev, I., Quan, J., Keck, T., Viola, F., and van Hasselt, H · 2021
Cited alongside, same era.
Chevalier-Boisvert, M., Dai, B., Towers, M., de Lazcano, R., Willems, L., Lahlou, S., Pal, S., Castro, P. S., and Terry, J · 2023
Closest in time.
Jax-lob: A gpu-accelerated limit order book simulator to unlock large scale reinforcement learning for trading
Frey, S. Y., Li, K., Nagy, P., Sapora, S., Lu, C., Zohren, S., Foerster, J., and Calinescu, A · 2023
Closest in time.
minimax: Efficient baselines for autocurricula in jax
Jiang, M., Dennis, M., Grefenstette, E., and Rocktäschel, T · 2023
Closest in time.
Towards general-purpose in-context learning agents
Kirsch, L., Harrison, J., Freeman, C., Sohl-Dickstein, J., and Schmidhuber, J · 2023
Closest in time.
Pgx: Hardware-accelerated parallel game simulation for reinforcement learning
Koyamada, S., Okano, S., Nishimori, S., Murata, Y., Habara, K., Kita, H., and Ishii, S · 2023
Closest in time.
Katakomba: Tools and benchmarks for data-driven nethack
Kurenkov, V., Nikulin, A., Tarasov, D., and Kolesnikov, S · 2023
Closest in time.
Gigastep - one billion steps per second multi-agent reinforcement learning
Lechner, M., Yin, L., Seyde, T., Wang, T.-H., Xiao, W., Hasani, R., Rountree, J., and Rus, D · 2023
Closest in time.
Supervised pretraining can learn in-context reinforcement learning
Lee, J. N., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., and Brunskill, E · 2023
Closest in time.
Parallel q q -learning: Scaling off-policy reinforcement learning under massively parallel simulation
Li, Z., Chen, T., Hong, Z.-W., Ajay, A., and Agrawal, P · 2023
Closest in time.
Learning to model the world with language
Lin, J., Du, Y., Watkins, O., Hafner, D., Abbeel, P., Klein, D., and Dragan, A · 2023
Closest in time.
Structured state space models for in-context reinforcement learning
Lu, C., Schroecker, Y., Gu, A., Parisotto, E., Foerster, J., Singh, S., and Behbahani, F · 2023
Closest in time.
First-explore, then exploit: Meta-learning intelligent exploration
Norman, B. and Clune, J · 2023
Closest in time.
Jaxmarl: Multi-agent rl environments in jax
Rutherford, A., Ellis, B., Gallici, M., Cook, J., Lupu, A., Ingvarsson, G., Willi, T., Khan, A., de Witt, C. S., Souly, A., et al · 2023
Closest in time.
An extensible, data-oriented architecture for high-performance, many-world simulation
Shacklett, B., Rosenzweig, L. G., Xie, Z., Sarkar, B., Szot, A., Wijmans, E., Koltun, V., Batra, D., and Fatahalian, K · 2023
Closest in time.
In-context reinforcement learning for variable action spaces
Sinii, V., Nikulin, A., Kurenkov, V., Zisman, I., and Kolesnikov, S · 2023
Closest in time.
Revisiting the minimalist approach to offline reinforcement learning
Tarasov, D., Kurenkov, V., Nikulin, A., and Kolesnikov, S · 2023
Closest in time.
Human-timescale adaptation in an open-ended task space
Team, A. A., Bauer, J., Baumli, K., Baveja, S., Behbahani, F., Bhoopchand, A., Bradley-Schmieg, N., Chang, M., Clay, N., Collister, A., et al · 2023
Closest in time.
Emergence of in-context reinforcement learning from noise distillation
Zisman, I., Kurenkov, V., Nikulin, A., Sinii, V., and Kolesnikov, S · 2023
Closest in time.
Craftax: A lightning-fast benchmark for open-ended reinforcement learning
Matthews, M., Beukman, M., Ellis, B., Samvelyan, M., Jackson, M., Coward, S., and Foerster, J · 2024
Closest in time.