Fetching the paper…
Reading the bibliography…
Designing better deep networks and better reinforcement learning (RL) algorithms are both important for deep RL.
Deep recurrent q-learning for partially observable mdps
Hausknecht, M. and Stone, P · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Deep attention recurrent q-network
Sorokin, I., Seleznev, A., Pavlov, M., Fedorov, A., and Ignateva, A · 2015
Earlier work this paper cites.
Ga3c: Gpu-based a3c for deep reinforcement learning
Babaeizadeh, M., Frosio, I., Tyree, S., Clemons, J., and Kautz, J · 2016
Earlier work this paper cites.
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., et al · 2016
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Deepmind ai reduces google data centre cooling bill by 40%
Evans, R. and Gao, J · 2016
Earlier work this paper cites.
Lstm: A search space odyssey
Greff, K., Srivastava, R. K., Koutník, J., Steunebrink, B. R., and Schmidhuber, J · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Value iteration networks
Tamar, A., Wu, Y., Thomas, G., Levine, S., and Abbeel, P · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Van Hasselt, H., Lanctot, M., and De Freitas, N · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Distributional reinforcement learning with quantile regression
Dabney, W., Rowland, M., Bellemare, M., and Munos, R · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J. G., Le, Q., and Salakhutdinov, R · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2019
Cited alongside, same era.
Reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Later among the works it cites.
Dreaming: Model-based reinforcement learning by latent imagination without reconstruction
Okada, M. and Taniguchi, T · 2021
Later among the works it cites.
Training larger networks for deep reinforcement learning
Ota, K., Jha, D. K., and Kanezaki, A · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N · 2021
Later among the works it cites.
Rrl: Resnet as representation for reinforcement learning
Shah, R. M. and Kumar, V · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kenton, J. D. M.-W. C. and Toutanova, L. K · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Can increasing input dimensionality improve deep reinforcement learning?
Ota, K., Oiki, T., Jha, D., Mariyama, T., and Nikovski, D · 2020
Cited alongside, same era.
Stabilizing transformers for reinforcement learning
Parisotto, E., Song, F., Rae, J., Pascanu, R., Gulcehre, C., Jayakumar, S., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., et al · 2020
Cited alongside, same era.
Later among the works it cites.
d3rlpy: An offline deep reinforcement library
Takuma Seno, M. I · 2021
Later among the works it cites.
Coberl: Contrastive bert for reinforcement learning
Banino, A., Badia, A. P., Walker, J. C., Scholtes, T., Mitrovic, J., and Blundell, C · 2022
Closest in time.
Goulão, M. and Oliveira, A. L · 2022
Closest in time.
On transforming reinforcement learning by transformer: The development trajectory
Hu, S., Shen, L., Zhang, Y., Chen, Y., and Tao, D · 2022
Closest in time.
Dualformer: Local-global stratified transformer for efficient video recognition
Liang, Y., Zhou, P., Zimmermann, R., and Yan, S · 2022
Closest in time.
Transformers are meta-reinforcement learners
Melo, L. C · 2022
Closest in time.
Deep reinforcement learning with swin transformer
Meng, L., Goodwin, M., Yazidi, A., and Engelstad, P · 2022
Closest in time.
Transformers are sample efficient world models
Micheli, V., Alonso, E., and Fleuret, F · 2022
Closest in time.
You can’t count on luck: Why decision transformers fail in stochastic environments
Paster, K., McIlraith, S. A., and Ba, J · 2022
Closest in time.
Transformer-based deep reinforcement learning in vizdoom
Sopov, V. and Makarov, I · 2022
Closest in time.
Evaluating vision transformer methods for deep reinforcement learning from pixels
Tao, T., Reda, D., and van de Panne, M · 2022
Closest in time.
Cascade transformers for end-to-end person search
Yu, R., Du, D., LaLonde, R., Davila, D., Funk, C., Hoogs, A., and Clipp, B · 2022
Closest in time.
Zheng, Q., Zhang, A., and Grover, A · 2022
Closest in time.