Fetching the paper…
Reading the bibliography…
Learning representations for pixel-based control has garnered significant attention recently in reinforcement learning.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2004
Earlier work this paper cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Laskin, M., Srinivas, A., and Abbeel, P · 2004
Earlier work this paper cites.
Learning Invariant Representations for Reinforcement Learning without Reconstruction
Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S · 2006
Earlier work this paper cites.
Predictive information accelerates learning in rl
Lee, K.-H., Fischer, I., Liu, A., Guo, Y., Lee, H., Canny, J., and Guadarrama, S · 2007
Earlier work this paper cites.
Learning Robust State Abstractions for Hidden-Parameter Block MDPs
Zhang, A., Sodhani, S., Khetarpal, K., and Pineau, J · 2007
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, J · 2010
Earlier work this paper cites.
Bisimulation metrics for continuous markov decision processes
Ferns, N., Panangaden, P., and Precup, D · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R., and Singh, S · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
From pixels to torques: Policy learning with deep dynamical models
Wahlström, N., Schön, T. B., and Deisenroth, M. P · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J. T., Boedecker, J., and Riedmiller, M · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Pac reinforcement learning with rich observations
Krishnamurthy, A., Agarwal, A., and Langford, J · 2016
Earlier work this paper cites.
Value-aware loss function for model-based reinforcement learning
Farahmand, A.-m., Barreto, A., and Nikovski, D · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Ha, D. and Schmidhuber, J · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Cited alongside, same era.
Learning discrete state abstractions with deep variational inference
Biza, O., Platt, R., van de Meent, J.-W., and Wong, L. L · 2020
Later among the works it cites.
Scalable methods for computing state similarity in deterministic markov decision processes
Castro, P. S · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Later among the works it cites.
Ecological reinforcement learning
Co-Reyes, J. D., Sanjeev, S., Berseth, G., Gupta, A., and Levine, S · 2020
Later among the works it cites.
Bootstrap your own latent: A new approach to self-supervised learning
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. A., Guo, Z. D., Azar, M. G., et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Cited alongside, same era.
Natural environment benchmarks for reinforcement learning
Zhang, A., Wu, Y., and Pineau, J · 2018
Cited alongside, same era.
Provably efficient rl with rich observations via latent state decoding
Du, S., Krishnamurthy, A., Jiang, N., Agarwal, A., Dudik, M., and Langford, J · 2019
Cited alongside, same era.
Implementation matters in deep rl: A case study on ppo and trpo
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2019
Cited alongside, same era.
DeepMDP: Learning continuous latent space models for representation learning
Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2019
Cited alongside, same era.
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2020
Later among the works it cites.
Kielak, K · 2020
Later among the works it cites.
Reinforcement learning with augmented data
Laskin, M., Lee, K., Stooke, A., Pinto, L., Abbeel, P., and Srinivas, A · 2020
Later among the works it cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Lee, A. X., Nagabandi, A., Abbeel, P., and Levine, S · 2020
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2020
Later among the works it cites.
Castro, P. S., Kastner, T., Panangaden, P., and Rowland, M · 2021
Closest in time.
Exploring simple siamese representation learning
Chen, X. and He, K · 2021
Closest in time.
Learning task informed abstractions
Fu, X., Yang, G., Agrawal, P., and Jaakkola, T · 2021
Closest in time.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T. P., Norouzi, M., and Ba, J · 2021
Closest in time.
Temporal predictive coding for model-based planning in latent space
Nguyen, T., Shu, R., Pham, T., Bui, H., and Ermon, S · 2021
Closest in time.
The distracting control suite–a challenging benchmark for reinforcement learning from pixels
Stone, A., Ramirez, O., Konolige, K., and Jonschkowski, R · 2021
Closest in time.
Decoupling representation learning from reinforcement learning
Stooke, A., Lee, K., Abbeel, P., and Laskin, M · 2021
Closest in time.
Model-invariant state abstractions for model-based reinforcement learning
Tomar, M., Zhang, A., Calandra, R., Taylor, M. E., and Pineau, J · 2021
Closest in time.
Barlow twins: Self-supervised learning via redundancy reduction
Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S · 2021
Closest in time.