Fetching the paper…
Reading the bibliography…
Unsupervised visual representation learning offers the opportunity to leverage large corpora of unlabeled trajectories to form useful visual representations, which can benefit the training of reinforcement learning (RL) algorithms.
Supervise thyself: Examining self-supervised representations in interactive environments
Racah, E. and Pal, C. (2019) · 1906
Earlier work this paper cites.
Cyanure: An Open-Source Toolbox for Empirical Risk Minimization for Python, C++, and soon more
Mairal, J. (2019) · 1912
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Chen, X., Fan, H., Girshick, R., and He, K. (2020b) · 2003
Earlier work this paper cites.
Masked contrastive representation learning for reinforcement learning
Zhu, J., Xia, Y., Wu, L., Deng, J., gang Zhou, W., Qin, T., and Li, H. (2020) · 2010
Earlier work this paper cites.
Online bag-of-visual-words generation for unsupervised representation learning
Gidaris, S., Bursuc, A., Puy, G., Komodakis, N., Cord, M., and Pérez, P. (2020) · 2012
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M. (2012) · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A. (2013) · 2013
Earlier work this paper cites.
Incremental majorization-minimization optimization with application to large-scale machine learning
Mairal, J. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Earlier work this paper cites.
Deconvolution and checkerboard artifacts
Odena, A., Dumoulin, V., and Olah, C. (2016) · 2016
Earlier work this paper cites.
Control of memory, active perception, and action in minecraft
Oh, J., Chockalingam, V., Singh, S. P., and Lee, H. (2016) · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Lin, T., Goyal, P., Girshick, R. B., He, K., and Dollár, P. (2017) · 2017
Earlier work this paper cites.
Playing hard exploration games by watching youtube
Aytar, Y., Pfaff, T., Budden, D., Paine, T. L., Wang, Z., and de Freitas, N. (2018) · 2018
Earlier work this paper cites.
Neural predictive belief representations
Guo, Z. D., Azar, M. G., Piot, B., Pires, B. A., and Munos, R. (2018) · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D. (2018) · 2018
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M. G., and Silver, D. (2018) · 2018
Earlier work this paper cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M. (2018) · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., de Vries, H., Dumoulin, V., and Courville, A. C. (2018) · 2018
Earlier work this paper cites.
Forward modeling for partial observation strategy games - A starcraft defogger
Synnaeve, G., Lin, Z., Gehring, J., Gant, D., Mella, V., Khalidov, V., Carion, N., and Usunier, N. (2018) · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al. (2018) · 2018
Earlier work this paper cites.
Unsupervised state representation learning in atari
Anand, A., Racah, E., Ozair, S., Bengio, Y., Côté, M., and Hjelm, R. D. (2019) · 2019
Earlier work this paper cites.
Large-scale study of curiosity-driven learning
Burda, Y., Edwards, H., Pathak, D., Storkey, A. J., Darrell, T., and Efros, A. A. (2019a) · 2019
Earlier work this paper cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A. J., and Klimov, O. (2019b) · 2019
Cited alongside, same era.
Deepmdp: Learning continuous latent space models for representation learning
Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G. (2019) · 2019
Cited alongside, same era.
An inexact variable metric proximal point algorithm for generic quasi-newton acceleration
Lin, H., Mairal, J., and Harchaoui, Z. (2019) · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M. (2020) · 2020
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. (2020) · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Reinforcement learning with prototypical representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. (2021a) · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, D., Kostrikov, I., and Fergus, R. (2021b) · 2021
Later among the works it cites.
Mastering atari games with limited data
Ye, W., Liu, S., Kurutach, T., Abbeel, P., and Gao, Y. (2021) · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S. (2021) · 2021
Later among the works it cites.
Coberl: Contrastive BERT for reinforcement learning
Banino, A., Badia, A. P., Walker, J. C., Scholtes, T., Mitrovic, J., and Blundell, C. (2022) · 2022
Closest in time.
Vicreg: Variance-invariance-covariance regularization for self-supervised learning
Bardes, A., Ponce, J., and LeCun, Y. (2022) · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. E. (2020a) · 2020
Cited alongside, same era.
Probing emergent semantics in predictive agents via question answering
Das, A., Carnevale, F., Merzic, H., Rimell, L., Schneider, R., Abramson, J., Hung, A., Ahuja, A., Clark, S., Wayne, G., and Hill, F. (2020) · 2020
Cited alongside, same era.
Bootstrap your own latent - A new approach to self-supervised learning
Grill, J., Strub, F., Altché, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. Á., Guo, Z., Azar, M. G., Piot, B., Kavukcuoglu, K., Munos, R., and Valko, M. (2020) · 2020
Cited alongside, same era.
Model based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Mohiuddin, A., Sepassi, R., Tucker, G., and Michalewski, H. (2020) · 2020
Cited alongside, same era.
CURL: contrastive unsupervised representations for reinforcement learning
Laskin, M., Srinivas, A., and Abbeel, P. (2020) · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al. (2020) · 2020
Cited alongside, same era.
Planning to explore via self-supervised world models
Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., and Pathak, D. (2020) · 2020
Cited alongside, same era.
Closest in time.
Self-supervised learning with random-projection quantizer for speech recognition
Chiu, C., Qin, J., Zhang, Y., Yu, J., and Wu, Y. (2022) · 2022
Closest in time.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
D’Oro, P., Schwarzer, M., Nikishin, E., Bacon, P.-L., Bellemare, M. G., and Courville, A. (2022) · 2022
Closest in time.
Garrido, Q., Balestriero, R., Najman, L., and Lecun, Y. (2022) · 2022
Closest in time.
Temporal difference learning for model predictive control
Hansen, N., Su, H., and Wang, X. (2022) · 2022
Closest in time.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Ma, Y. J., Sodhani, S., Jayaraman, D., Bastani, O., Kumar, V., and Zhang, A. (2022) · 2022
Closest in time.
Transformers are sample efficient world models
Micheli, V., Alonso, E., and Fleuret, F. (2022) · 2022
Closest in time.
R3m: A universal visual representation for robot manipulation
Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A. (2022) · 2022
Closest in time.
Reinforcement learning with action-free pre-training from videos
Seo, Y., Lee, K., James, S. L., and Abbeel, P. (2022) · 2022
Closest in time.
Mask-based latent reconstruction for reinforcement learning
Yu, T., Zhang, Z., Lan, C., Chen, Z., and Lu, Y. (2022) · 2022
Closest in time.
Light-weight probing of unsupervised representations for reinforcement learning
Zhang, W., GX-Chen, A., Sobal, V., LeCun, Y., and Carion, N. (2022) · 2022
Closest in time.
Self-supervised learning from images with a joint-embedding predictive architecture
Assran, M., Duval, Q., Misra, I., Bojanowski, P., Vincent, P., Rabbat, M., LeCun, Y., and Ballas, N. (2023) · 2023
Closest in time.
Reinforcement learning from passive data via latent intentions
Ghosh, D., Bhateja, C. A., and Levine, S. (2023) · 2023
Closest in time.
Lee, H., Lee, K., Hwang, D., Lee, H., Lee, B., and Choo, J. (2023) · 2023
Closest in time.
Dinov2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P. (2023) · 2023
Closest in time.
Transformer-based world models are happy with 100k interactions
Robine, J., Höftmann, M., Uelwer, T., and Harmeling, S. (2023) · 2023
Closest in time.
Bigger, better, faster: Human-level atari with human-level efficiency
Schwarzer, M., Ceron, J. S. O., Courville, A., Bellemare, M. G., Agarwal, R., and Castro, P. S. (2023) · 2023
Closest in time.
Become a proficient player with limited data through watching pure videos
Ye, W., Zhang, Y., Abbeel, P., and Gao, Y. (2023) · 2023
Closest in time.
Bridging state and history representations: Understanding self-predictive rl
Ni, T., Eysenbach, B., Seyedsalehi, E., Ma, M., Gehring, C., Mahajan, A., and Bacon, P.-L. (2024) · 2024
Closest in time.