Fetching the paper…
Reading the bibliography…
Self-supervised learning of image representations by predicting future frames is a promising direction but still remains a challenge.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A · 1988
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., and Vincent, P · 2000
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Towards understanding action recognition
Jhuang, H., Gall, J., Zuffi, S., Schmid, C., and Black, M. J · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Doersch, C., Gupta, A., and Efros, A. A · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Srivastava, N., Mansimov, E., and Salakhudinov, R · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos
Wang, X. and Gupta, A · 2015
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Finn, C., Goodfellow, I., and Levine, S · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Misra, I., Zitnick, C. L., and Hebert, M · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Noroozi, M. and Favaro, P · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Pathak, D., Krähenbühl, P., Donahue, J., Darrell, T., and Efros, A. A · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
Vondrick, C., Pirsiavash, H., and Torralba, A · 2016
Earlier work this paper cites.
Colorful image colorization
Zhang, R., Isola, P., and Efros, A. A · 2016
Earlier work this paper cites.
Stochastic variational video prediction
Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al · 2017
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
Pont-Tuset, J., Perazzi, F., Caelles, S., Arbeláez, P., Sorkine-Hornung, A., and Van Gool, L · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Stochastic video generation with a learned prior
Denton, E. and Fergus, R · 2018
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations
Gidaris, S., Singh, P., and Komodakis, N · 2018
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G · 2018
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Cited alongside, same era.
Large scale adversarial representation learning
Donahue, J. and Simonyan, K · 2019
Cited alongside, same era.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Gupta, A., Kumar, V., Lynch, C., Levine, S., and Hausman, K · 2019
Cited alongside, same era.
Self-supervised learning for video correspondence flow
Lai, Z. and Xie, W · 2019
Cited alongside, same era.
Videomoco: Contrastive video representation learning with temporally adversarial examples
Pan, T., Song, Y., Yang, T., Jiang, W., and Liu, W · 2021
Later among the works it cites.
Spatiotemporal contrastive video representation learning
Qian, R., Meng, T., Gong, B., Yang, M.-H., Wang, H., Belongie, S., and Cui, Y · 2021
Later among the works it cites.
Rethinking self-supervised correspondence learning: A video frame-level similarity perspective
Xu, J. and Wang, X · 2021
Later among the works it cites.
Videogpt: Video generation using vq-vae and transformers
Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A · 2021
Later among the works it cites.
Real robot challenge: A robotics competition in the cloud
Bauer, S., Wüthrich, M., Widmaier, F., Buchholz, A., Stark, S., Goyal, A., Steinbrenner, T., Akpo, J., Joshi, S., Berenz, V., et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Joint-task self-supervised learning for temporal correspondence
Li, X., Liu, S., De Mello, S., Wang, X., Kautz, J., and Yang, M.-H · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Learning correspondence from the cycle-consistency of time
Wang, X., Jabri, A., and Efros, A. A · 2019
Cited alongside, same era.
Self-supervised spatiotemporal learning via video clip order prediction
Xu, D., Xiao, J., Zhao, Z., Shao, J., Xie, D., and Zhuang, Y · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V · 2019
Cited alongside, same era.
Speednet: Learning the speediness in videos
Benaim, S., Ephrat, A., Lang, O., Mosseri, I., Freeman, W. T., Rubinstein, M., Irani, M., and Dekel, T · 2020
Cited alongside, same era.
Masked autoencoders as spatiotemporal learners
Feichtenhofer, C., Li, Y., He, K., et al · 2022
Later among the works it cites.
Cross-architecture self-supervised video representation learning
Guo, S., Xiong, Z., Zhong, Y., Wang, L., Guo, X., Han, B., and Huang, W · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Later among the works it cites.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al · 2022
Later among the works it cites.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
James, S. and Davison, A. J · 2022
Later among the works it cites.
Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation
James, S., Wada, K., Laidlow, T., and Davison, A. J · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation
Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., and Wang, L · 2022
Later among the works it cites.
Phenaki: Variable length video generation from open domain textual description
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D · 2022
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H · 2022
Later among the works it cites.
Adaptive temporal encoding network for video instance-level human parsing
Zhou, Q., Liang, X., Gong, K., and Lin, L · 2022
Later among the works it cites.
Siamese masked autoencoders
Gupta, A., Wu, J., Deng, J., and Fei-Fei, L · 2023
Later among the works it cites.
Soda: Bottleneck diffusion models for representation learning
Hudson, D. A., Zoran, D., Malinowski, M., Lampinen, A. K., Jaegle, A., McClelland, J. L., Matthey, L., Hill, F., and Lerchner, A · 2023
Later among the works it cites.
Mage: Masked generative encoder to unify representation learning and image synthesis
Li, T., Chang, H., Mishra, S. K., Zhang, H., Katabi, D., and Krishnan, D · 2023
Later among the works it cites.
Where are we in the search for an artificial visual cortex for embodied intelligence?
Majumdar, A., Yadav, K., Arnaud, S., Ma, Y. J., Chen, C., Silwal, S., Jain, A., Berges, V.-P., Abbeel, P., Malik, J., et al · 2023
Later among the works it cites.
Harnessing discrete representations for continual reinforcement learning
Meyer, E., White, A., and Machado, M. C · 2023
Later among the works it cites.
Multi-view masked world models for visual robotic manipulation
Seo, Y., Kim, J., James, S., Lee, K., Shin, J., and Abbeel, P · 2023
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Video probabilistic diffusion models in projected latent space
Yu, S., Sohn, K., Kim, S., and Shin, J · 2023
Later among the works it cites.