Fetching the paper…
Reading the bibliography…
Generating representations of video data is of key importance in advancing the field of machine perception.
Learning video representations using contrastive bidirectional transformer
Sun, C., Baradel, F., Murphy, K., and Schmid, C · 1906
Earlier work this paper cites.
Multi-modal self-supervision from generalized data transformations
Patrick, M., Asano, Y. M., Kuznetsova, P., Fong, R., Henriques, J. F., Zweig, G., and Vedaldi, A · 2003
Earlier work this paper cites.
Audio-visual instance discrimination with cross-modal agreement
Morgado, P., Vasconcelos, N., and Misra, I · 2004
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Earlier work this paper cites.
Support-set bottlenecks for video-text representation learning
Patrick, M., Huang, P.-Y., Asano, Y., Metze, F., Hauptmann, A., Henriques, J., and Vedaldi, A · 2010
Earlier work this paper cites.
Self-supervised spatiotemporal feature learning via video rotation prediction
Jing, L., Yang, X., Liu, J., and Tian, Y · 2018
Earlier work this paper cites.
Cooperative learning of audio and video models from self-supervised synchronization
Korbar, B., Tran, D., and Torresani, L · 2018
Earlier work this paper cites.
Self-supervised multimodal versatile networks
Alayrac, J.-B., Recasens, A., Schneider, R., Arandjelović, R., Ramapuram, J., De Fauw, J., Smaira, L., Dieleman, S., and Zisserman, A · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Learning spatiotemporal features via video and text pair discrimination
Li, T. and Wang, L · 2020
Cited alongside, same era.
End-to-end learning of visual representations from uncurated instructional videos
Miech, A., Alayrac, J.-B., Smaira, L., Laptev, I., Sivic, J., and Zisserman, A · 2020
Cited alongside, same era.
Evolving losses for unsupervised video representation learning
Piergiovanni, A., Angelova, A., and Ryoo, M. S · 2020
Cited alongside, same era.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Akbari, H., Yuan, L., Qian, R., Chuang, W.-H., Chang, S.-F., Cui, Y., and Gong, B · 2021
Later among the works it cites.
Vivit: A video vision transformer
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., and Schmid, C · 2021
Later among the works it cites.
Is space-time attention all you need for video understanding
Bertasius, G., Wang, H., and Torresani, L · 2021
Later among the works it cites.
Multimodal clustering networks for self-supervised learning from unlabeled videos
Chen, B., Rouditchenko, A., Duarte, K., Kuehne, H., Thomas, S., Boggust, A., Panda, R., Kingsbury, B., Feris, R., Harwath, D., et al · 2021
Later among the works it cites.
Space-time crop & attend: Improving cross-modal video representation learning
Patrick, M., Huang, P.-Y., Misra, I., Metze, F., Vedaldi, A., Asano, Y. M., and Henriques, J. F · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning video representations from textual web supervision
Stroud, J. C., Lu, Z., Sun, C., Deng, J., Sukthankar, R., Schmid, C., and Ross, D. A · 2020
Cited alongside, same era.
Audiovisual slowfast networks for video recognition
Xiao, F., Lee, Y. J., Grauman, K., Malik, J., and Feichtenhofer, C · 2020
Cited alongside, same era.
Memory-augmented dense predictive coding for video representation learning
Han, T., Xie, W., and Zisserman, A
Cited in the paper.
Self-supervised co-training for video representation learning
Han, T., Xie, W., and Zisserman, A
Cited in the paper.
Learning representations from audio-visual spatial alignment
Morgado, P., Li, Y., and Nvasconcelos, N
Cited in the paper.
Videobert: A joint model for video and language representation learning
Sun, C., Myers, A., Vondrick, C., Murphy, K., and Schmid, C
Cited in the paper.
Spatiotemporal contrastive video representation learning
Qian, R., Meng, T., Gong, B., Yang, M.-H., Wang, H., Belongie, S., and Cui, Y · 2021
Later among the works it cites.
Broaden your views for self-supervised video learning
Recasens, A., Luc, P., Alayrac, J.-B., Wang, L., Strub, F., Tallec, C., Malinowski, M., Pătrăucean, V., Altché, F., Valko, M., et al · 2021
Later among the works it cites.