Fetching the paper…
Reading the bibliography…
Most transformer-based video encoders are limited to short temporal contexts due to their quadratic complexity.
Remembering: A study in experimental and social psychology
Bartlett, F. C · 1932
Earlier work this paper cites.
Possible principles underlying the transformation of sensory messages
Barlow, H. B · 1961
Earlier work this paper cites.
Simple memory: a theory for archicortex
Marr, D · 1971
Earlier work this paper cites.
Memory and consciousness
Tulving, E · 1985
Earlier work this paper cites.
Geometric approximation via coresets
Agarwal, P. K., Har-Peled, S., Varadarajan, K. R., et al · 2005
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
Good semi-supervised learning that requires a bad gan
Dai, Z., Yang, Z., Yang, F., Cohen, W. W., and Salakhutdinov, R. R · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C., Shrivastava, A., Singh, S., and Gupta, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
RESOUND: Towards action recognition without representation bias
Li, Y., Li, Y., and Vasconcelos, N · 2018
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Earlier work this paper cites.
SlowFast networks for video recognition
Feichtenhofer, C., Fan, H., Malik, J., and He, K · 2019
Earlier work this paper cites.
Howto100M: Learning a text-video embedding by watching hundred million narrated video clips
Miech, A., Zhukov, D., Alayrac, J.-B., Tapaswi, M., Laptev, I., and Sivic, J · 2019
Earlier work this paper cites.
Video object segmentation using space-time memory networks
Oh, S. W., Lee, J.-Y., Xu, N., and Kim, S. J · 2019
Earlier work this paper cites.
Long-term feature banks for detailed video understanding
Wu, C.-Y., Feichtenhofer, C., Fan, H., He, K., Krahenbuhl, P., and Girshick, R · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Memory-augmented dense predictive coding for video representation learning
Han, T., Xie, W., and Zisserman, A · 2020
Earlier work this paper cites.
MAST: A memory-augmented self-supervised tracker
Lai, Z., Lu, E., and Xie, W · 2020
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., and Lillicrap, T. P · 2020
Earlier work this paper cites.
Big Bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Earlier work this paper cites.
ViViT: A video vision transformer
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., and Schmid, C · 2021
Cited alongside, same era.
Is space-time attention all you need for video understanding?
Bertasius, G., Wang, H., and Torresani, L · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Cited alongside, same era.
Can an image classifier suffice for action recognition?
Fan, Q., Panda, R., et al · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
PaLI: A jointly-scaled multilingual language-image model
Chen, X., Wang, X., Changpinyo, S., Piergiovanni, A., Padlewski, P., Salz, D., Goodman, S., Grycner, A., Mustafa, B., Beyer, L., et al · 2023
Later among the works it cites.
Bard [large language model], 2023
Google AI · 2023
Later among the works it cites.
Semi-parametric video-grounded text generation
Kim, S., Kim, J.-H., Lee, J., and Seo, M · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L · 2021
Cited alongside, same era.
Next-QA: Next phase of question-answering to explaining temporal actions
Xiao, J., Shang, X., Yao, A., and Chua, T.-S · 2021
Cited alongside, same era.
Flamingo: A visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Hydra attention: Efficient attention with many heads
Bolya, D., Fu, C.-Y., Dai, X., Zhang, P., and Hoffman, J · 2022
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G. B., Lespiau, J.-B., Damoc, B., Clark, A., et al · 2022
Cited alongside, same era.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
Later among the works it cites.
VideoChat: Chat-centric video understanding
Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., and Qiao, Y · 2023
Later among the works it cites.
MM-VID: Advancing video understanding with GPT-4V(ision)
Lin, K., Ahmed, F., Li, L., Lin, C.-C., Azarnasab, E., Yang, Z., Wang, J., Liang, L., Liu, Z., Lu, Y., et al · 2023
Later among the works it cites.
EgoSchema: A diagnostic benchmark for very long-form video language understanding
Mangalam, K., Akshulakov, R., and Malik, J · 2023
Later among the works it cites.
Perception Test: A diagnostic benchmark for multimodal video models
Pătrăucean, V., Smaira, L., Gupta, A., Continente, A. R., Markeeva, L., Banarse, D., Koppula, S., Heyward, J., Malinowski, M., Yang, Y., et al · 2023
Later among the works it cites.
Rethinking video ViTs: Sparse video tubes for joint image and video learning
Piergiovanni, A., Kuo, W., and Angelova, A · 2023
Later among the works it cites.
Hiera: A hierarchical vision transformer without the bells-and-whistles
Ryali, C., Hu, Y.-T., Bolya, D., Wei, C., Fan, H., Huang, P.-Y., Aggarwal, V., Chowdhury, A., Poursaeed, O., Hoffman, J., et al · 2023
Later among the works it cites.
Token Turing machines
Ryoo, M. S., Gopalakrishnan, K., Kahatapitiya, K., Xiao, T., Rao, K., Stone, A., Lu, Y., Ibarz, J., and Arnab, A · 2023
Later among the works it cites.
Contrastive video question answering via video graph transformer
Xiao, J., Zhou, P., Yao, A., Li, Y., Hong, R., Yan, S., and Chua, T.-S · 2023
Later among the works it cites.
AIM: Adapting image models for efficient video action recognition
Yang, T., Zhu, Y., Xie, Y., Zhang, A., Chen, C., and Li, M · 2023
Later among the works it cites.
HiTeA: Hierarchical temporal-aware video-language pre-training
Ye, Q., Xu, G., Yan, M., Xu, H., Qian, Q., Zhang, J., and Huang, F · 2023
Later among the works it cites.
Self-chained image-language model for video localization and question answering
Yu, S., Cho, J., Yadav, P., and Bansal, M · 2023
Later among the works it cites.
A simple LLM framework for long-range video question-answering
Zhang, C., Lu, T., Islam, M. M., Wang, Z., Yu, S., Bansal, M., and Bertasius, G · 2023
Later among the works it cites.
A simple recipe for contrastively pre-training video-first encoders beyond 16 frames
Papalampidi, P., Koppula, S., Pathak, S., Chiu, J., Heyward, J., Patraucean, V., Shen, J., Miech, A., Zisserman, A., and Nematzdeh, A · 2024
Closest in time.
Mirasol3B: A multimodal autoregressive model for time-aligned and contextual modalities
Piergiovanni, A., Nobel, I., Kim, D., Ryoo, M. S., Gomes, V., and Angelova, A · 2024
Closest in time.
MovieChat: From dense token to sparse memory for long video understanding
Song, E., Chai, W., Wang, G., Zhang, Y., Zhou, H., Wu, F., Guo, X., Ye, T., Lu, Y., Hwang, J.-N., et al · 2024
Closest in time.
A generative model of memory construction and consolidation
Spens, E. and Burgess, N · 2024
Closest in time.