Fetching the paper…
Reading the bibliography…
The ability to perceive how objects change over time is a crucial ingredient in human intelligence.
1908
Earlier work this paper cites.
Evans, V.: 22 how we conceptualise time: language, meaning and temporal cognition. The cognitive linguistics reader p. 733 (2004)
2004
Earlier work this paper cites.
Chen, D.L., Dolan, W.B.: Collecting highly parallel data for paraphrase evaluation. In: Annual Meeting of the Association for Computational Linguistics (2011)
2011
Earlier work this paper cites.
Klein, W.: Time in language. routledge (2013)
2013
Earlier work this paper cites.
Rohrbach, A., Rohrbach, M., Tandon, N., Schiele, B.: A dataset for movie description. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) pp. 3202–3212 (2015)
2015
Earlier work this paper cites.
Xu, J., Mei, T., Yao, T., Rui, Y.: Msr-vtt: A large video description dataset for bridging video and language. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) pp. 5288–5296 (2016)
2016
Earlier work this paper cites.
Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fründ, I., Yianilos, P.N., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., Memisevic, R.: The “something something” video database for learning and evaluating visual common sense. 2017 IEEE International Conference on Computer Vision (ICCV) pp. 5843–5851 (2017)
2017
Earlier work this paper cites.
Krishna, R., Hata, K., Ren, F., Fei-Fei, L., Niebles, J.C.: Dense-captioning events in videos. 2017 IEEE International Conference on Computer Vision (ICCV) pp. 706–715 (2017)
2017
Earlier work this paper cites.
Zhou, L., Xu, C., Corso, J.J.: Towards automatic learning of procedures from web instructional videos. In: AAAI Conference on Artificial Intelligence (2017)
2017
Earlier work this paper cites.
Ghodrati, A., Gavves, E., Snoek, C.G.M.: Video time: Properties, encoders and evaluation. In: British Machine Vision Conference (2018)
2018
Earlier work this paper cites.
Hendricks, L.A., Wang, O., Shechtman, E., Sivic, J., Darrell, T., Russell, B.C.: Localizing moments in video with temporal language. In: Conference on Empirical Methods in Natural Language Processing (2018)
2018
Earlier work this paper cites.
Huang, D.A., Ramanathan, V., Mahajan, D.K., Torresani, L., Paluri, M., Fei-Fei, L., Niebles, J.C.: What makes a video a video: Analyzing temporal information in video understanding models and datasets. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition pp. 7366–7375 (2018)
2018
Earlier work this paper cites.
Li, Y., Li, Y., Vasconcelos, N.: RESOUND: towards action recognition without representation bias. In: ECCV (6). Lecture Notes in Computer Science, vol. 11210, pp. 520–535. Springer (2018)
2018
Earlier work this paper cites.
Wang, Y., Hoai, M.: Pulling actions out of context: Explicit separation for effective combination. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition pp. 7044–7053 (2018)
2018
Earlier work this paper cites.
Wei, D., Lim, J.J., Zisserman, A., Freeman, W.T.: Learning and using the arrow of time. In: CVPR. pp. 8052–8060. Computer Vision Foundation / IEEE Computer Society (2018)
2018
Earlier work this paper cites.
Choi, J., Gao, C., Messou, J.C., Huang, J.B.: Why can’t i dance in the mall? learning to mitigate scene bias in action recognition. In: Neural Information Processing Systems (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Price, W., Damen, D.: Retro-actions: Learning ’close’ by time-reversing ’open’ videos. In: ICCV Workshops. pp. 1371–1380. IEEE (2019)
2019
Earlier work this paper cites.
Sevilla-Lara, L., Zha, S., Yan, Z., Goswami, V., Feiszli, M., Torresani, L.: Only time can tell: Discovering temporal data for temporal modeling. 2021 IEEE Winter Conference on Applications of Computer Vision (WACV) pp. 535–544 (2019)
2019
Earlier work this paper cites.
Wang, X.E., Wu, J., Chen, J., Li, L., fang Wang, Y., Wang, W.Y.: Vatex: A large-scale, high-quality multilingual dataset for video-and-language research. 2019 IEEE/CVF International Conference on Computer Vision (ICCV) pp. 4580–4590 (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Benaim, S., Ephrat, A., Lang, O., Mosseri, I., Freeman, W.T., Rubinstein, M., Irani, M., Dekel, T.: Speednet: Learning the speediness in videos. In: CVPR. pp. 9919–9928. Computer Vision Foundation / IEEE (2020)
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
Li, L., Chen, Y.C., Cheng, Y., Gan, Z., Yu, L., Liu, J.: Hero: Hierarchical encoder for video+language omni-representation pre-training. In: Conference on Empirical Methods in Natural Language Processing (2020)
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Shao, D., Zhao, Y., Dai, B., Lin, D.: Finegym: A hierarchical video dataset for fine-grained action understanding. In: CVPR. pp. 2613–2622. Computer Vision Foundation / IEEE (2020)
2022
Later among the works it cites.
Li, J., Li, D., Xiong, C., Hoi, S.C.H.: Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: International Conference on Machine Learning (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Wang, J., Gao, Y., Li, K., Lin, Y., Ma, A.J., Sun, X.: Removing the background by adding the background: Towards background robust self-supervised video representation learning. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp. 11799–11808 (2020)
2020
Cited alongside, same era.
Zhang, Z., Yin, Z., Ren, S., Li, X., Li, S.: Dca: Diversified co-attention towards informative live video commenting. In: Natural Language Processing and Chinese Computing (2020)
2020
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Huang, L., Liu, Y., Wang, B., Pan, P., Xu, Y., Jin, R.: Self-supervised video representation learning by context and motion decoupling. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp. 13881–13890 (2021)
2021
Cited alongside, same era.
Lei, J., Li, L., Zhou, L., Gan, Z., Berg, T.L., Bansal, M., Liu, J.: Less is more: Clipbert for video-and-language learning via sparse sampling. In: CVPR (2021)
2021
Cited alongside, same era.
Li, D., Li, J., Li, H., Niebles, J.C., Hoi, S.C.H.: Align and prompt: Video-and-language pre-training with entity prompts. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp. 4943–4953 (2021)
2021
Cited alongside, same era.
Luo, H., Ji, L., Zhong, M., Chen, Y., Lei, W., Duan, N., Li, T.: Clip4clip: An empirical study of clip for end to end video clip retrieval. Neurocomputing 508
2021
Cited alongside, same era.
Ma, Y., Xu, G., Sun, X., Yan, M., Zhang, J.C., Ji, R.: X-clip: End-to-end multi-grained contrastive learning for video-text retrieval. Proceedings of the 30th ACM International Conference on Multimedia (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
OpenAI: Introducing ChatGPT (2022), https://openai.com/blog/chatgpt
2022
Later among the works it cites.
Park, J.S., Shen, S., Farhadi, A., Darrell, T., Choi, Y., Rohrbach, A.: Exposing the limits of video-text models through contrast sets. In: North American Chapter of the Association for Computational Linguistics (2022)
2022
Later among the works it cites.
Thrush, T., Jiang, R., Bartolo, M., Singh, A., Williams, A., Kiela, D., Ross, C.: Winoground: Probing vision and language models for visio-linguistic compositionality. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) pp. 5228–5238 (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Ye, J., Gao, J., Feng, J., Wu, Z., Yu, T., Kong, L.: Progen: Progressive zero-shot dataset generation via in-context feedback. In: Conference on Empirical Methods in Natural Language Processing (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
Liu, Y., Li, L., Ren, S., Gao, R., Li, S., Chen, S., Sun, X., Hou, L.: FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation. In: NeurIPS (2023)
2023
Closest in time.
2023
Closest in time.
Ren, S., Chen, S., Li, S., Sun, X., Hou, L.: TESTA: temporal-spatial token aggregation for long-form video-language understanding. In: EMNLP (Findings). pp. 932–947. Association for Computational Linguistics (2023)
2023
Closest in time.
2023
Closest in time.