Fetching the paper…
Reading the bibliography…
Modelling and understanding time remains a challenge in contemporary video understanding models.
Towards a general theory of action and time
James F. Allen · 1984
Earlier work this paper cites.
Temporal prepositions and temporal generalized quantifiers
Ianthe Pratt and Nissim Francez · 2001
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Gregory S. Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Seeing the arrow of time
Lyndsey C. Pickup, Zheng Pan, Donglai Wei, YiChang Shih, Changshui Zhang, Andrew Zisserman, Bernhard Scholkopf, and William T. Freeman · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Book2movie: Aligning video scenes with book chapters
Makarand Tapaswi, Martin Bauml, and Rainer Stiefelhagen · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Earlier work this paper cites.
The “something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, João Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Earlier work this paper cites.
Dense-Captioning Events in Videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
A structured learning approach to temporal relation extraction
Qiang Ning, Zhili Feng, and Dan Roth · 2017
Earlier work this paper cites.
What actions are needed for understanding human actions in videos?
Gunnar A. Sigurdsson, Olga Russakovsky, and Abhinav Kumar Gupta · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Rethinking spatiotemporal feature learning for video understanding
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin P. Murphy · 2017
Earlier work this paper cites.
Video time: Properties, encoders and evaluation
Amir Ghodrati, Efstratios Gavves, and Cees G. M. Snoek · 2018
Earlier work this paper cites.
Localizing moments in video with temporal language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2018
Earlier work this paper cites.
Localizing moments in video with temporal language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell · 2018
Earlier work this paper cites.
What makes a video a video: Analyzing temporal information in video understanding models and datasets
De-An Huang, Vignesh Ramanathan, Dhruv Mahajan, Lorenzo Torresani, Manohar Paluri, Li Fei-Fei, and Juan Carlos Niebles · 2018
Earlier work this paper cites.
Joint reasoning for temporal and causal relations
Qiang Ning, Zhili Feng, Hao Wu, and Dan Roth · 2018
Earlier work this paper cites.
Cogcomptime: A tool for understanding time in natural language
Qiang Ning, Ben Zhou, Zhili Feng, Haoruo Peng, and Dan Roth · 2018
Earlier work this paper cites.
Actor and observer: Joint modeling of first and third-person videos
Gunnar A. Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Earlier work this paper cites.
Learning spatiotemporal 3d convolution with video order self-supervision
Tomoyuki Suzuki, Takahiro Itazuri, Kensho Hara, and Hirokatsu Kataoka · 2018
Earlier work this paper cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Learning and using the arrow of time
Donglai Wei, Joseph J Lim, Andrew Zisserman, and William T Freeman · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Earlier work this paper cites.
Video jigsaw: Unsupervised learning of spatiotemporal context for video action recognition
Unaiza Ahsan, Rishi Madhok, and Irfan Essa · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Joint event and temporal relation extraction with shared representations and structured prediction
Rujun Han, Qiang Ning, and Nanyun Peng · 2019
Earlier work this paper cites.
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, and In So Kweon · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
An improved neural baseline for temporal relation extraction
Qiang Ning, Sanjay Subramanian, and Dan Roth · 2019
Earlier work this paper cites.
Retro-actions: Learning ’close’ by time-reversing ’open’ videos
Will Price and Dima Damen · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Earlier work this paper cites.
Only time can tell: Discovering temporal data for temporal modeling
Laura Sevilla-Lara, Shengxin Zha, Zhicheng Yan, Vedanuj Goswami, Matt Feiszli, and Lorenzo Torresani · 2019
Earlier work this paper cites.
Deep Multimodal Feature Encoding for Video Ordering
Vivek Sharma, Makarand Tapaswi, and Rainer Stiefelhagen · 2019
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin P. Murphy, and Cordelia Schmid · 2019
Cited alongside, same era.
Learning correspondence from the cycle-consistency of time
X. Wang, A. Jabri, and Alexei A. Efros · 2019
Cited alongside, same era.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Cited alongside, same era.
Ben Zhou, Daniel Khashabi, Qiang Ning, and Dan Roth · 2019
Cited alongside, same era.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, H. Wang, Serge J. Belongie, and Yin Cui · 2021
Later among the works it cites.
Timedial: Temporal commonsense reasoning in dialog
Lianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He, Yejin Choi, and Manaal Faruqui · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
Broaden your views for self-supervised video learning
Adrià Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Patraucean, Florent Altch’e, Michael Valko, Jean-Bastien Grill, Aäron van den Oord, and Andrew Zisserman · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovi’c, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman · 2020
Cited alongside, same era.
Speednet: Learning the speediness in videos
Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T Freeman, Michael Rubinstein, Michal Irani, and Tali Dekel · 2020
Cited alongside, same era.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Alahari Karteek, and Cordelia Schmid · 2020
Cited alongside, same era.
Coot: Cooperative hierarchical transformer for video-text representation learning
Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, and Thomas Brox · 2020
Cited alongside, same era.
Deer: A data efficient language model for event temporal reasoning
Rujun Han, Xiang Ren, and Nanyun Peng · 2020
Cited alongside, same era.
Space-time correspondence as a contrastive random walk
A. Jabri, Andrew Owens, and Alexei A. Efros · 2020
Cited alongside, same era.
Video representation learning by recognizing temporal transformations
Simon Jenni, Givi Meishvili, and Paolo Favaro · 2020
Cited alongside, same era.
Later among the works it cites.
Do image classifiers generalize across time?
Vaishaal Shankar, Achal Dave, Rebecca Roelofs, Deva Ramanan, Benjamin Recht, and Ludwig Schmidt · 2021
Later among the works it cites.
Semi-supervised action recognition with temporal contrastive learning
Ankit Singh, Omprakash Chakraborty, Ashutosh Varshney, Rameswar Panda, Rogério Schmidt Feris, Kate Saenko, and Abir Das · 2021
Later among the works it cites.
Skeleton-contrastive 3d action representation learning
Fida Mohammad Thoker, Hazel Doughty, and Cees G. M. Snoek · 2021
Later among the works it cites.
Probing language models for understanding of temporal expressions
Shivin Thukral, Kunal Kukreja, and Christian Kavouras · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Unsupervised visual representation learning by tracking patches in video
Guangting Wang, Yizhou Zhou, Chong Luo, Wenxuan Xie, Wenjun Zeng, and Zhiwei Xiong · 2021
Later among the works it cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Later among the works it cites.
VideoCLIP: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel C. F. Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, Ce Liu, Mengchen Liu, Zicheng Liu, Yumao Lu, Yu Shi, Lijuan Wang, Jianfeng Wang, Bin Xiao, Zhen Xiao, Jianwei Yang, Michael Zeng, Luowei Zhou, and Pengchuan Zhang · 2021
Later among the works it cites.
Temporal reasoning on implicit events from distant supervision
Ben Zhou, Kyle Richardson, Qiang Ning, Tushar Khot, Ashish Sabharwal, and Dan Roth · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan · 2022
Later among the works it cites.
A clip-hitchhiker’s guide to long video retrieval, 2022
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2022
Later among the works it cites.
Revisiting the “Video” in Video-Language Understanding
Shyamal Buch, Cristobal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles · 2022
Later among the works it cites.
Are 3d convolutional networks inherently biased towards appearance?
Petr Byvshev, Pascal Mettes, and Yu Xiao · 2022
Later among the works it cites.
Locvtp: Video-text pre-training for temporal localization
Meng Cao, Tianyu Yang, Junwu Weng, Can Zhang, Jue Wang, and Yuexian Zou · 2022
Later among the works it cites.
Vindlu: A recipe for effective video-and-language pretraining
Feng Cheng, Xizi Wang, Jie Lei, David Crandall, Mohit Bansal, and Gedas Bertasius · 2022
Later among the works it cites.
Tclr: Temporal contrastive learning for video representation
Ishan Rajendra Dave, Rohit Gupta, Mamshad Nayeem Rizve, and Mubarak Shah · 2022
Later among the works it cites.
Time-aware language models as temporal knowledge bases
Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen · 2022
Later among the works it cites.
Bridging video-text retrieval with multiple choice questions
Yuying Ge, Yixiao Ge, Xihui Liu, Dian Li, Ying Shan, Xiaohu Qie, and Ping Luo · 2022
Later among the works it cites.
Omnivore: A single model for many visual modalities
Rohit Girdhar, Mannat Singh, Nikhil Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra · 2022
Later among the works it cites.
Temporal alignment networks for long-term video
Tengda Han, Weidi Xie, and Andrew Zisserman · 2022
Later among the works it cites.
Revealing single frame bias for video-and-language learning, 2022
Jie Lei, Tamara L. Berg, and Mohit Bansal · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi · 2022
Later among the works it cites.
Compositional temporal grounding with structured variational cross-graph correspondence learning
Juncheng Li, Junlin Xie, Long Qian, Linchao Zhu, Siliang Tang, Fei Wu, Yi Yang, Yueting Zhuang, and Xin Eric Wang · 2022
Later among the works it cites.
Self-supervised spatiotemporal representation learning by exploiting video continuity
Hanwen Liang, Niamul Quader, Zhixiang Chi, Lizhe Chen, Peng Dai, Juwei Lu, and Yang Wang · 2022
Later among the works it cites.
Egocentric video-language pretraining
Kevin Lin, Alex Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z. Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, Chengfei Cai, Hongfa Wang, Dima Damen, Bernard Ghanem, Wei Liu, and Mike Zheng Shou · 2022
Later among the works it cites.
Frozen clip models are efficient video learners
Ziyi Lin, Shijie Geng, Renrui Zhang, Peng Gao, Gerard de Melo, Xiaogang Wang, Jifeng Dai, Y. Qiao, and Hongsheng Li · 2022
Later among the works it cites.
Expanding language-image pretrained models for general video recognition
Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling · 2022
Later among the works it cites.
Self-supervised learning for videos: A survey
Madeline Chantry Schiappa, Yogesh Singh Rawat, and Mubarak Shah · 2022
Later among the works it cites.
Long-form video-language pre-training with multimodal temporal contrastive learning
Yuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu, Huan Yang, and Jianlong Fu · 2022
Later among the works it cites.
How severe is benchmark-sensitivity in video self-supervised learning?
Fida Mohammad Thoker, Hazel Doughty, Piyush Bagad, and Cees G. M. Snoek · 2022
Later among the works it cites.
Omnivl: One foundation model for image-language and video-language tasks
Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, and Lu Yuan · 2022
Later among the works it cites.
Omnivl: One foundation model for image-language and video-language tasks
Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, and Lu Yuan · 2022
Later among the works it cites.
Bevt: Bert pretraining of video transformers
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, and Lu Yuan · 2022
Later among the works it cites.
Clip-vip: Adapting pre-trained image-text model to video-language representation alignment
Hongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu, Rui Song, Houqiang Li, and Jiebo Luo · 2022
Later among the works it cites.
Zero-shot video question answering via frozen bidirectional language models
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2022
Later among the works it cites.
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng, Maria Attarian, Brian Ichter, Krzysztof Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, and Pete Florence · 2022
Later among the works it cites.
Centerclip: Token clustering for efficient text-video retrieval
Shuai Zhao, Linchao Zhu, Xiaohan Wang, and Yi Yang · 2022
Later among the works it cites.
Tubelet-contrastive self-supervision for video-efficient generalization, 2023
Fida Mohammad Thoker, Hazel Doughty, and Cees Snoek · 2023
Closest in time.