Fetching the paper…
Reading the bibliography…
The remarkable success of deep learning in various domains relies on the availability of large-scale annotated datasets.
Use what you have: Video retrieval using representations from collaborative experts
Y. Liu, S. Albanie, A. Nagrani, and A. Zisserman · 1907
Earlier work this paper cites.
“cloze procedure”: A new tool for measuring readability
Wilson L Taylor · 1953
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
The ava-kinetics localized human actions video dataset
Ang Li, Meghana Thotakuri, David A Ross, João Carreira, Alexander Vostrikov, and Andrew Zisserman · 2005
Earlier work this paper cites.
Contrastive estimation: Training log-linear models on unlabeled data
Noah A Smith and Jason Eisner · 2005
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Ronan Collobert and Jason Weston · 2008
Earlier work this paper cites.
Self-supervised video representation using pretext-contrastive learning
Li Tao, Xueting Wang, and Toshihiko Yamasaki · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander Hauptmann, Joao Henriques, and Andrea Vedaldi · 2010
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Learning fine-grained image similarity with deep ranking
Jiang Wang, Yang Song, Thomas Leung, Chuck Rosenberg, Jingbin Wang, James Philbin, Bo Chen, and Ying Wu · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A Efros · 2015
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Learning large-scale automatic image colorization
Aditya Deshpande, Jason Rock, and David Forsyth · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Spatial transformer networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition
Gregory Koch, Richard Zemel, Ruslan Salakhutdinov, et al · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
Shuffle and Learn: Unsupervised Learning Using Temporal Order Verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Tut database for acoustic scene classification and sound event detection
Annamaria Mesaros, Toni Heittola, and Tuomas Virtanen · 2016
Earlier work this paper cites.
Let there be color! joint end-to-end learning of global and local image priors for automatic image colorization with simultaneous classification
Satoshi Iizuka, Edgar Simo-Serra, and Hiroshi Ishikawa · 2016
Earlier work this paper cites.
Learning representations for automatic colorization
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michaël Mathieu, Camille Couprie, and Yann LeCun · 2016
Earlier work this paper cites.
Robust scene text recognition with automatic rectification
Baoguang Shi, Xinggang Wang, Pengyuan Lyu, Cong Yao, and Xiang Bai · 2016
Earlier work this paper cites.
Online action detection
Roeland De Geest, Efstratios Gavves, Amir Ghodrati, Zhenyang Li, Cees Snoek, and Tinne Tuytelaars · 2016
Earlier work this paper cites.
Netvlad: Cnn architecture for weakly supervised place recognition
Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Suggestive annotation: A deep active learning framework for biomedical image segmentation
Lin Yang, Yizhe Zhang, Jianxu Chen, Siyuan Zhang, and Danny Z Chen · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Unsupervised representation learning by sorting sequences
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2017
Earlier work this paper cites.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Lip reading sentences in the wild
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2017
Earlier work this paper cites.
Colorization as a proxy task for visual understanding
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich · 2017
Earlier work this paper cites.
Learning features by watching objects move
Deepak Pathak, Ross Girshick, Piotr Dollár, Trevor Darrell, and Bharath Hariharan · 2017
Earlier work this paper cites.
Deeppermnet: Visual permutation learning
Rodrigo Santa Cruz, Basura Fernando, Anoop Cherian, and Stephen Gould · 2017
Earlier work this paper cites.
Dual Motion GAN for Future-Flow Embedded Video Prediction
Xiaodan Liang, Lisa Lee, Wei Dai, and Eric P. Xing · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Lukasz Kaiser, Aidan N Gomez, Noam Shazeer, Ashish Vaswani, Niki Parmar, Llion Jones, and Jakob Uszkoreit · 2017
Earlier work this paper cites.
See, hear, and read: Deep aligned representations
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2017
Earlier work this paper cites.
Deep learning for biometrics: A survey
Kalaivani Sundararajan and Damon L. Woodard · 2018
Earlier work this paper cites.
Deep learning from crowds
Filipe Rodrigues and Francisco Pereira · 2018
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations
Nikos Komodakis and Spyros Gidaris · 2018
Earlier work this paper cites.
Learning image representations by completing damaged jigsaw puzzles
Dahun Kim, Donghyeon Cho, Donggeun Yoo, and In So Kweon · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Cited alongside, same era.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Cited alongside, same era.
What have we learned from deep representations for action recognition?
Christoph Feichtenhofer, Axel Pinz, Richard P Wildes, and Andrew Zisserman · 2018
Cited alongside, same era.
What makes a video a video: Analyzing temporal information in video understanding models and datasets
De-An Huang, Vignesh Ramanathan, Dhruv Mahajan, Lorenzo Torresani, Manohar Paluri, Li Fei-Fei, and Juan Carlos Niebles · 2018
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Coot: Cooperative hierarchical transformer for video-text representation learning
Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, and Thomas Brox · 2020
Later among the works it cites.
Temporally coherent embeddings for self-supervised video representation learning
Joshua Knights, Ben Harwood, Daniel Ward, Anthony Vanderkop, Olivia Mackenzie-Ross, and Peyman Moghadam · 2020
Later among the works it cites.
Self-supervised video representation learning by maximizing mutual information
Fei Xue, Hongbing Ji, Wenbo Zhang, and Yi Cao · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al · 2018
Cited alongside, same era.
Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos
Serena Yeung, Olga Russakovsky, Ning Jin, Mykhaylo Andriluka, Greg Mori, and Li Fei-Fei · 2018
Cited alongside, same era.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Cited alongside, same era.
VoxCeleb2: Deep Speaker Recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Cited alongside, same era.
Improvements to context based self-supervised learning
T Nathan Mundhenk, Daniel Ho, and Barry Y Chen · 2018
Cited alongside, same era.
Boosting self-supervised learning via knowledge transfer
Mehdi Noroozi, Ananth Vinjimoor, Paolo Favaro, and Hamed Pirsiavash · 2018
Cited alongside, same era.
Learning and using the arrow of time
Donglai Wei, Joseph J Lim, Andrew Zisserman, and William T Freeman · 2018
Cited alongside, same era.
Representation learning with video deep infomax
R Devon Hjelm and Philip Bachman · 2020
Later among the works it cites.
Temporal contrastive pretraining for video action recognition
Guillaume Lorre, Jaonary Rabarisoa, Astrid Orcesi, Samia Ainouz, and Stephane Canu · 2020
Later among the works it cites.
Self-supervised learning of pretext-invariant representations
Ishan Misra and Laurens van der Maaten · 2020
Later among the works it cites.
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al · 2020
Later among the works it cites.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Later among the works it cites.
Self-supervised learning by cross-modal audio-video clustering
Humam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani, Bernard Ghanem, and Du Tran · 2020
Later among the works it cites.
Unsupervised learning of video representations via dense trajectory clustering
Pavel Tokmakov, Martial Hebert, and Cordelia Schmid · 2020
Later among the works it cites.
Unsupervised learning from video with deep neural embeddings
Chengxu Zhuang, Tianwei She, Alex Andonian, Max Sobol Mark, and Daniel Yamins · 2020
Later among the works it cites.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Later among the works it cites.
AVLnet: Learning Audio-Visual Language Representations from Instructional Videos
Andrew Rouditchenko, Angie Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogerio Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, and James Glass · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Later among the works it cites.
Self-supervised Learning of Audio-Visual Objects from Video , volume 12363 LNCS
Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman · 2020
Later among the works it cites.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman · 2020
Later among the works it cites.
Revisiting superpixels for active learning in semantic segmentation with realistic annotation costs
Lile Cai, Xun Xu, Jun Hao Liew, and Chuan Sheng Foo · 2021
Later among the works it cites.
Understanding and mitigating annotation bias in facial expression recognition
Yunliang Chen and Jungseock Joo · 2021
Later among the works it cites.
Cds: Cross-domain self-supervised pre-training
Donghyun Kim, Kuniaki Saito, Tae-Hyun Oh, Bryan A. Plummer, Stan Sclaroff, and Kate Saenko · 2021
Later among the works it cites.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Spatio-temporal self-supervised representation learning for 3d point clouds
Siyuan Huang, Yichen Xie, Song-Chun Zhu, and Yixin Zhu · 2021
Later among the works it cites.
Self-supervised video representation learning with constrained spatiotemporal jigsaw
Yuqi Huo, Mingyu Ding, Haoyu Lu, Ziyuan Huang, Mingqian Tang, Zhiwu Lu, and Tao Xiang · 2021
Later among the works it cites.
Temporally coherent embeddings for self-supervised video representation learning
Joshua Knights, Ben Harwood, Daniel Ward, Anthony Vanderkop, Olivia Mackenzie-Ross, and Peyman Moghadam · 2021
Later among the works it cites.
A survey on contrastive self-supervised learning
Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh, Debapriya Banerjee, and Fillia Makedon · 2021
Later among the works it cites.
VLM: Task-agnostic video-language model pre-training for video understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Prahal Arora, Masoumeh Aminzadeh, Christoph Feichtenhofer, Florian Metze, and Luke Zettlemoyer · 2021
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Later among the works it cites.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Later among the works it cites.
Deep audio-visual learning: A survey
Hao Zhu, Man-Di Luo, Rui Wang, Ai-Hua Zheng, and Ran He · 2021
Later among the works it cites.
Value: A multi-task benchmark for video-and-language understanding evaluation
Linjie Li, Jie Lei, Zhe Gan, Licheng Yu, Yen-Chun Chen, Rohit Pillai, Yu Cheng, Luowei Zhou, Xin Eric Wang, William Yang Wang, et al · 2021
Later among the works it cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Later among the works it cites.
Enhancing unsupervised video representation learning by decoupling the scene and the motion
Jinpeng Wang, Yuting Gao, Ke Li, Xinyang Jiang, Xiaowei Guo, Rongrong Ji, and Xing Sun · 2021
Later among the works it cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Later among the works it cites.
Representation learning via global temporal alignment and cycle-consistency
Isma Hadji, Konstantinos G. Derpanis, and Allan D. Jepson · 2021
Later among the works it cites.
Videomoco: Contrastive video representation learning with temporally adversarial examples
Tian Pan, Yibing Song, Tianyu Yang, Wenhao Jiang, and Wei Liu · 2021
Later among the works it cites.
A large-scale study on unsupervised spatiotemporal representation learning
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He · 2021
Later among the works it cites.
Seco: Exploring sequence supervision for unsupervised representation learning
Ting Yao, Yiheng Zhang, Zhaofan Qiu, Yingwei Pan, and Tao Mei · 2021
Later among the works it cites.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui · 2021
Later among the works it cites.
Cocon: Cooperative-contrastive learning
Nishant Rai, Ehsan Adeli, Kuan-Hui Lee, Adrien Gaidon, and Juan Carlos Niebles · 2021
Later among the works it cites.
Noise estimation using density estimation for self-supervised multimodal learning
Elad Amrani, Rami Ben-Ari, Daniel Rotman, and Alex Bronstein · 2021
Later among the works it cites.
Deep video action clustering via spatio-temporal feature learning
Bo Peng, Jianjun Lei, Huazhu Fu, Yalong Jia, Zongqian Zhang, and Yi Li · 2021
Later among the works it cites.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Linagzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Later among the works it cites.
Audio-visual instance discrimination with cross-modal agreement
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra · 2021
Later among the works it cites.
Multimodal clustering networks for self-supervised learning from unlabeled videos
Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie Boggust, Rameswar Panda, Brian Kingsbury, Rogerio Feris, David Harwath, James Glass, Michael Picheny, and Shih-Fu Chang · 2021
Later among the works it cites.
W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu · 2021
Later among the works it cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Later among the works it cites.
Understanding robustness of transformers for image classification
Srinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li, Thomas Unterthiner, and Andreas Veit · 2021
Later among the works it cites.
Multibench: Multiscale benchmarks for multimodal representation learning
Paul Pu Liang, Yiwei Lyu, Xiang Fan, Zetian Wu, Yun Cheng, Jason Wu, Leslie Chen, Peter Wu, Michelle A Lee, Yuke Zhu, et al · 2021
Later among the works it cites.
Graph self-supervised learning: A survey
Yixin Liu, Ming Jin, Shirui Pan, Chuan Zhou, Yu Zheng, Feng Xia, and Philip Yu · 2022
Closest in time.
Revisiting the" video" in video-language understanding
Shyamal Buch, Cristóbal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles · 2022
Closest in time.
Robustness analysis of video-language models against visual and language perturbations
Madeline C. Schiappa, Shruti Vyas, Hamid Palangi, Yogesh S. Rawat, and Vibhav Vineet · 2022
Closest in time.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Closest in time.
Masked feature prediction for self-supervised visual pre-training
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer · 2022
Closest in time.
Self-supervised video representation learning with motion-aware masked autoencoders, 2022
Haosen Yang, Deng Huang, Bin Wen, Jiannan Wu, Hongxun Yao, Yi Jiang, Xiatian Zhu, and Zehuan Yuan · 2022
Closest in time.
Masked autoencoders as spatiotemporal learners
Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He · 2022
Closest in time.
VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Closest in time.
Tclr: Temporal contrastive learning for video representation
Ishan Dave, Rohit Gupta, Mamshad Nayeem Rizve, and Mubarak Shah · 2022
Closest in time.
Self-supervised video transformer
Kanchana Ranasinghe, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan, and Michael S Ryoo · 2022
Closest in time.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Closest in time.