Learning Video Representations using Contrastive Bidirectional Transformer
Original
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 1906
Earlier work this paper cites.
UNITER: UNiversal Image-TExt Representation Learning
Original
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 1909
Earlier work this paper cites.
Pruning Algorithms-A Survey
Russell Reed · 1993
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Noise-Contrastive Estimation: A New Estimation Principle for Unnormalized Statistical Models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Multimodal Deep Learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng · 2011
Earlier work this paper cites.
Matrix Computations , volume 3
Gene H Golub and Charles F Van Loan · 2012
Earlier work this paper cites.
UCF101: A Dataset of 101 Human Action Classes From Videos in The Wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Multimodal Learning with Deep Boltzmann Machines
Nitish Srivastava and Russ R Salakhutdinov · 2012
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
FaceNet: A Unified Embedding for Face Recognition and Clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
SoundNet: Learning Sound Representations from Unlabeled Video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Layer Normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs)
Original
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size
Original
Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer · 2016
Earlier work this paper cites.
Visually Indicated Sounds
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman · 2016
Earlier work this paper cites.
XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi · 2016
Earlier work this paper cites.
Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Look, Listen and Learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Audio Set: An Ontology and Human-Labeled Dataset for Audio Events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Original
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
The Kinetics Human Action Video Dataset
Original
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Asynchronous Temporal Fields for Action Recognition
Gunnar A Sigurdsson, Santosh Divvala, Ali Farhadi, and Abhinav Gupta · 2017
Earlier work this paper cites.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Objects that Sound
Relja Arandjelovic and Andrew Zisserman · 2018
Earlier work this paper cites.
Multimodal Machine Learning: A Survey and Taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.