Fetching the paper…
Reading the bibliography…
Videos on the Internet are paired with pieces of text, such as titles and descriptions.
Annotating images by mining image search results
Xin-Jing Wang, Lei Zhang, Xirong Li, and Wei-Ying Ma · 2008
Earlier work this paper cites.
Exploiting weakly-labeled web images to improve object classification: a domain adaptation approach
Alessandro Bergamo and Lorenzo Torresani · 2010
Earlier work this paper cites.
Learning object categories from internet image searches
Rob Fergus, Li Fei-Fei, Pietro Perona, and Andrew Zisserman · 2010
Earlier work this paper cites.
Harvesting image databases from the web
Florian Schroff, Antonio Criminisi, and Andrew Zisserman · 2010
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara L Berg · 2011
Earlier work this paper cites.
A dataset of 101 human action classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and M Shah · 2012
Earlier work this paper cites.
Neil: Extracting visual knowledge from web data
Xinlei Chen, Abhinav Shrivastava, and Abhinav Gupta · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Learning everything about anything: Webly-supervised visual concept learning
Santosh K Divvala, Ali Farhadi, and Carlos Guestrin · 2014
Earlier work this paper cites.
A multi-view embedding space for modeling internet images, tags, and their semantics
Yunchao Gong, Qifa Ke, Michael Isard, and Svetlana Lazebnik · 2014
Earlier work this paper cites.
Improving image-sentence embeddings using large weakly annotated photo collections
Yunchao Gong, Liwei Wang, Micah Hodosh, Julia Hockenmaier, and Svetlana Lazebnik · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Instructional videos for unsupervised harvesting and learning of action examples
Shoou-I Yu, Lu Jiang, and Alexander Hauptmann · 2014
Earlier work this paper cites.
Webly supervised learning of convolutional networks
Xinlei Chen and Abhinav Gupta · 2015
Earlier work this paper cites.
Deep classifiers from image tags in the wild
Hamid Izadinia, Bryan C Russell, Ali Farhadi, Matthew D Hoffman, and Aaron Hertzmann · 2015
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien · 2016
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Webly-supervised video recognition by mutually voting for relevant web images and web video frames
Chuang Gan, Chen Sun, Lixin Duan, and Boqing Gong · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning visual features from large weakly supervised data
Armand Joulin, Laurens van der Maaten, Allan Jabri, and Nicolas Vasilache · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
William Lotter, Gabriel Kreiman, and David Cox · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Cited alongside, same era.
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba · 2016
Cited alongside, same era.
Anticipating visual representations from unlabeled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Later among the works it cites.
Self-supervised learning by cross-modal audio-video clustering
Humam Alwassel, Dhruv Mahajan, Lorenzo Torresani, Bernard Ghanem, and Du Tran · 2019
Later among the works it cites.
A short note on the kinetics-700 human action dataset
Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman · 2019
Later among the works it cites.
Temporal cycle-consistency learning
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zisserman · 2019
Later among the works it cites.
Large-scale weakly-supervised pre-training for video action recognition
Deepti Ghadiyaram, Du Tran, and Dhruv Mahajan · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
Unsupervised representation learning by sorting sequences
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2017
Cited alongside, same era.
Learning visual n-grams from web data
Ang Li, Allan Jabri, Armand Joulin, and Laurens van der Maaten · 2017
Cited alongside, same era.
Learning features by watching objects move
Deepak Pathak, Ross Girshick, Piotr Dollár, Trevor Darrell, and Bharath Hariharan · 2017
Cited alongside, same era.
A short note about kinetics-600
Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman · 2018
Cited alongside, same era.
More Than 500 Hours Of Content Are Now Being Uploaded To YouTube Every Minute, 2019
James Hale · 2019
Later among the works it cites.
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman · 2019
Later among the works it cites.
Quantifying Diminishing Returns of Annotated Data, 2019
Gal Hyams, Dan Malowany, Ariel Biller, and Gregory Axler · 2019
Later among the works it cites.
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, and In So Kweon · 2019
Later among the works it cites.
Mining youtube-a dataset for learning fine-grained action concepts from webly supervised video data
Hilde Kuehne, Ahsan Iqbal, Alexander Richard, and Juergen Gall · 2019
Later among the works it cites.
Self-supervised learning for video correspondence flow
Zihang Lai and Weidi Xie · 2019
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2019
Later among the works it cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Later among the works it cites.
Self-supervised audio-visual co-segmentation
Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh McDermott, and Antonio Torralba · 2019
Later among the works it cites.
Contrastive bidirectional transformer for temporal representation learning
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Yunhui Liu, and Wei Liu · 2019
Later among the works it cites.
Learning correspondence from the cycle-consistency of time
Xiaolong Wang, Allan Jabri, and Alexei A Efros · 2019
Later among the works it cites.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Later among the works it cites.
Cross-task weakly supervised learning from instructional videos
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic · 2019
Later among the works it cites.
Watching the world go by: Representation learning from unlabeled videos
Daniel Gordon, Kiana Ehsani, Dieter Fox, and Ali Farhadi · 2020
Closest in time.
Learning spatiotemporal features via video and text pair discrimination
Tianhao Li and Limin Wang · 2020
Closest in time.
Speech2action: Cross-modal supervision for action recognition
Arsha Nagrani, Sun Chen, David Ross, Rahul Sukthankar, Cordelia Schmid, and Andrew Zisserman · 2020
Closest in time.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui · 2020
Closest in time.
Video representation learning with visual tempo consistency
Ceyuan Yang, Yinghao Xu, Bo Dai, and Bolei Zhou · 2020
Closest in time.