Fetching the paper…
Reading the bibliography…
Annotating videos is cumbersome, expensive and not scalable.
Solving the multiple instance problem with axis-parallel rectangles
Thomas G Dietterich, Richard H Lathrop, and Tomás Lozano-Pérez · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jurgen Schmidhuber · 1997
Earlier work this paper cites.
Support vector machines for multiple-instance learning
Stuart Andrews, Ioannis Tsochantaridis, and Thomas Hofmann · 2003
Earlier work this paper cites.
DIFFRAC: a discriminative and flexible framework for clustering
Francis Bach and Zaïd Harchaoui · 2007
Earlier work this paper cites.
Visual tracking with online multiple instance learning
Boris Babenko, Ming-Hsuan Yang, and Serge Belongie · 2009
Earlier work this paper cites.
Automatic annotation of human actions in video
Olivier Duchenne, Ivan Laptev, Josef Sivic, Francis Bach, and Jean Ponce · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Handling label noise in video classification via multiple instance learning
Thomas Leung, Yang Song, and John Zhang · 2011
Earlier work this paper cites.
Similarity constrained latent support vector machine: An application to weakly supervised action classification
Nataliya Shapovalova, Arash Vahdat, Kevin Cannons, Tian Lan, and Greg Mori · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Finding Actors and Actions in Movies
Piotr Bojanowski, Francis Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, and Josef Sivic · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
A multi-view embedding space for modeling internet images, tags, and their semantics
Yunchao Gong, Qifa Ke, Michael Isard, and Svetlana Lazebnik · 2014
Earlier work this paper cites.
Improving image-sentence embeddings using large weakly annotated photo collections
Yunchao Gong, Liwei Wang, Micah Hodosh, Julia Hockenmaier, and Svetlana Lazebnik · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Fei Fei F Li · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Linking people in videos with “their” names using coreference resolution
Vignesh Ramanathan, Armand Joulin, Percy Liang, and Li Fei-Fei · 2014
Earlier work this paper cites.
Instructional videos for unsupervised harvesting and learning of action examples
Shoou-I Yu, Lu Jiang, and Alexander Hauptmann · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors
Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf · 2015
Earlier work this paper cites.
What’s cookin’? interpreting cooking videos using text, speech and vision
Jonathan Malmaud, Jonathan Huang, Vivek Rathod, Nick Johnston, Andrew Rabinovich, and Kevin Murphy · 2015
Earlier work this paper cites.
Is object localization for free? - weakly-supervised learning with convolutional neural networks
Maxime Oquab, Leon Bottou, Ivan Laptev, and Josef Sivic · 2015
Earlier work this paper cites.
It’s in the bag: Stronger supervision for automated face labelling
Omkar Parkhi, Esa Rahtu, and Andrew Zisserman · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos
Xiaolong Wang and Abhinav Gupta · 2015
Earlier work this paper cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
Ran Xu, Caiming Xiong, Wei Chen, and Jason J Corso · 2015
Earlier work this paper cites.
YouTube-8M: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Ivan Laptev, Josef Sivic, and Simon Lacoste-Julien · 2016
Earlier work this paper cites.
NetVLAD: CNN architecture for weakly supervised place recognition
Relja Arandjelović, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Cited alongside, same era.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Cited alongside, same era.
Jointly modeling embedding and translation to bridge video and language
Yingwei Pan, Tao Mei, Ting Yao, Houqiang Li, and Yong Rui · 2016
Cited alongside, same era.
Anticipating visual representations from unlabeled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
Real-world anomaly detection in surveillance videos
Waqas Sultani, Chen Chen, and Mubarak Shah · 2018
Later among the works it cites.
Tracking emerges by colorizing videos
Carl Vondrick, Abhinav Shrivastava, Alireza Fathi, Sergio Guadarrama, and Kevin Murphy · 2018
Later among the works it cites.
Learning two-branch neural networks for image-text matching tasks
Liwei Wang, Yin Li, Jing Huang, and Svetlana Lazebnik · 2018
Later among the works it cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Later among the works it cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2016
Cited alongside, same era.
Towards weakly-supervised action localization
Philippe Weinzaepfel, Xavier Martin, and Cordelia Schmid · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Cited alongside, same era.
Look, listen and learn
Relja Arandjelović and Andrew Zisserman · 2017
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
João Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Discover and learn new objects from documentaries
Kai Chen, Hang Song, Chen Change Loy, and Dahua Lin · 2017
Cited alongside, same era.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Cited alongside, same era.
Elad Amrani, Rami Ben-Ari, Tal Hakim, and Alex Bronstein · 2019
Closest in time.
Grounding spoken words in unlabeled video
Angie Boggust, Kartik Audhkhasi, Dhiraj Joshi, David Harwath, Samuel Thomas, Rogerio Feris, Dan Gutfreund, Yang Zhang, Antonio Torralba, Michael Picheny, et al · 2019
Closest in time.
Unsupervised pre-training of image features on non-curated data
Mathilde Caron, Piotr Bojanowski, Julien Mairal, and Armand Joulin · 2019
Closest in time.
A short note on the kinetics-700 human action dataset
João Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Closest in time.
Dual encoding for zero-example video retrieval
Jianfeng Dong, Xirong Li, Chaoxi Xu, Shouling Ji, Yuan He, Gang Yang, and Xun Wang · 2019
Closest in time.
Temporal cycle-consistency learning
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zisserman · 2019
Closest in time.
Geometry guided convolutional neural networks for self-supervised video representation learning
Chuang Gan, Boqing Gong, Kun Liu, Hao Su, and Leonidas J Guibas · 2019
Closest in time.
Large-scale weakly-supervised pre-training for video action recognition
Deepti Ghadiyaram, Du Tran, and Dhruv Mahajan · 2019
Closest in time.
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman · 2019
Closest in time.
Data-efficient image recognition with contrastive predictive coding
Olivier J Hénaff, Ali Razavi, Carl Doersch, SM Eslami, and Aaron van den Oord · 2019
Closest in time.
A case study on combining asr and visual features for generating instructional video captions
Jack Hessel, Bo Pang, Zhenhai Zhu, and Radu Soricut · 2019
Closest in time.
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, and In So Kweon · 2019
Closest in time.
Mining youtube-a dataset for learning fine-grained action concepts from webly supervised video data
Hilde Kuehne, Ahsan Iqbal, Alexander Richard, and Juergen Gall · 2019
Closest in time.
Howto100M: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Closest in time.
Grounding object detections with transcriptions
Yasufumi Moriya, Ramon Sanabria, Florian Metze, and Gareth JF Jones · 2019
Closest in time.
Multimodal abstractive summarization for how2 videos
Shruti Palaskar, Jindrich Libovickỳ, Spandana Gella, and Florian Metze · 2019
Closest in time.
Contrastive bidirectional transformer for temporal representation learning
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Closest in time.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Closest in time.
Coin: A large-scale dataset for comprehensive instructional video analysis
Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou · 2019
Closest in time.
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Closest in time.
Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Yunhui Liu, and Wei Liu · 2019
Closest in time.
Fine-grained action retrieval through multiple parts-of-speech embeddings
Michael Wray, Diane Larlus, Gabriela Csurka, and Dima Damen · 2019
Closest in time.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Closest in time.
Cross-task weakly supervised learning from instructional videos
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic · 2019
Closest in time.