Fetching the paper…
Reading the bibliography…
Learning text-video embeddings usually requires a dataset of video clips with manually provided captions.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
A multi-view embedding space for modeling internet images, tags, and their semantics
Y. Gong, Q. Ke, M. Isard, and S. Lazebnik · 2014
Earlier work this paper cites.
Improving image-sentence embeddings using large weakly annotated photo collections
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and F. F. F. Li · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Earlier work this paper cites.
Instructional videos for unsupervised harvesting and learning of action examples
S.-I. Yu, L. Jiang, and A. Hauptmann · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors
B. Klein, G. Lev, G. Sadeh, and L. Wolf · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Earlier work this paper cites.
What’s cookin’? interpreting cooking videos using text, speech and vision
J. Malmaud, J. Huang, V. Rathod, N. Johnston, A. Rabinovich, and K. Murphy · 2015
Earlier work this paper cites.
A dataset for movie description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Earlier work this paper cites.
Unsupervised semantic parsing of video collections
O. Sener, A. R. Zamir, S. Savarese, and A. Saxena · 2015
Earlier work this paper cites.
Using descriptive video services to create a large data source for video annotation research
A. Torabi, C. Pal, H. Larochelle, and A. Courville · 2015
Earlier work this paper cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
R. Xu, C. Xiong, W. Chen, and J. J. Corso · 2015
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
J.-B. Alayrac, P. Bojanowski, N. Agrawal, I. Laptev, J. Sivic, and S. Lacoste-Julien · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Connectionist temporal modeling for weakly supervised action labeling
D.-A. Huang, L. Fei-Fei, and J. C. Niebles · 2016
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Earlier work this paper cites.
TGIF: A New Dataset and Benchmark on Animated GIF Description
Y. Li, Y. Song, L. Cao, J. Tetreault, L. Goldberg, A. Jaimes, and J. Luo · 2016
Earlier work this paper cites.
Hierarchical recurrent neural encoder for video representation with application to captioning
P. Pan, Z. Xu, Y. Yang, F. Wu, and Y. Zhuang · 2016
Earlier work this paper cites.
Jointly modeling embedding and translation to bridge video and language
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Cited alongside, same era.
Learning language-visual embedding for movie understanding with natural-language
A. Torabi, N. Tandon, and L. Sigal · 2016
Cited alongside, same era.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Cited alongside, same era.
Image captioning with semantic attention
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Cited alongside, same era.
Video paragraph captioning using hierarchical recurrent neural networks
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
K. Hara, H. Kataoka, and Y. Satoh · 2018
Later among the works it cites.
Jointly discovering visual objects and spoken words from raw sensory input
D. Harwath, A. Recasens, D. Surís, G. Chuang, A. Torralba, and J. Glass · 2018
Later among the works it cites.
Finding ”it”: Weakly-supervised reference-aware visual grounding in instructional video
D.-A. Huang, V. Ramanathan, D. Mahajan, L. Torresani, M. Paluri, L. Fei-Fei, and J. C. Niebles · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
K. L. Jacob Devlin, Ming-Wei Chang and K. Toutanova · 2018
Later among the works it cites.
Exploring the limits of weakly supervised pretraining
D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. van der Maaten · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu · 2016
Cited alongside, same era.
Video captioning and retrieval models with semantic attention
Y. Yu, H. Ko, J. Choi, and G. Kim · 2016
Cited alongside, same era.
Joint discovery of object states and manipulation actions
J.-B. Alayrac, J. Sivic, I. Laptev, and S. Lacoste-Julien · 2017
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
Discover and learn new objects from documentaries
K. Chen, H. Song, C. Change Loy, and D. Lin · 2017
Cited alongside, same era.
Localizing moments in video with natural language
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell · 2017
Cited alongside, same era.
Unsupervised visual-linguistic reference resolution in instructional videos
D.-A. Huang, J. J. Lim, L. Fei-Fei, and J. C. Niebles · 2017
Cited alongside, same era.
A. Miech, I. Laptev, and J. Sivic · 2018
Later among the works it cites.
Learning joint embedding with multimodal cues for cross-modal video-text retrieval
N. C. Mithun, J. Li, F. Metze, and A. K. Roy-Chowdhury · 2018
Later among the works it cites.
Improving Language Understandingby Generative Pre-Training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Later among the works it cites.
Action sets: Weakly supervised action segmentation without ordering constraints
A. Richard, H. Kuehne, and J. Gall · 2018
Later among the works it cites.
How2: a large-scale dataset for multimodal language understanding
R. Sanabria, O. Caglayan, S. Palaskar, D. Elliott, L. Barrault, L. Specia, and F. Metze · 2018
Later among the works it cites.
Unsupervised learning and segmentation of complex activities from video
F. Sener and A. Yao · 2018
Later among the works it cites.
Learning two-branch neural networks for image-text matching tasks
L. Wang, Y. Li, J. Huang, and S. Lazebnik · 2018
Later among the works it cites.
Learning to compose topic-aware mixture of experts for zero-shot video captioning
X. Wang, J. Wu, D. Zhang, Y. Su, and W. Y. Wang · 2018
Later among the works it cites.
A joint sequence fusion model for video question answering and retrieval
Y. Yu, J. Kim, and G. Kim · 2018
Later among the works it cites.
Towards automatic learning of procedures from web instructional videos
L. Zhou, X. Chenliang, and J. J. Corso · 2018
Later among the works it cites.
Towards automatic learning of procedures from web instructional videos
L. Zhou, C. Xu, and J. J. Corso · 2018
Later among the works it cites.
https://www.di.ens.fr/willow/research/howto100m/ , 2019
Project webpage · 2019
Closest in time.
Dual encoding for zero-example video retrieval
J. Dong, X. Li, C. Xu, S. Ji, Y. He, G. Yang, and X. Wang · 2019
Closest in time.
Howto100M: Learning a text-video embedding by watching hundred million narrated video clips
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic · 2019
Closest in time.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Closest in time.
Coin: A large-scale dataset for comprehensive instructional video analysis
Y. Tang, D. Ding, Y. Rao, Y. Zheng, D. Zhang, L. Zhao, J. Lu, and J. Zhou · 2019
Closest in time.
Cross-task weakly supervised learning from instructional videos
D. Zhukov, J.-B. Alayrac, R. G. Cinbis, D. Fouhey, I. Laptev, and J. Sivic · 2019
Closest in time.