Fetching the paper…
Reading the bibliography…
Joint understanding of video and language is an active research area with many applications.
Hierarchical mixtures of experts and the em algorithm
Jordan, M.I.: · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., Schmidhuber, J.: · 1997
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: · 2009
Earlier work this paper cites.
Online learning for matrix factorization and sparse coding
Mairal, J., Bach, F., Ponce, J., Sairo, G.: · 2010
Earlier work this paper cites.
Product Quantization for Nearest Neighbor Search
Jégou, H., Douze, M., Schmid, C.: · 2011
Earlier work this paper cites.
Action Recognition with Improved Trajectories
Wang, H., Schmid, C.: · 2013
Earlier work this paper cites.
Low-rank matrix completion using alternating minimization
Jain, P., Netrapalli, P., Sanghavi, S.: · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., Dean, J.: · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: · 2014
Earlier work this paper cites.
A multi-view embedding space for modeling internet images, tags, and their semantics
Gong, Y., Ke, Q., Isard, M., Lazebnik, S.: · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Simonyan, K., Zisserman, A.: · 2014
Earlier work this paper cites.
On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
Cho, K., van Merrienboer, B., Bahdanau, D., Bengio, Y.: · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Karpathy, A., Joulin, A., Li, F.F.F.: · 2014
Earlier work this paper cites.
Weakly-supervised alignment of video with text
Bojanowski, P., Lajugie, R., Grave, E., Bach, F., Laptev, I., Ponce, J., Schmid, C.: · 2015
Earlier work this paper cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
Xu, R., Xiong, C., Chen, W., Corso, J.J.: · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: · 2015
Earlier work this paper cites.
A dataset for movie description
Rohrbach, A., Rohrbach, M., Tandon, N., Schiele, B.: · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B.A., Wang, L., Cervantes, C.M., Caicedo, J.C., Hockenmaier, J., Lazebnik, S.: · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors
Klein, B., Lev, G., Sadeh, G., Wolf, L.: · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
Malinowski, M., Rohrbach, M., Fritz, M.: · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D.P., Ba, J.: · 2015
Cited alongside, same era.
Movieqa: Understanding stories in movies through question-answering
Tapaswi, M., Zhu, Y., Stiefelhagen, R., Torralba, A., Urtasun, R., Fidler, S.: · 2016
Cited alongside, same era.
Unsupervised learning from narrated instruction videos
Alayrac, J.B., Bojanowski, P., Agrawal, N., Laptev, I., Sivic, J., Lacoste-Julien, S.: · 2016
Cited alongside, same era.
Jointly modeling embedding and translation to bridge video and language
Pan, Y., Mei, T., Yao, T., Li, H., Rui, Y.: · 2016
Cited alongside, same era.
Video paragraph captioning using hierarchical recurrent neural networks
Yu, H., Wang, J., Huang, Z., Yang, Y., Xu, W.: · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., Rui, Y.: · 2016
Learning from Video and Text via Large-Scale Discriminative Clustering
Miech, A., Alayrac, J.B., Bojanowski, P., Laptev, I., Sivic, J.: · 2017
Later among the works it cites.
Localizing moments in video with natural language
Hendricks, L.A., Wang, O., Shechtman, E., Sivic, J., Darrell, T., Russell, B.: · 2017
Later among the works it cites.
Enhancing video summarization via vision-language embedding
Plummer, B.A., Brown, M., Lazebnik, S.: · 2017
Later among the works it cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J., Zisserman, A.: · 2017
Later among the works it cites.
Sampling matters in deep embedding learning
Wu, C.Y., Manmatha, R., Smola, A.J., Krähenbühl, P.: · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., Bernstein, M., Fei-Fei, L.: · 2016
Cited alongside, same era.
Learning deep structure-preserving image-text embeddings
Wang, L., Li, Y., Lazebnik, S.: · 2016
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning
Johnson, J., Karpathy, A., Fei-Fei, L.: · 2016
Cited alongside, same era.
Hierarchical recurrent neural encoder for video representation with application to captioning
Pan, P., Xu, Z., Yang, Y., Wu, F., Zhuang, Y.: · 2016
Cited alongside, same era.
Image captioning with semantic attention
You, Q., Jin, H., Wang, Z., Fang, C., Luo, J.: · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., Rohrbach, M.: · 2016
Cited alongside, same era.
Aytar, Y., Vondrick, C., Torralba, A.: · 2017
Later among the works it cites.
Phrase localization and visual relationship detection with comprehensive image-language cues
Plummer, B.A., Mallya, A., Cervantes, C.M., Hockenmaier, J., Lazebnik, S.: · 2017
Later among the works it cites.
Spatiotemporal multiplier networks for video action recognition
Feichtenhofer, C., Pinz, A., Wildes, R.P.: · 2017
Later among the works it cites.
Learnable pooling with context gating for video classification
Miech, A., Laptev, I., Sivic, J.: · 2017
Later among the works it cites.
Actionvlad: Learning spatio-temporal aggregation for action classification
Girdhar, R., Ramanan, D., Gupta, A., Sivic, J., Russell, B.: · 2017
Later among the works it cites.
Long-term Temporal Convolutions for Action Recognition
Varol, G., Laptev, I., Schmid, C.: · 2017
Later among the works it cites.
Look, listen and learn
Arandjelovic, R., Zisserman, A.: · 2017
Later among the works it cites.
Semantic image inpainting with deep generative models
Yeh, R.A., Chen, C., Lim, T.Y., Schwing, A.G., Hasegawa-Johnson, M., Do, M.N.: · 2017
Later among the works it cites.
Ulyanov, D., Vedaldi, A., Lempitsky, V.: · 2017
Later among the works it cites.
Ubernet: Training a universal convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory
Kokkinos, I.: · 2017
Later among the works it cites.
Missing modalities imputation via cascaded residual autoencoder
Tran, L., Liu, X., Zhou, J., Jin, R.: · 2017
Later among the works it cites.
CNN architectures for large-scale audio classification
Hershey, S., Chaudhuri, S., Ellis, D.P.W., Gemmeke, J.F., Jansen, A., Moore, C., Plakal, M., Platt, D., Saurous, R.A., Seybold, B., Slaney, M., Weiss, R., Wilson, K.: · 2017
Later among the works it cites.
Learning two-branch neural networks for image-text matching tasks
Wang, L., Li, Y., Huang, J., Lazebnik, S.: · 2018
Closest in time.
A joint sequence fusion model for video question answering and retrieval
Yu, Y., Kim, J., Kim, G.: · 2018
Closest in time.