Fetching the paper…
Reading the bibliography…
One of the key factors of enabling machine learning models to comprehend and solve real-world tasks is to leverage multimodal data.
Contrastive bidirectional transformer for temporal representation learning
Sun, C.; Baradel, F.; Murphy, K.; and Schmid, C. 2019a · 1906
Earlier work this paper cites.
Tian, Y.; Krishnan, D.; and Isola, P. 2019 · 1906
Earlier work this paper cites.
Learning from noisy labels with distillation
Li, Y.; Yang, J.; Song, Y.; Cao, L.; Luo, J.; and Li, L.-J. 2017 · 1918
Earlier work this paper cites.
Polysemous visual-semantic embedding for cross-modal retrieval
Song, Y.; and Soleymani, M. 2019 · 1988
Earlier work this paper cites.
Variable kernel density estimation
Terrell, G. R.; and Scott, D. W. 1992 · 1992
Earlier work this paper cites.
The use of polynomial splines and their tensor products in multivariate function estimation
Stone, C. J. 1994 · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Heterogeneous memory enhanced multimodal attention model for video question answering
Fan, C.; Zhang, X.; Zhang, S.; Wang, W.; Zhang, C.; and Huang, H. 2019 · 2007
Earlier work this paper cites.
The EM algorithm and extensions , volume 382
McLachlan, G. J.; and Krishnan, T. 2007 · 2007
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Koller, D.; and Friedman, N. 2009 · 2009
Earlier work this paper cites.
Deep boltzmann machines
Salakhutdinov, R.; and Hinton, G. 2009 · 2009
Earlier work this paper cites.
Large scale image annotation: learning to rank with joint word-image embeddings
Weston, J.; Bengio, S.; and Usunier, N. 2010 · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D. L.; and Dolan, W. B. 2011 · 2011
Earlier work this paper cites.
Learning representations for multimodal data with deep belief nets
Srivastava, N.; and Salakhutdinov, R. 2012 · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Frome, A.; Corrado, G. S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G. S.; and Dean, J. 2013 · 2013
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Mao, J.; Xu, W.; Yang, Y.; Wang, J.; Huang, Z.; and Yuille, A. 2014 · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Socher, R.; Karpathy, A.; Le, Q. V.; Manning, C. D.; and Ng, A. Y. 2014 · 2014
Earlier work this paper cites.
A deep and tractable density estimator
Uria, B.; Murray, I.; and Larochelle, H. 2014 · 2014
Earlier work this paper cites.
Learning fine-grained image similarity with deep ranking
Wang, J.; Song, Y.; Leung, T.; Rosenberg, C.; Wang, J.; Philbin, J.; Chen, B.; and Wu, Y. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Training deep neural networks on noisy labels with bootstrapping
Reed, S.; Lee, H.; Anguelov, D.; Szegedy, C.; Erhan, D.; and Rabinovich, A. 2015 · 2015
Earlier work this paper cites.
The long-short story of movie description
Rohrbach, A.; Rohrbach, M.; and Schiele, B. 2015 · 2015
Cited alongside, same era.
A dataset for movie description
Rohrbach, A.; Rohrbach, M.; Tandon, N.; and Schiele, B. 2015 · 2015
Cited alongside, same era.
Facenet: A unified embedding for face recognition and clustering
Schroff, F.; Kalenichenko, D.; and Philbin, J. 2015 · 2015
Cited alongside, same era.
Unsupervised learning of video representations using lstms
Srivastava, N.; Mansimov, E.; and Salakhudinov, R. 2015 · 2015
Cited alongside, same era.
Training convolutional networks with noisy labels
Sukhbaatar, S.; Estrach, J. B.; Paluri, M.; Bourdev, L.; and Fergus, R. 2015 · 2015
Cited alongside, same era.
Learning with symmetric label noise: The importance of being unhinged
Van Rooyen, B.; Menon, A.; and Williamson, R. C. 2015 · 2015
Cited alongside, same era.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Jang, Y.; Song, Y.; Yu, Y.; Kim, Y.; and Kim, G. 2017 · 2017
Later among the works it cites.
Masked autoregressive flow for density estimation
Papamakarios, G.; Pavlakou, T.; and Murray, I. 2017 · 2017
Later among the works it cites.
Video question answering via gradually refined attention over appearance and motion
Xu, D.; Zhao, Z.; Xiao, J.; Wu, F.; Zhang, H.; He, X.; and Zhuang, Y. 2017 · 2017
Later among the works it cites.
Motion-appearance co-memory networks for video question answering
Gao, J.; Ge, R.; Chen, K.; and Nevatia, R. 2018 · 2018
Later among the works it cites.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Hara, K.; Kataoka, H.; and Satoh, Y. 2018 · 2018
Later among the works it cites.
Learning answer embeddings for visual question answering
Hu, H.; Chao, W.-L.; and Sha, F. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Translating Videos to Natural Language Using Deep Recurrent Neural Networks
Venugopalan, S.; Xu, H.; Donahue, J.; Rohrbach, M.; Mooney, R.; and Saenko, K. 2015 · 2015
Cited alongside, same era.
Nonparametric multivariate density estimation using mixtures
Wang, X.; and Wang, Y. 2015 · 2015
Cited alongside, same era.
Learning from massive noisy labeled data for image classification
Xiao, T.; Xia, T.; Yang, Y.; Huang, C.; and Wang, X. 2015 · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Shuffle and learn: unsupervised learning using temporal order verification
Misra, I.; Zitnick, C. L.; and Hebert, M. 2016 · 2016
Cited alongside, same era.
Unsupervised learning of visual representations by solving jigsaw puzzles
Noroozi, M.; and Favaro, P. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels
Jiang, L.; Zhou, Z.; Leung, T.; Li, L.; and Fei-Fei, L. 2018 · 2018
Later among the works it cites.
Cooperative learning of audio and video models from self-supervised synchronization
Korbar, B.; Tran, D.; and Torresani, L. 2018 · 2018
Later among the works it cites.
Learning a text-video embedding from incomplete and heterogeneous data
Miech, A.; Laptev, I.; and Sivic, J. 2018 · 2018
Later among the works it cites.
Learning joint embedding with multimodal cues for cross-modal video-text retrieval
Mithun, N. C.; Li, J.; Metze, F.; and Roy-Chowdhury, A. K. 2018 · 2018
Later among the works it cites.
Density estimation for statistics and data analysis
Silverman, B. W. 2018 · 2018
Later among the works it cites.
Joint optimization framework for learning with noisy labels
Tanaka, D.; Ikami, D.; Yamasaki, T.; and Aizawa, K. 2018 · 2018
Later among the works it cites.
Learning and using the arrow of time
Wei, D.; Lim, J. J.; Zisserman, A.; and Freeman, W. T. 2018 · 2018
Later among the works it cites.
A joint sequence fusion model for video question answering and retrieval
Yu, Y.; Kim, J.; and Kim, G. 2018 · 2018
Later among the works it cites.
Learning to Detect and Retrieve Objects From Unlabeled Videos
Amrani, E.; Ben-Ari, R.; Hakim, T.; and Bronstein, A. 2019 · 2019
Later among the works it cites.
Billion-scale similarity search with GPUs
Johnson, J.; Douze, M.; and Jégou, H. 2019 · 2019
Later among the works it cites.
Learning to learn from noisy labeled data
Li, J.; Wong, Y.; Zhao, Q.; and Kankanhalli, M. S. 2019 · 2019
Later among the works it cites.
Use What You Have: Video Retrieval Using Representations From Collaborative Experts
Liu, Y.; Albanie, S.; Nagrani, A.; and Zisserman, A. 2019 · 2019
Later among the works it cites.
Howto100M: Learning a text-video embedding by watching hundred million narrated video clips
Miech, A.; Zhukov, D.; Alayrac, J.-B.; Tapaswi, M.; Laptev, I.; and Sivic, J. 2019 · 2019
Later among the works it cites.
Grounding Object Detections With Transcriptions
Moriya, Y.; Sanabria, R.; Metze, F.; and Jones, G. J. 2019 · 2019
Later among the works it cites.
Nonparametric density estimation for high-dimensional data—Algorithms and applications
Wang, Z.; and Scott, D. W. 2019 · 2019
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
Miech, A.; Alayrac, J.-B.; Smaira, L.; Laptev, I.; Sivic, J.; and Zisserman, A. 2020 · 2020
Closest in time.