Fetching the paper…
Reading the bibliography…
Recent advances in representation learning have demonstrated an ability to represent information from different modalities such as video, text, and audio in a single high-level embedding vector.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M. Gutmann and A. Hyvärinen · 2010
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Y. Bengio, N. Léonard, and A. Courville · 2013
Earlier work this paper cites.
Learning words from images and speech
G. Synnaeve, M. Versteegh, and E. Dupoux · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Earlier work this paper cites.
Deep multimodal semantic embeddings for speech and images
D. Harwath and J. Glass · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
F. Schroff, D. Kalenichenko, and J. Philbin · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier · 2017
Earlier work this paper cites.
Learning word-like units from joint audio-visual analysis
D. Harwath and J. R. Glass · 2017
Earlier work this paper cites.
Unsupervised learning of spoken language with visual context
D. Harwath, A. Torralba, and J. R. Glass · 2017
Earlier work this paper cites.
Learning modality-invariant representations for speech and images
K. Leidal, D. Harwath, and J. Glass · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. v. d. Oord, O. Vinyals, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Jointly discovering visual objects and spoken words from raw sensory input
D. Harwath, A. Recasens, D. Surís, G. Chuang, A. Torralba, and J. Glass · 2018
Earlier work this paper cites.
Semantic speech retrieval with a visually grounded model of untranscribed speech
H. Kamper, G. Shakhnarovich, and K. Livescu · 2018
Cited alongside, same era.
Learning a text-video embedding from incomplete and heterogeneous data
A. Miech, I. Laptev, and J. Sivic · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
A joint sequence fusion model for video question answering and retrieval
Y. Yu, J. Kim, and G. Kim · 2018
Cited alongside, same era.
Grounding spoken words in unlabeled video
A. W. Boggust, K. Audhkhasi, D. Joshi, D. Harwath, S. Thomas, R. S. Feris, D. Gutfreund, Y. Zhang, A. Torralba, M. Picheny, et al · 2019
Cited alongside, same era.
Learning hierarchical discrete linguistic units from visually-grounded speech
D. Harwath, W.-N. Hsu, and J. Glass · 2020
Later among the works it cites.
Jointly discovering visual objects and spoken words from raw sensory input
D. Harwath, A. Recasens, D. Surís, G. Chuang, A. Torralba, and J. Glass · 2020
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman · 2020
Later among the works it cites.
Support-set bottlenecks for video-text representation learning
M. Patrick, P.-Y. Huang, Y. Asano, F. Metze, A. Hauptmann, J. Henriques, and A. Vedaldi · 2020
Later among the works it cites.
Avlnet: Learning audio-visual language representations from instructional videos
A. Rouditchenko, A. Boggust, D. Harwath, D. Joshi, S. Thomas, K. Audhkhasi, R. Feris, B. Kingsbury, M. Picheny, A. Torralba, et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Ilharco, Y. Zhang, and J. Baldridge · 2019
Cited alongside, same era.
Tsm: Temporal shift module for efficient video understanding
J. Lin, C. Gan, and S. Han · 2019
Cited alongside, same era.
Improving referring expression grounding with cross-modal attention-guided erasing
X. Liu, Z. Wang, J. Shao, X. Wang, and H. Li · 2019
Cited alongside, same era.
Use what you have: Video retrieval using representations from collaborative experts
Y. Liu, S. Albanie, A. Nagrani, and A. Zisserman · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic · 2019
Cited alongside, same era.
Multi-moments in time: Learning and interpreting models for multi-action video understanding
M. Monfort, K. Ramakrishnan, A. Andonian, B. A. McNamara, A. Lascelles, B. Pan, Q. Fan, D. Gutfreund, R. Feris, and A. Oliva · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
Later among the works it cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
M. Bain, A. Nagrani, G. Varol, and A. Zisserman · 2021
Closest in time.
Teachtext: Crossmodal generalized distillation for text-video retrieval
I. Croitoru, S.-V. Bogolin, Y. Liu, S. Albanie, M. Leordeanu, H. Jin, and A. Zisserman · 2021
Closest in time.
Mdmmt: Multidomain multimodal transformer for video retrieval
M. Dzabraev, M. Kalashnikov, S. Komkov, and A. Petiushko · 2021
Closest in time.
Less is more: Clipbert for video-and-language learning via sparse sampling
J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu · 2021
Closest in time.
Hit: Hierarchical transformer with momentum contrast for video-text retrieval
S. Liu, H. Fan, S. Qian, Y. Chen, W. Ding, and Z. Wang · 2021
Closest in time.
Clip4clip: An empirical study of clip for end to end video clip retrieval
H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, and T. Li · 2021
Closest in time.
Spoken moments: Learning joint audio-visual representations from video descriptions
M. Monfort, S. Jin, A. Liu, D. Harwath, R. Feris, J. Glass, and A. Oliva · 2021
Closest in time.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Closest in time.
Talk, don’t write: A study of direct speech-based image retrieval
R. Sanabria, A. Waters, and J. Baldridge · 2021
Closest in time.
Learning audio-visual correlations from variational cross-modal generation
Y. Zhu, Y. Wu, H. Latapie, Y. Yang, and Y. Yan · 2021
Closest in time.