Fetching the paper…
Reading the bibliography…
Speech-based image retrieval has been studied as a proxy for joint representation learning, usually without emphasis on retrieval itself.
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier, “Collecting image annotations using Amazon’s Mechanical Turk,” in
2010
Earlier work this paper cites.
G. Synnaeve, M. Versteegh, and E. Dupoux, “Learning words from images and speech,” in
2014
Earlier work this paper cites.
D. Harwath and J. Glass, “Deep multimodal semantic embeddings for speech and images,” in
2015
Earlier work this paper cites.
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt
2015
Earlier work this paper cites.
A. Karpathy and F.-F. Li, “Deep visual-semantic alignments for generating image descriptions,” in
2015
Earlier work this paper cites.
L. Gelderloos and G. Chrupała, “From phonemes to images: levels of representation in a recurrent neural model of visually-grounded language learning,” in
2016
Earlier work this paper cites.
D. Harwath, A. Torralba, and J. Glass, “Unsupervised learning of spoken language with visual context,” in
2016
Earlier work this paper cites.
G. Chrupała, L. Gelderloos, and A. Alishahi, “Representations of language in a model of visually grounded speech signal,” in
2017
Earlier work this paper cites.
H. Kamper, S. Settle, G. Shakhnarovich, and K. Livescu, “Visually grounded learning of keyword prediction from untranscribed speech,” in
2017
Earlier work this paper cites.
D. Harwath and J. R. Glass, “Learning word-like units from joint audio-visual analysis,” in
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in
2017
Earlier work this paper cites.
D. Harwath, A. Recasens, D. Surís, G. Chuang, A. Torralba, and J. Glass, “Jointly discovering visual objects and spoken words from raw sensory input,” in
2018
Earlier work this paper cites.
H. Kamper and M. Roth, “Visually grounded cross-lingual keyword spotting in speech,” in
2018
Cited alongside, same era.
D. Gillick, A. Presta, and G. S. Tomar, “End-to-end retrieval in continuous space,”
2018
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in
2018
Cited alongside, same era.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in
2018
Cited alongside, same era.
D. Harwath, W.-N. Hsu, and J. Glass, “Learning hierarchical discrete linguistic units from visually-grounded speech,” in
2019
Cited alongside, same era.
G. Chrupała, “Symbolic inductive bias for visually grounded learning of spoken language,”
2019
Later among the works it cites.
D. Merkx and S. L. Frank, “Learning semantic sentence representations from visually grounded language without lexical knowledge,”
2019
Later among the works it cites.
K. Li, Y. Zhang, K. Li, Y. Li, and Y. Fu, “Visual semantic reasoning for image-text matching,” in
2019
Later among the works it cites.
B. Higy, D. Eliott, and G. Chrupała, “Textual supervision for visually grounded spoken language understanding,” in
2020
Later among the works it cites.
Y. Ohishi, A. Kimura, T. Kawanishi, K. Kashino, D. Harwath, and J. Glass, “Trilingual semantic embeddings of visually grounded speech with self-attention mechanisms,” in
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Harwath and J. Glass, “Towards visually grounded sub-word speech unit discovery,” in
2019
Cited alongside, same era.
W. N. Havard, J.-P. Chevrot, and L. Besacier, “Word recognition, competition, and activation in a model of visually grounded speech,”
2019
Cited alongside, same era.
G. Ilharco, Y. Zhang, and J. Baldridge, “Large-scale representation learning from visually grounded untranscribed speech,” in
2019
Cited alongside, same era.
E. Azuh, D. Harwath, and J. R. Glass, “Towards bilingual lexicon discovery from visually grounded speech audio.” in
2019
Cited alongside, same era.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,”
2019
Cited alongside, same era.
M. Tan and Q. V. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in
2019
Cited alongside, same era.
D. Merkx, S. L. Frank, and M. Ernestus, “Language learning using speech to image retrieval,” in
2019
Cited alongside, same era.
J. Pont-Tuset, J. Uijlings, S. Changpinyo, R. Soricut, and V. Ferrari, “Connecting vision and language with localized narratives,” in
2020
Later among the works it cites.
K. Kawakami, L. Wang, C. Dyer, P. Blunsom, and A. van den Oord, “Unsupervised learning of efficient and robust speech representations,” in
2020
Later among the works it cites.
D. Harwath, A. Recasens, D. Surís, G. Chuang, A. Torralba, and J. Glass, “Jointly discovering visual objects and spoken words from raw sensory input,” in
2020
Later among the works it cites.
W. Havard, L. Besacier, and J.-P. Chevrot, “Catplayinginthesnow: Impact of prior segmentation on a model of visually grounded speech,” in
2020
Later among the works it cites.
Z. Parekh, J. Baldridge, D. Cer, A. Waters, and Y. Yang, “Crisscrossed Captions: Extended intramodal and intermodal semantic similarity judgments for MS-COCO,”
2020
Later among the works it cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in
2020
Later among the works it cites.
A. Rouditchenko, A. Boggust, D. Harwath, D. Joshi, S. Thomas, K. Audhkhasi, R. Feris, B. Kingsbury, M. Picheny, A. Torralba, and J. Glass, “AVLnet: Learning audio-visual language representations from instructional videos,” in
2021
Closest in time.