Fetching the paper…
Reading the bibliography…
This paper presents an augmentation of MSCOCO dataset where speech is added to image and text.
H. Bortfeld, S. Leon, J. Bloom, M. Schober, and S. Brennan, “Disfluency Rates in Conversation: Effects of Age, Relationship, Topic, Role, and Gender,”
2001
Earlier work this paper cites.
D. Roy, “Grounded spoken language acquisition: experiments in word learning,”
2003
Earlier work this paper cites.
D. Schwarz, “Corpus-based concatenative synthesis : Assembling sounds by content-based selection of units from large sound databases,”
2007
Earlier work this paper cites.
A. Jansen and B. Van Durme, “Efficient spoken term discovery using randomized algorithms,” in
2011
Earlier work this paper cites.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,”
2013
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,”
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in
2014
Earlier work this paper cites.
B. Ludusan, M. Versteegh, A. Jansen, G. Gravier, X.-N. Cao, M. Johnson, and E. Dupoux, “Bridging the gap between speech technology and natural language processing: an evaluation toolbox for term discovery systems,” May 2014. [Online]. Available:
2014
Cited alongside, same era.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-Based Models for Speech Recognition,” in
2015
Cited alongside, same era.
A. Karpathy and F. Li, “Deep visual-semantic alignments for generating image descriptions,” in
2015
Cited alongside, same era.
L. Wang, Y. Li, and S. Lazebnik, “Learning deep structure-preserving image-text embeddings,”
2015
Cited alongside, same era.
D. Harwath and J. Glass, “Deep multimodal semantic embeddings for speech and images,” in
2015
2016
Later among the works it cites.
2016
Later among the works it cites.
T. Miyazaki and N. Shimizu, “Cross-lingual image caption generation,” in
2016
Later among the works it cites.
D. F. Harwath and J. R. Glass, “Learning word-like units from joint audio-visual analysis,”
2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2015
Cited alongside, same era.
R. Bernardi, R. Cakici, D. Elliott, A. Erdem, E. Erdem, N. Ikizler-Cinbis, F. Keller, A. Muscat, and B. Plank, “Automatic description generation from images: A survey of models, datasets, and evaluation measures,”
2016
Cited alongside, same era.
2017
Closest in time.
G. Chrupała, L. Gelderloos, and A. Alishahi, “Representations of language in a model of visually grounded speech signal,” 2017, arXiv170201991 Cs
2017
Closest in time.