Fetching the paper…
Reading the bibliography…
We present the visually-grounded language modelling track that was introduced in the Zero-Resource Speech challenge, 2021 edition, 2nd round.
D. Roy and A. Pentland, “Learning words from sights and sounds: a computational model,” Cognitive Science , vol. 26, pp. 113–146, 2002
2002
Earlier work this paper cites.
C. Yu and D. Ballard, “A multimodal learning interface for grounding spoken language in sensory perceptions,” ACM Transactions on Applied Perceptions , vol. 1, pp. 57–80, 2004
2004
Earlier work this paper cites.
L. ten Bosch, H. Van hamme, L. Boves, and R. K. Moore, “A computational model of language acquisition: the emergence of words,” Fundamenta Informaticae , vol. 90, pp. 229–249, 2008
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition , Jun. 2009, pp. 248–255, iSSN: 1063-6919
2009
Earlier work this paper cites.
J. Driesen and H. Van hamme, “Modeling vocabulary acquisition, adaptation and generalization in infants using adaptive Bayesian PLSA,” Neurocomputing , vol. 74, pp. 1874–1882, 2011
2011
Earlier work this paper cites.
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng, “Grounded compositional semantics for finding and describing images with sentences,” Transactions of the Association for Computational Linguistics , vol. 2, pp. 207–218, 2014
2014
Earlier work this paper cites.
G. Synnaeve, M. Versteegh, and E. Dupoux, “Learning words from images and speech,” in 28th Conference on Neural Information Processing Systems (NIPS) Workshop on Learning Semantics , December 8–13, 2014, Montreal, Canada., 2014
2014
Earlier work this paper cites.
K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Doha, Qatar: Association for Computational Linguistics, 2014, pp. 1724–1734
2014
Earlier work this paper cites.
O. Räsänen and H. Rasilo, “A joint model of word segmentation and meaning acquisition through cross-situational learning,” Psychological Review , vol. 122, pp. 792–829, 2015
2015
Earlier work this paper cites.
O. Mangin, D. Filliat, L. ten Bosch, and P.-Y. Oudeyer, “MCA-NMF: Multimodal concept acquisition with non-negative matrix factorization,” PLOS One , 2015
2015
Earlier work this paper cites.
A. Karpathy and F. Li, “Deep visual-semantic alignments for generating image descriptions,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2015) , June 7–12, 2015, Boston, MA, pp. 3128–3137., 2015, pp. 3128–3137
2015
Earlier work this paper cites.
D. Harwath and J. Glass, “Deep multimodal semantic embeddings for speech and images,” in 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) . IEEE, 2015, pp. 237–244
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
M. VanDam, “HomeBank VanDam Public 5-minute Corpus,” 2015, type: dataset. [Online]. Available: http://homebank.talkbank.org/access/Public/VanDam-5minute.html
2015
Cited alongside, same era.
——, “HomeBank VanDam Public Daylong Corpus,” 2015, type: dataset. [Online]. Available: http://homebank.talkbank.org/access/Public/VanDam-Daylong.html
2015
Cited alongside, same era.
D. F. Harwath, A. Torralba, and J. R. Glass, “Unsupervised learning of spoken language with visual context,” in Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems (NIPS 2016) , December 5–10, 2016, Barcelona, Spain, pp. 1858–1866., 2016, pp. 1858–1866
2016
Cited alongside, same era.
M. Versteegh, X. Anguera, A. Jansen, and E. Dupoux, “The zero resource speech challenge 2015: Proposed approaches and results,” Procedia Computer Science , vol. 81, pp. 67–72, 12 2016
2016
Cited alongside, same era.
2019
Later among the works it cites.
W. N. Havard, J. Chevrot, and L. Besacier, “Models of visually grounded speech signal pay attention to nouns: A bilingual experiment on english and japanese,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2019) , May 12–17, 2019, Brighton, UK, pp. 8618–8622., 2019, pp. 8618–8622
2019
Later among the works it cites.
——, “Word recognition, competition, and activation in a model of visually grounded speech,” in Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL 2019) , November 3–4, 2019, Hong Kong, China, pp. 339–348., 2019, pp. 339–348
2019
Later among the works it cites.
Y. A. Chung, W. N. Hsu, H. Tang, and J. Glass, “An unsupervised autoregressive model for speech representation learning,” Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH , pp. 146–150, 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition . Seattle, WA, USA: IEEE, Jun. 2016, pp. 770–778
2016
Cited alongside, same era.
G. Chrupała, L. Gelderloos, and A. Alishahi, “Representations of language in a model of visually grounded speech signal,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , July 30–August 4, 2017, Vancover, Canada, pp. 613–622, 2017, pp. 613–622
2017
Cited alongside, same era.
A. Alishahi, M. Barking, and G. Chrupała, “Encoding of phonology in a recurrent neural model of grounded speech,” in Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017) , August 3–4, 2017, Vancouver, Canada, pp. 368–378, 2017, pp. 368–378
2017
Cited alongside, same era.
E. Dunbar, X. N. Cao, J. Benjumea, J. Karadayi, M. Bernard, L. Besacier, X. Anguera, and E. Dupoux, “The zero resource speech challenge 2017,” 2017
2017
Cited alongside, same era.
2018
Cited alongside, same era.
L. Zhou, C. Xu, and J. Corso, “Towards automatic learning of procedures from web instructional videos,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
E. Dunbar, R. Algayres, J. Karadayi, M. Bernard, J. Benjumea, X.-N. Cao, L. Miskic, C. Dugrain, L. Ondel, A. W. Black, L. Besacier, S. Sakti, and E. Dupoux, “The zero resource speech challenge 2019: Tts without t,” 2019
2019
Later among the works it cites.
D. Merkx, S. L. Frank, and M. Ernestus, “Language Learning Using Speech to Image Retrieval,” in Proc. Interspeech 2019 , 2019, pp. 1841–1845
2019
Later among the works it cites.
G. Chrupała, “Symbolic Inductive Bias for Visually Grounded Learning of Spoken Language,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Florence, Italy: Association for Computational Linguistics, Jul. 2019, pp. 6452–6462
2019
Later among the works it cites.
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips,” in ICCV , 2019
2019
Later among the works it cites.
2020
Later among the works it cites.
W.-N. Hsu, D. Harwath, C. Song, and J. Glass, “Text-Free Image-to-Speech Synthesis Using Learned Segmental Units,” in 34th Conference on Neural Information Processing Systems (NeurIPS) Workshop on Self-Supervised Learning for Speech and Audio Processing , Dec. 2020
2020
Later among the works it cites.
B. Higy, D. Elliott, and G. Chrupała, “Textual Supervision for Visually Grounded Spoken Language Understanding,” in Findings of the Association for Computational Linguistics: EMNLP 2020 . Online: Association for Computational Linguistics, Nov. 2020, pp. 2698–2709
2020
Later among the works it cites.