Fetching the paper…
Reading the bibliography…
In this paper, we explore neural network models that learn to associate segments of spoken audio captions with the semantically relevant portions of natural images that they refer to.
Principles of object perception
Spelke, E.S.: · 1990
Earlier work this paper cites.
Signature verification using a ”siamese” time delay neural network
Bromley, J., Guyon, I., LeCun, Y., Säckinger, E., Shah, R.: · 1994
Earlier work this paper cites.
Birch: an efficient data clustering method for very large databases
Zhang, T., Ramakrishnan, R., Livny, M.: · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., Haffner, P.: · 1998
Earlier work this paper cites.
Learning words from sights and sounds: a computational model
Roy, D., Pentland, A.: · 2002
Earlier work this paper cites.
Grounded spoken language acquisition: Experiments in word learning
Roy, D.: · 2003
Earlier work this paper cites.
Using multiple segmentations to discover objects and their extent in image collections
Russell, B., Efros, A., Sivic, J., Freeman, W., Zisserman, A.: · 2006
Earlier work this paper cites.
Unsupervised pattern discovery in speech
Park, A., Glass, J.: · 2008
Earlier work this paper cites.
Towards automatic discovery of object categories
Weber, M., Welling, M., Perona, P.: · 2010
Earlier work this paper cites.
Toward spoken term discovery at scale with zero resources
Jansen, A., Church, K., Hermansky, H.: · 2010
Earlier work this paper cites.
Efficient spoken term discovery using randomized algorithms
Jansen, A., Van Durme, B.: · 2011
Earlier work this paper cites.
A nonparametric Bayesian approach to acoustic model discovery
Lee, C., Glass, J.: · 2012
Earlier work this paper cites.
Resource configurable spoken query detection using deep boltzmann machines
Zhang, Y., Salakhutdinov, R., Chang, H.A., Glass, J.: · 2012
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., Malik, J.: · 2013
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Malinowski, M., Fritz, M.: · 2014
Earlier work this paper cites.
Self-taught object localization with deep networks
Bergamo, A., Bazzani, L., Anguelov, D., Torresani, L.: · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., Oliva, A.: · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Karpathy, A., Joulin, A., Fei-Fei, L.: · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K., Zisserman, A.: · 2014
Earlier work this paper cites.
Object detectors emerge in deep scene CNNs
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., Torralba, A.: · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A., Li, F.F.: · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., Erhan, D.: · 2015
Cited alongside, same era.
From captions to visual concepts and back
Fang, H., Gupta, S., Iandola, F., Rupesh, S., Deng, L., Dollar, P., Gao, J., He, X., Mitchell, M., C., P.J., Zitnick, C.L., Zweig, G.: · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R., Bengio, Y.: · 2015
Cited alongside, same era.
VQA: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence, Z., Parikh, D.: · 2015
Cited alongside, same era.
Ask your neurons: A neural-based approach to answering questions about images
Malinowski, M., Rohrbach, M., Fritz, M.: · 2015
Soundnet: Learning sound representations from unlabeled video
Aytar, Y., Vondrick, C., Torralba, A.: · 2016
Later among the works it cites.
Unsupervised learning of spoken language with visual context
Harwath, D., Torralba, A., Glass, J.R.: · 2016
Later among the works it cites.
You only look once: Unified, real-time object detection
Redmon, J., Divvala, S., Girshick, R., Farhadi, A.: · 2016
Later among the works it cites.
Weakly supervised object localization with multi-fold multiple instance learning
Cinbis, R., Verbeek, J., Schmid, C.: · 2016
Later among the works it cites.
Ethnologue: Languages of the World, Nineteenth edition
Lewis, M.P., Simon, G.F., Fennig, C.D.: · 2016
Later among the works it cites.
Variational inference for acoustic unit discovery
Ondel, L., Burget, L., Cernocky, J.: · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Are you talking to a machine? dataset and methods for multilingual image question answering
Gao, H., Mao, J., Zhou, J., Huang, Z., Yuille, A.: · 2015
Cited alongside, same era.
Exploring models and data for image question answering
Ren, M., Kiros, R., Zemel, R.: · 2015
Cited alongside, same era.
Unsupervised object discovery and localization in the wild: Part-based matching with bottom-up region proposals
Cho, M., Kwak, S., Schmid, C., Ponce, J.: · 2015
Cited alongside, same era.
Object detectors emerge in deep scene CNNs
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., Torralba, A.: · 2015
Cited alongside, same era.
Unsupervised visual representation learning by context prediction
Doersch, C., Gupta, A., Efros, A.A.: · 2015
Cited alongside, same era.
A comparison of neural network methods for unsupervised representation learning on the zero resource speech challenge
Renshaw, D., Kamper, H., Jansen, A., Goldwater, S.: · 2015
Cited alongside, same era.
Unsupervised word segmentation and lexicon discovery using acoustic word embeddings
Kamper, H., Jansen, A., Goldwater, S.: · 2016
Later among the works it cites.
Gelderloos, L., Chrupała, G.: · 2016
Later among the works it cites.
Learning deep features for discriminative localization
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., Torralba, A.: · 2016
Later among the works it cites.
Guesswhat?! visual object discovery through multi-modal dialogue
de Vries, H., Strub, F., Chandar, S., Pietquin, O., Larochelle, H., Courville, A.C.: · 2017
Later among the works it cites.
Look, listen, and learn
Arandjelovic, R., Zisserman, A.: · 2017
Later among the works it cites.
CNN features are also great at unsupervised classification
Guérin, J., Gibaru, O., Thiery, S., Nyiri, E.: · 2017
Later among the works it cites.
Learning wor d-like units from joint audio-visual analysis
Harwath, D., Glass, J.: · 2017
Later among the works it cites.
Analysis of audio-visual features for unsupervised speech recognition
Drexler, J., Glass, J.: · 2017
Later among the works it cites.
Representations of language in a model of visually grounded speech signal
Chrupala, G., Gelderloos, L., Alishahi, A.: · 2017
Later among the works it cites.
Encoding of phonology in a recurrent neural model of grounded speech
Alishahi, A., Barking, M., Chrupala, G.: · 2017
Later among the works it cites.
A segmental framework for fully-unsupervised large-vocabulary speech recognition
Kamper, H., Jansen, A., Goldwater, S.: · 2017
Later among the works it cites.
Scene parsing through ADE20K dataset
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., Torralba, A.: · 2017
Later among the works it cites.
Cognitive science in the era of artificial intelligence: A roadmap for reverse-engineering the infant language-learner
Dupoux, E.: · 2018
Closest in time.