Fetching the paper…
Reading the bibliography…
In this paper, we explore the learning of neural network embeddings for natural images and speech waveforms describing the content of those images.
“Signature verification using a ”siamese” time delay neural network,”
J. Bromley, I. Guyon, Y. LeCun, E. Säckinger, and R. Shah, · 1994
Earlier work this paper cites.
“Matching words and pictures,”
K. Barnard, P. Duygulu, D. Forsyth, N. DeFreitas, D. M. Blei, and M. I. Jordan, · 2003
Earlier work this paper cites.
“Grounded spoken language acquisition: Experiments in word learning,”
D. Roy, · 2003
Earlier work this paper cites.
“Unsupervised pattern discovery in speech,”
A. Park and J. Glass, · 2008
Earlier work this paper cites.
“Imagenet: A large scale hierarchical image database,”
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and F. Li, · 2009
Earlier work this paper cites.
“Toward spoken term discovery at scale with zero resources,”
A. Jansen, K. Church, and H. Hermansky, · 2010
Earlier work this paper cites.
“Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora,”
R. Socher and F. Li, · 2010
Earlier work this paper cites.
“Efficient spoken term discovery using randomized algorithms,”
A. Jansen and B. Van Durme, · 2011
Earlier work this paper cites.
“A nonparametric Bayesian approach to acoustic model discovery,”
C. Lee and J. Glass, · 2012
Earlier work this paper cites.
“Resource configurable spoken query detection using deep boltzmann machines,”
Y. Zhang, R. Salakhutdinov, H.-A. Chang, and J. Glass, · 2012
Earlier work this paper cites.
“A joint model of language and perception for grounded attribute learning,”
C. Matuszek, N. Fitzgerald, L. Zettlemoyer, L. Bo, and D. Fox, · 2012
Earlier work this paper cites.
“Devise: A deep visual-semantic embedding model,”
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov, · 2013
Earlier work this paper cites.
“Statistical phrase-based translation,”
P. Koehn, F. Och, and D. Marcu, · 2013
Earlier work this paper cites.
“Visual semantic search: Retrieving videos via complex textual queries,”
D. Lin, S. Fidler, C. Kong, and R. Urtasun, · 2014
Earlier work this paper cites.
“What are you talking about? text-to-image coreference,”
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler, · 2014
Earlier work this paper cites.
“Very deep convolutional networks for large-scale image recognition,”
K. Simonyan and A. Zisserman, · 2014
Cited alongside, same era.
“Learning deep features for scene recognition using places database,”
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva, · 2014
Cited alongside, same era.
“Deep multimodal semantic embeddings for speech and images,”
D. Harwath and J. Glass, · 2015
Cited alongside, same era.
“A comparison of neural network methods for unsupervised representation learning on the zero resource speech challenge,”
D. Renshaw, H. Kamper, A. Jansen, and S. Goldwater, · 2015
Cited alongside, same era.
“Unsupervised neural network based feature extraction using weak top-down constraints,”
H. Kamper, M. Elsner, A. Jansen, and S. Goldwater, · 2015
Cited alongside, same era.
“Variational inference for acoustic unit discovery,”
L. Ondel, L. Burget, and J. Cernocky, · 2016
Later among the works it cites.
“Unsupervised word segmentation and lexicon discovery using acoustic word embeddings,”
H. Kamper, A. Jansen, and S. Goldwater, · 2016
Later among the works it cites.
“Generative adversarial text to image synthesis,”
S. E. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee, · 2016
Later among the works it cites.
“A shared task on multimodal machine translation and crosslingual image description,”
L. Specia, S. Frank, K. Sima’an, and D. Elliott, · 2016
Later among the works it cites.
“An attentional model for speech translation without transcription,”
L. Duong, A. Anastasopoulos, D. Chiang, S. Bird, and T. Cohn, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“A hybrid dynamic time warping-deep neural net- work architecture for unsupervised acoustic modeling,”
R. Thiolliere, E. Dunbar, G. Synnaeve, M. Versteegh, and E. Dupoux, · 2015
Cited alongside, same era.
“Deep visual-semantic alignments for generating image descriptions,”
A. Karpathy and F. Li, · 2015
Cited alongside, same era.
“Show and tell: A neural image caption generator,”
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, · 2015
Cited alongside, same era.
“From captions to visual concepts and back,”
H. Fang, S. Gupta, F. Iandola, S. Rupesh, L. Deng, P. Dollar, J. Gao, X. He, M. Mitchell, P. J. C., C. L. Zitnick, and G. Zweig, · 2015
Cited alongside, same era.
“VQA: Visual question answering,”
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, Z. Lawrence, and D. Parikh, · 2015
Cited alongside, same era.
“Neural machine translation by jointly learning to align and translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2015
Cited alongside, same era.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
S. Ioffe and C. Szegedy, · 2015
Cited alongside, same era.
D. Harwath and J. Glass, · 2017
Later among the works it cites.
“Analysis of audio-visual features for unsupervised speech recognition,”
J. Drexler and J. Glass, · 2017
Later among the works it cites.
“Representations of language in a model of visually grounded speech signal,”
G. Chrupala, L. Gelderloos, and A. Alishahi, · 2017
Later among the works it cites.
“Encoding of phonology in a recurrent neural model of grounded speech,”
A. Alishahi, M. Barking, and G. Chrupala, · 2017
Later among the works it cites.
“A segmental framework for fully-unsupervised large-vocabulary speech recognition,”
H. Kamper, A. Jansen, and S. Goldwater, · 2017
Later among the works it cites.
“Guesswhat?! visual object discovery through multi-modal dialogue,”
H. de Vries, F. Strub, S. Chandar, O. Pietquin, H. Larochelle, and A. C. Courville, · 2017
Later among the works it cites.
“Sequence-to-sequence models can directly translate foreign speech,”
R. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z. Chen, · 2017
Later among the works it cites.
“Towards speech-to-text translation without speech recognition,”
A. L. S. G. Sameer Bansal, Herman Kamper, · 2017
Later among the works it cites.
“Image pivoting for learning multilingual multimodal representations,”
S. Gella, R. Sennrich, F. Keller, and M. Lapata, · 2017
Later among the works it cites.