Fetching the paper…
Reading the bibliography…
Given a collection of images and spoken audio captions, we present a method for discovering word-like acoustic units in the continuous speech signal and grounding them to semantically relevant image regions.
The TIMIT acoustic-phonetic continuous speech corpus
John Garofolo, Lori Lamel, William Fisher, Jonathan Fiscus, David Pallet, Nancy Dahlgren, and Victor Zue. 1993 · 1993
Earlier work this paper cites.
WordNet: An Electronic Lexical Database
Christiane Fellbaum. 1998 · 1998
Earlier work this paper cites.
Grounded spoken language acquisition: Experiments in word learning
Deb Roy. 2003 · 2003
Earlier work this paper cites.
Unsupervised word segmentation for sesotho using adaptor grammars
Mark Johnson. 2008 · 2008
Earlier work this paper cites.
Unsupervised pattern discovery in speech
Alex Park and James Glass. 2008 · 2008
Earlier work this paper cites.
Visualizing high-dimensional data using t-sne
Laurens van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
A Bayesian framework for word segmentation: exploring the effects of context
Sharon Goldwater, Thomas Griffiths, and Mark Johnson. 2009 · 2009
Earlier work this paper cites.
Unsupervised spoken keyword spotting via segmental DTW on Gaussian posteriorgrams
Yaodong Zhang and James Glass. 2009 · 2009
Earlier work this paper cites.
NLP on spoken documents without ASR
Mark Dredze, Aren Jansen, Glen Coppersmith, and Kenneth Church. 2010 · 2010
Earlier work this paper cites.
Toward spoken term discovery at scale with zero resources
Aren Jansen, Kenneth Church, and Hynek Hermansky. 2010 · 2010
Earlier work this paper cites.
Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora
Richard Socher and Fei-Fei Li. 2010 · 2010
Earlier work this paper cites.
Efficient spoken term discovery using randomized algorithms
Aren Jansen and Benjamin Van Durme. 2011 · 2011
Earlier work this paper cites.
The Kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely. 2011 · 2011
Cited alongside, same era.
Zero resource spoken audio corpus analysis
David Harwath, Timothy J. Hazen, and James Glass. 2012 · 2012
Cited alongside, same era.
A nonparametric Bayesian approach to acoustic model discovery
Chia-Ying Lee and James Glass. 2012 · 2012
Cited alongside, same era.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomas Mikolov. 2013 · 2013
Cited alongside, same era.
Self-taught object localization with deep networks
Alessandro Bergamo, Loris Bazzani, Dragomir Anguelov, and Lorenzo Torresani. 2014 · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li. 2015 · 2015
Later among the works it cites.
Unsupervised lexicon discovery from acoustic input
Chia-Ying Lee, Timothy J. O’Donnell, and James Glass. 2015 · 2015
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015 · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dimitru Erhan. 2015 · 2015
Later among the works it cites.
Object detectors emerge in deep scene CNNs
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. 2015 · 2015
Later among the works it cites.
Weakly supervised object localization with multi-fold multiple instance learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrej Karpathy, Armand Joulin, and Fei-Fei Li. 2014 · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Cited alongside, same era.
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Andrej Karpathy, Quoc V. Le, Christopher D. Manning, and Andrew Y. Ng. 2014 · 2014
Cited alongside, same era.
Learning deep features for scene recognition using places database
Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva. 2014 · 2014
Cited alongside, same era.
Unsupervised object discovery and localization in the wild: Part-based matching with bottom-up region proposals
Minsu Cho, Suha Kwak, Cordelia Schmid, and Jean Ponce. 2015 · 2015
Cited alongside, same era.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Srivastava Rupesh, Li Deng, Piotr Dollar, Jianfeng Gao, Xiaodong He, Margaret Mitchell, Platt John C., C. Lawrence Zitnick, and Geoffrey Zweig. 2015 · 2015
Cited alongside, same era.
Deep multimodal semantic embeddings for speech and images
David Harwath and James Glass. 2015 · 2015
Cited alongside, same era.
Ramazan Cinbis, Jakob Verbeek, and Cordelia Schmid. 2016 · 2016
Later among the works it cites.
Lieke Gelderloos and Grzegorz Chrupała. 2016 · 2016
Later among the works it cites.
Unsupervised learning of spoken language with visual context
David Harwath, Antonio Torralba, and James R. Glass. 2016 · 2016
Later among the works it cites.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei. 2016 · 2016
Later among the works it cites.
Ethnologue: Languages of the World, Nineteenth edition
M. Paul Lewis, Gary F. Simon, and Charles D. Fennig. 2016 · 2016
Later among the works it cites.
Variational inference for acoustic unit discovery
Lucas Ondel, Lukas Burget, and Jan Cernocky. 2016 · 2016
Later among the works it cites.