Fetching the paper…
Reading the bibliography…
Systems that can associate images with their spoken audio captions are an important step towards visually grounded language learning.
Towards visually grounded sub-word speech unit discovery
David Harwath and James Glass. 2019 · 1902
Earlier work this paper cites.
Yinfei Yang, Gustavo Hernandez Abrego, Steve Yuan, Mandy Guo, Qinlan Shen, Daniel Cer, Yun-hsuan Sung, Brian Strope, and Ray Kurzweil. 2019 · 1902
Earlier work this paper cites.
Jasper: An end-to-end convolutional neural acoustic model
Jason Li, Vitaly Lavrukhin, Boris Ginsburg, Ryan Leary, Oleksii Kuchaiev, Jonathan M Cohen, Huyen Nguyen, and Ravi Teja Gadde. 2019 · 1904
Earlier work this paper cites.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le. 2019 · 1904
Earlier work this paper cites.
The symbol grounding problem
Stevan Harnad. 1990 · 1990
Earlier work this paper cites.
Signature verification using a ”siamese” time delay neural network
Jane Bromley, Isabelle Guyon, Yann LeCun, Eduard Säckinger, and Roopak Shah. 1994 · 1994
Earlier work this paper cites.
Infants′ detection of the sound patterns of words in fluent speech
PW Jusczyk and RN Aslin. 1995 · 1995
Earlier work this paper cites.
Word segmentation: The role of distributional cues
Jenny R. Saffran, Elissa L. Newport, and Richard N. Aslin. 1996 · 1996
Earlier work this paper cites.
Scaling to very very large corpora for natural language disambiguation
Michele Banko and Eric Brill. 2001 · 2001
Earlier work this paper cites.
Unsupervised pattern discovery in speech
Alex S Park and James R Glass. 2007 · 2007
Earlier work this paper cites.
Unsupervised learning of acoustic sub-word units
Balakrishnan Varadarajan, Sanjeev Khudanpur, and Emmanuel Dupoux. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
The unreasonable effectiveness of data
Alon Halevy, Peter Norvig, and Fernando Pereira. 2009 · 2009
Earlier work this paper cites.
Amazon’s mechanical turk: A new source of inexpensive, yet high-quality, data?
Michael Buhrmester, Tracy Kwang, and Samuel D Gosling. 2011 · 2011
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013 · 2013
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li F Fei-Fei. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Cited alongside, same era.
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Andrej Karpathy, Quoc V Le, Christopher D Manning, and Andrew Y Ng. 2014 · 2014
Cited alongside, same era.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015 · 2015
Cited alongside, same era.
Deep multimodal semantic embeddings for speech and images
David Harwath and James Glass. 2015 · 2015
Cited alongside, same era.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Cited alongside, same era.
Weakly supervised spoken term discovery using cross-lingual side information
Sameer Bansal, Herman Kamper, Sharon Goldwater, and Adam Lopez. 2017 · 2017
Later among the works it cites.
Representations of language in a model of visually grounded speech signal
Grzegorz Chrupała, Lieke Gelderloos, and Afra Alishahi. 2017 · 2017
Later among the works it cites.
Very deep convolutional networks for text classification
Alexis Conneau, Holger Schwenk, Loïc Barrault, and Yann Lecun. 2017 · 2017
Later among the works it cites.
Efficient natural language response suggestion for smart reply
Matthew Henderson, Rami Al-Rfou, Brian Strope, Yun-hsuan Sung, Laszlo Lukacs, Ruiqi Guo, Sanjiv Kumar, Balint Miklos, and Ray Kurzweil. 2017 · 2017
Later among the works it cites.
Incidental supervision: Moving beyond supervised learning
Dan Roth. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
Raffaella Bernardi, Ruket Cakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, and Barbara Plank. 2016 · 2016
Cited alongside, same era.
Fully-convolutional siamese networks for object tracking
Luca Bertinetto, Jack Valmadre, Joao F Henriques, Andrea Vedaldi, and Philip HS Torr. 2016 · 2016
Cited alongside, same era.
Unsupervised learning of spoken language with visual context
David Harwath, Antonio Torralba, and James Glass. 2016 · 2016
Cited alongside, same era.
Unsupervised word segmentation and lexicon discovery using acoustic word embeddings
Herman Kamper, Aren Jansen, and Sharon Goldwater. 2016 · 2016
Cited alongside, same era.
Siamese recurrent architectures for learning sentence similarity
Jonas Mueller and Aditya Thyagarajan. 2016 · 2016
Cited alongside, same era.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. 2017 · 2017
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. 2017 · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Later among the works it cites.
Phoneme based embedded segmental k-means for unsupervised term discovery
Saurabhch Bhati, Herman Kamper, and K Sri Rama Murty. 2018 · 2018
Later among the works it cites.
Benchmark analysis of representative deep neural network architectures
Simone Bianco, Remi Cadene, Luigi Celona, and Paolo Napoletano. 2018 · 2018
Later among the works it cites.
Temporally grounding natural sentence in video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua. 2018 · 2018
Later among the works it cites.
Jointly discovering visual objects and spoken words from raw sensory input
David Harwath, Adria Recasens, Dídac Surís, Galen Chuang, Antonio Torralba, and James Glass. 2018 · 2018
Later among the works it cites.
Illustrative language understanding: Large-scale visual grounding with image search
Jamie Kiros, William Chan, and Geoffrey Hinton. 2018 · 2018
Later among the works it cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Later among the works it cites.
ParaNMT-50M: Pushing the limits of paraphrastic sentence embeddings with millions of machine translations
John Wieting and Kevin Gimpel. 2018 · 2018
Later among the works it cites.
Fast and accurate reading comprehension by combining self-attention and convolution
Adams Wei Yu, David Dohan, Quoc Le, Thang Luong, Rui Zhao, and Kai Chen. 2018 · 2018
Later among the works it cites.
Symbolic inductive bias for visually grounded learning of spoken language
Grzegorz Chrupała. 2019 · 2019
Closest in time.