Fetching the paper…
Reading the bibliography…
This paper introduces speech-based visual question answering (VQA), the task of generating an answer given an image and a spoken question.
Show&Tell: a semi-automated image annotation system
R.K. Srihari and Zhongfei Zhang. 2000 · 2000
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd international conference on Machine learning
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
An Application of Recurrent Neural Networks to Discriminative Keyword Spotting
Santiago Fernández, Alex Graves, and Jürgen Schmidhuber. 2007 · 2007
Earlier work this paper cites.
Speech-based annotation and retrieval of digital photographs. In INTERSPEECH
Timothy J. Hazen, Brennan Sherry, and Mark Adler. 2007 · 2007
Earlier work this paper cites.
VizWiz: nearly real-time answers to visual questions. In Proceedings of the 23nd annual ACM symposium on User interface software and technology
Jeffrey P Bigham, Chandrika Jayant, Hanjie Ji, Greg Little, Andrew Miller, Robert C Miller, Robin Miller, Aubrey Tatarowicz, Brandyn White, Samual White, and others. 2010 · 2010
Earlier work this paper cites.
A Semantics-Based Approach for Speech Annotation of Images
D.V. Kalashnikov, S. Mehrotra, Jie Xu, and N. Venkatasubramanian. 2011 · 2011
Earlier work this paper cites.
The Kaldi Speech Recognition Toolkit. In IEEE 2011 Workshop on Automatic Speech Recognition and Understanding
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely. 2011 · 2011
Earlier work this paper cites.
PixelTone: a multimodal interface for image editing. In CHI
Gierad P Laput, Mira Dontcheva, Gregg Wilensky, Walter Chang, Aseem Agarwala, Jason Linder, and Eytan Adar. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
ImageSpirit: Verbal Guided Image Parsing
Ming-Ming Cheng, Shuai Zheng, Wen-Yan Lin, Vibhav Vineet, Paul Sturgess, Nigel Crook, Niloy J. Mitra, and Philip Torr. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input. In Advances in Neural Information Processing Systems
Mateusz Malinowski and Mario Fritz. 2014 · 2014
Cited alongside, same era.
Recurrent models of visual attention. In Advances in neural information processing systems
Volodymyr Mnih, Nicolas Heess, Alex Graves, and others. 2014 · 2014
Cited alongside, same era.
A Dataset and Taxonomy for Urban Sound Research. In Proceedings of the ACM International Conference on Multimedia, MM ’14, Orlando, FL, USA, November 03 - 07, 2014
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello. 2014 · 2014
Cited alongside, same era.
Smile - Smart Photo Annotation
2015 · 2015
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, and others. 2015 · 2015
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE International Conference on Computer Vision
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Later among the works it cites.
SoundNet: Learning Sound Representations from Unlabeled Video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba. 2016 · 2016
Later among the works it cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016 · 2016
Later among the works it cites.
Unsupervised Learning of Spoken Language with Visual Context. In Advances in Neural Information Processing Systems
David Harwath, Antonio Torralba, and James Glass. 2016 · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
VQA: Visual Question Answering. In International Conference on Computer Vision (ICCV)
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, and others. 2015 · 2015
Cited alongside, same era.
Very Deep Convolutional Networks for Large-Scale Image Recognition. In ICLR
K. Simonyan and A. Zisserman. 2015 · 2015
Cited alongside, same era.
Show and Tell: A Neural Image Caption Generator. In CVPR
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016 · 2016
Later among the works it cites.
Dual Attention Networks for Multimodal Reasoning and Matching
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. 2016 · 2016
Later among the works it cites.
Speech Is 3x Faster than Typing for English and Mandarin Text Entry on Mobile Devices
Sherry Ruan, Jacob O Wobbrock, Kenny Liou, Andrew Ng, and James Landay. 2016 · 2016
Later among the works it cites.
Dynamic memory networks for visual and textual question answering
Caiming Xiong, Stephen Merity, and Richard Socher. 2016 · 2016
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering. In European Conference on Computer Vision
Huijuan Xu and Kate Saenko. 2016 · 2016
Later among the works it cites.
Stacked attention networks for image question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. 2016 · 2016
Later among the works it cites.