Fetching the paper…
Reading the bibliography…
In this paper we present the first model for directly synthesizing fluent, natural-sounding spoken audio captions for images that does not require natural language text as an intermediate representation or source of supervision.
VQVAE unsupervised unit discovery and multi-scale code2spec inverter for zerospeech challenge 2019
Andros Tjandra, Berrak Sisman, Mingyang Zhang, Sakriani Sakti, Haizhou Li, and Satoshi Nakamura. 2019b · 1905
Earlier work this paper cites.
Effectiveness of self-supervised pre-training for speech recognition
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed. 2019 · 1911
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim. 1984 · 1984
Earlier work this paper cites.
Voice conversion through vector quantization
Masanobu Abe, Satoshi Nakamura, Kiyohiro Shikano, and Hisao Kuwabara. 1990 · 1990
Earlier work this paper cites.
Continuous probabilistic transform for voice conversion
Yannis Stylianou, Olivier Cappé, and Eric Moulines. 1998 · 1998
Earlier work this paper cites.
Handbook of the International Phonetic Association: A guide to the use of the International Phonetic Alphabet
International Phonetic Association. 1999 · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Data augmenting contrastive learning of speech representations in the time domain
Eugene Kharitonov, Morgane Rivière, Gabriel Synnaeve, Lior Wolf, Pierre-Emmanuel Mazaré, Matthijs Douze, and Emmanuel Dupoux. 2020 · 2007
Earlier work this paper cites.
Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory
Tomoki Toda, Alan W Black, and Keiichi Tokuda. 2007 · 2007
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Distance metric learning for large margin nearest neighbor classification
Kilian Q. Weinberger and Lawrence K. Saul. 2009 · 2009
Earlier work this paper cites.
Collecting image annotations using Amazon’s Mechanical Turk
Cyrus Rashtchian, Peter Young, Micah Hodosh, and Julia Hockenmaier. 2010 · 2010
Earlier work this paper cites.
Show and speak: Directly synthesize spoken description of images
Xinsheng Wang, Siyuan Feng, Jihua Zhu, Mark Hasegawa-Johnson, and Odette Scharenborg. 2020b · 2010
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie. 2014 · 2014
Earlier work this paper cites.
Multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Rich Zemel. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Learning words from images and speech
Gabriel Synnaeve, Maarten Versteegh, and Emmanuel Dupoux. 2014 · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
Bolei Zhou, Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and Aude Oliva. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Deep multimodal semantic embeddings for speech and images
David Harwath and James Glass. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille. 2015 · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Breaking the unwritten language barrier: The BULB project
Gilles Adda, Sebastian Stüker, Martine Adda-Decker, Odette Ambouroue, Laurent Besacier, David Blachon, Hélene Bonneau-Maynard, Pierre Godard, Fatima Hamlaoui, Dmitry Idiatov, et al. 2016 · 2016
Earlier work this paper cites.
SPICE: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Earlier work this paper cites.
Unsupervised learning of spoken language with visual context
David Harwath, Antonio Torralba, and James R. Glass. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Voice conversion from non-parallel corpora using variational auto-encoder
Chin-Cheng Hsu, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao, and Hsin-Min Wang. 2016 · 2016
Cited alongside, same era.
Ethnologue: Languages of the World, Nineteenth edition
M. Paul Lewis, Gary F. Simon, and Charles D. Fennig. 2016 · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
Encoding of phonology in a recurrent neural model of grounded speech
Afra Alishahi, Marie Barking, and Grzegorz Chrupała. 2017 · 2017
Cited alongside, same era.
The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods
Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda, Daisuke Saito, Fernando Villavicencio, Tomi Kinnunen, and Zhenhua Ling. 2018 · 2018
Later among the works it cites.
Neural baby talk
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2018 · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Later among the works it cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al. 2018 · 2018
Later among the works it cites.
Diverse beam search for improved description of complex scenes
Ashwin K Vijayakumar, Michael Cogswell, Ramprasaath R Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Representations of language in a model of visually grounded speech signal
Grzegorz Chrupała, Lieke Gelderloos, and Afra Alishahi. 2017 · 2017
Cited alongside, same era.
Contrastive learning for image captioning
Bo Dai and Dahua Lin. 2017 · 2017
Cited alongside, same era.
Analysis of audio-visual features for unsupervised speech recognition
Jennifer Drexler and James Glass. 2017 · 2017
Cited alongside, same era.
Learning word-like units from joint audio-visual analysis
David Harwath and James Glass. 2017 · 2017
Cited alongside, same era.
Image2speech: Automatically generating audio descriptions of images
Mark Hasegawa-Johnson, Alan Black, Lucas Ondel, Odette Scharenborg, and Francesco Ciannella. 2017 · 2017
Cited alongside, same era.
Speech-coco: 600k visually grounded spoken captions aligned to mscoco data set
William Havard, Laurent Besacier, and Olivier Rosec. 2017 · 2017
Cited alongside, same era.
The LJ speech dataset
Keith Ito. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Fei Ren, Ye Jia, and Rif A Saurous. 2018 · 2018
Later among the works it cites.
VQVAE with speaker adversarial training
Suhee Cho, Yeonjung Hong, Yookyunk Shin, and Youngsun Cho. 2019 · 2019
Later among the works it cites.
Unsupervised speech representation learning using wavenet autoencoders
Jan Chorowski, Ron J. Weiss, Samy Bengio, and Aäron van den Oord. 2019 · 2019
Later among the works it cites.
An unsupervised autoregressive model for speech representation learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James R. Glass. 2019 · 2019
Later among the works it cites.
The zero resource speech challenge 2019: TTS without T
Ewan Dunbar, Robin Algayres, Julien Karadayi, Mathieu Bernard, Juan Benjumea, Xuan-Nga Cao, Lucie Miskic, Charlotte Dugrain, Lucas Ondel, Alan W. Black, Laurent Besacier, Sakriani Sakti, and Emmanuel Dupoux. 2019 · 2019
Later among the works it cites.
Multimodal one-shot learning of speech and images
Ryan Eloff, Herman Engelbrecht, and Herman Kamper. 2019 · 2019
Later among the works it cites.
Towards visually grounded sub-word speech unit discovery
David Harwath and James Glass. 2019 · 2019
Later among the works it cites.
Jointly discovering visual objects and spoken words from raw sensory input
David Harwath, Adrià Recasens, Dídac Surís, Galen Chuang, Antonio Torralba, and James Glass. 2019 · 2019
Later among the works it cites.
Hierarchical generative modeling for controllable speech synthesis
Wei-Ning Hsu, Yu Zhang, Ron Weiss, Heiga Zen, Yonghui Wu, Yuan Cao, and Yuxuan Wang. 2019 · 2019
Later among the works it cites.
Large-scale representation learning from visually grounded untranscribed speech
Gabriel Ilharco, Yuan Zhang, and Jason Baldridge. 2019 · 2019
Later among the works it cites.
A factorial deep markov model for unsupervised disentangled representation learning from speech
Sameer Khurana, Shafiq Rayhan Joty, Ahmed Ali, and James Glass. 2019 · 2019
Later among the works it cites.
Importance of search and evaluation strategies in neural dialogue modeling
Ilia Kulikov, Alexander Miller, Kyunghyun Cho, and Jason Weston. 2019 · 2019
Later among the works it cites.
Unpaired image-to-speech synthesis with multimodal information bottleneck
Shuang Ma, Daniel McDuff, and Yale Song. 2019 · 2019
Later among the works it cites.
Language learning using speech to image retrieval
Danny Merkx, Stefan L. Frank, and Mirjam Ernestus. 2019 · 2019
Later among the works it cites.
Waveglow: A flow-based generative network for speech synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. 2019 · 2019
Later among the works it cites.
Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion
Joan Serrà, Santiago Pascual, and Carlos Segura Perales. 2019 · 2019
Later among the works it cites.
Learning words by drawing images
Dídac Surís, Adrià Recasens, David Bau, David Harwath, James Glass, and Antonio Torralba. 2019 · 2019
Later among the works it cites.
Speech-to-speech translation between untranscribed unknown languages
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2019a · 2019
Later among the works it cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli. 2020 · 2020
Closest in time.
Learning hierarchical discrete linguistic units from visually-grounded speech
David Harwath, Wei-Ning Hsu, and James Glass. 2020 · 2020
Closest in time.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Closest in time.
On mutual information maximization for representation learning
Michael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly, and Mario Lucic. 2020 · 2020
Closest in time.
Neural text generation with unlikelihood training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020 · 2020
Closest in time.