Fetching the paper…
Reading the bibliography…
In this paper, we explore a method for training speech-to-speech translation tasks without any transcription or linguistic supervision.
“Signal estimation from modified short-time Fourier transform,”
Daniel Griffin and Jae Lim, · 1984
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Bleu: a method for automatic evaluation of machine translation,”
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, · 2002
Earlier work this paper cites.
“Creating corpora for speech-to-speech translation,”
Genichiro Kikui, Eiichiro Sumita, Toshiyuki Takezawa, and Seiichi Yamamoto, · 2003
Earlier work this paper cites.
“METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,”
Satanjeev Banerjee and Alon Lavie, · 2005
Earlier work this paper cites.
“Comparative study on corpora for speech translation,”
Gen-ichiro Kikui, Seiichi Yamamoto, Toshiyuki Takezawa, and Eiichiro Sumita, · 2006
Earlier work this paper cites.
“Extracting and composing robust features with denoising autoencoders,”
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol, · 2008
Earlier work this paper cites.
“Better hypothesis testing for statistical machine translation: Controlling for optimizer instability,”
Jonathan H Clark, Chris Dyer, Alon Lavie, and Noah A Smith, · 2011
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
Ilya Sutskever, Oriol Vinyals, and Quoc V Le, · 2014
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Cited alongside, same era.
“The zero resource speech challenge 2015,”
Maarten Versteegh, Roland Thiolliere, Thomas Schatz, Xuan Nga Cao, Xavier Anguera, Aren Jansen, and Emmanuel Dupoux, · 2015
Cited alongside, same era.
“Effective approaches to attention-based neural machine translation,”
Thang Luong, Hieu Pham, and Christopher D. Manning, · 2015
Cited alongside, same era.
“Empirical evaluation of rectified activations in convolutional network,”
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li, · 2015
Cited alongside, same era.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
Sergey Ioffe and Christian Szegedy, · 2015
Cited alongside, same era.
“Neural discrete representation learning,”
Aaron van den Oord, Oriol Vinyals, et al., · 2017
Later among the works it cites.
“Listening while speaking: Speech chain by deep learning,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2017
Later among the works it cites.
“Fast decoding in sequence models using discrete latent variables,”
Lukasz Kaiser, Samy Bengio, Aurko Roy, Ashish Vaswani, Niki Parmar, Jakob Uszkoreit, and Noam Shazeer, · 2018
Later among the works it cites.
“Towards a better understanding of vector quantized autoencoders,”
Aurko Roy, Ashish Vaswani, Niki Parmar, and Arvind Neelakantan, · 2018
Later among the works it cites.
“Multi-scale alignment and contextual history for attention mechanism in sequence-to-sequence model,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2018
Later among the works it cites.
“Machine speech chain with one-shot speaker adaptation,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“librosa: Audio and music signal analysis in Python,”
Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto, · 2015
Cited alongside, same era.
“An attentional model for speech translation without transcription,”
Long Duong, Antonios Anastasopoulos, David Chiang, Steven Bird, and Trevor Cohn, · 2016
Cited alongside, same era.
“Listen and translate: A proof of concept for end-to-end speech-to-text translation,”
Alexandre Bérard, Olivier Pietquin, Christophe Servan, and Laurent Besacier, · 2016
Cited alongside, same era.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Cited alongside, same era.
“Sequence-to-sequence models can directly translate foreign speech,”
Ron J. Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen, · 2017
Cited alongside, same era.
“Structured-based curriculum learning for end-to-end english-japanese speech translation,”
Takatomo Kano, Sakriani Sakti, and Satoshi Nakamura, · 2017
Cited alongside, same era.
“The zero resource speech challenge 2017,”
Ewan Dunbar, Xuan Nga Cao, Juan Benjumea, Julien Karadayi, Mathieu Bernard, Laurent Besacier, Xavier Anguera, and Emmanuel Dupoux, · 2017
Cited alongside, same era.
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2018
Later among the works it cites.
“Direct speech-to-speech translation with a sequence-to-sequence model,”
Ye Jia, Ron J Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu, · 2019
Closest in time.
“The zero resource speech challenge 2019: TTS without T,”
Ewan Dunbar, Robin Algayres, Julien Karadayi, Mathieu Bernard, Juan Benjumea, Xuan-Nga Cao, Lucie Miskic, Charlotte Dugrain, Lucas Ondel, Alan W. Black, Laurent Besacier, Sakriani Sakti, and Emmanuel Dupoux, · 2019
Closest in time.
“Unsupervised end-to-end learning of discrete linguistic units for voice conversion,”
Andy T. Liu, Po-chun Hsu, and Hung-yi Lee, · 2019
Closest in time.
“VQVAE with speaker adversarial training,” https://github.com/Suhee05/Zerospeech2019, 2019
Suhee Cho, Yeonjung Hong, Yookyung Shin, and Youngsun Cho, · 2019
Closest in time.
Andros Tjandra, Berrak Sisman, Mingyang Zhang, Sakriani Sakti, Haizhou Li, and Satoshi Nakamura, · 2019
Closest in time.
“End-to-end feedback loss in speech chain framework via straight-through estimator,”
A. Tjandra, S. Sakti, and S. Nakamura, · 2019
Closest in time.