Fetching the paper…
Reading the bibliography…
Existing speech to speech translation systems heavily rely on the text of target language: they usually translate source language either to target text and then synthesize target speech from text, or directly to target speech with target text for auxiliary training.
MASS: Masked Sequence to Sequence Pre-training for Language Generation
Song, K.; Tan, X.; Qin, T.; Lu, J.; and Liu, T.-Y. 2019 · 1905
Earlier work this paper cites.
Exploring phoneme-level speech representations for end-to-end speech translation
Salesky, E.; Sperber, M.; and Black, A. W. 2019 · 1906
Earlier work this paper cites.
vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
Baevski, A.; Schneider, S.; and Auli, M. 2019 · 1910
Earlier work this paper cites.
Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning
Liu, A. H.; Tu, T.; Lee, H.-y.; and Lee, L.-s. 2019 · 1910
Earlier work this paper cites.
Speech-to-speech Translation between Untranscribed Unknown Languages
Tjandra, A.; Sakti, S.; and Nakamura, S. 2019 · 1910
Earlier work this paper cites.
Signal estimation from modified short-time Fourier transform
Griffin, D.; and Lim, J. 1984 · 1984
Earlier work this paper cites.
The evolutionary history of the human speech organs
Wind, J. 1989 · 1989
Earlier work this paper cites.
JANUS-III: Speech-to-speech translation in multiple languages
Lavie, A.; Waibel, A.; Levin, L.; Finke, M.; Gates, D.; Gavalda, M.; Zeppenfeld, T.; and Zhan, P. 1997 · 1997
Earlier work this paper cites.
Handbook of the International Phonetic Association: A guide to the use of the International Phonetic Alphabet
Association, I. P.; Staff, I. P. A.; et al. 1999 · 1999
Earlier work this paper cites.
Speech translation: Coupling of recognition and translation
Ney, H. 1999 · 1999
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
Representing the meanings of object and action words: The featural and unitary semantic space hypothesis
Vigliocco, G.; Vinson, D. P.; Lewis, W.; and Garrett, M. F. 2004 · 2004
Earlier work this paper cites.
On the integration of speech recognition and statistical machine translation
Matusov, E.; Kanthak, S.; and Ney, H. 2005 · 2005
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A.; Fernández, S.; Gomez, F.; and Schmidhuber, J. 2006 · 2006
Earlier work this paper cites.
The ATR multilingual speech-to-speech translation system
Nakamura, S.; Markov, K.; Nakaiwa, H.; Kikui, G.-i.; Kawai, H.; Jitsuhiro, T.; Zhang, J.-S.; Yamamoto, H.; Sumita, E.; and Yamamoto, S. 2006 · 2006
Earlier work this paper cites.
An introduction to phonetics and phonology
Yallop, C.; and Fletcher, J. 2007 · 2007
Earlier work this paper cites.
Phonetic learning as a pathway to language: new data and native language magnet theory expanded (NLM-e)
Kuhl, P. K.; Conboy, B. T.; Coffey-Corina, S.; Padden, D.; Rivera-Gaxiola, M.; and Nelson, T. 2008 · 2008
Earlier work this paper cites.
Improved speech-to-text translation with the Fisher and Callhome Spanish–English speech translation corpus
Post, M.; Kumar, G.; Lopez, A.; Karakos, D.; Callison-Burch, C.; and Khudanpur, S. 2013 · 2013
Cited alongside, same era.
Verbmobil: foundations of speech-to-speech translation
Wahlster, W. 2013 · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D.; Cho, K.; and Bengio, Y. 2014 · 2014
Cited alongside, same era.
Automatic discovery of a phonetic inventory for unwritten languages for statistical speech synthesis
Muthukumar, P. K.; and Black, A. W. 2014 · 2014
Cited alongside, same era.
Parallel inference of Dirichlet process Gaussian mixture models for unsupervised acoustic modeling: A feasibility study
Chen, H.; Leung, C.-C.; Xie, L.; Ma, B.; and Li, H. 2015 · 2015
Cited alongside, same era.
An embedded segmental k-means model for unsupervised segmentation and clustering of speech
Kamper, H.; Livescu, K.; and Goldwater, S. 2017 · 2017
Later among the works it cites.
Neural discrete representation learning
van den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Later among the works it cites.
Sequence-to-Sequence Models Can Directly Translate Foreign Speech
Weiss, R. J.; Chorowski, J.; Jaitly, N.; Wu, Y.; and Chen, Z. 2017 · 2017
Later among the works it cites.
Unsupervised Machine Translation Using Monolingual Corpora Only
Lample, G.; Conneau, A.; Denoyer, L.; and Ranzato, M. 2018 · 2018
Later among the works it cites.
Tensor2Tensor for Neural Machine Translation
Vaswani, A.; Bengio, S.; Brevdo, E.; Chollet, F.; Gomez, A.; Gouws, S.; Jones, L.; Kaiser, Ł.; Kalchbrenner, N.; Parmar, N.; et al. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simons, and Charles D. Fennig (eds.).(2015). Ethnologue: Languages of the World, Dallas, Texas: SIL International
Lewis, M. P.; and Gary, F. 2013 · 2015
Cited alongside, same era.
Effective Approaches to Attention-based Neural Machine Translation
Luong, M.-T.; Pham, H.; and Manning, C. D. 2015 · 2015
Cited alongside, same era.
The zero resource speech challenge 2015: Proposed approaches and results
Versteegh, M.; Anguera, X.; Jansen, A.; and Dupoux, E. 2016 · 2015
Cited alongside, same era.
Listen and translate: A proof of concept for end-to-end speech-to-text translation
Bérard, A.; Pietquin, O.; Servan, C.; and Besacier, L. 2016 · 2016
Cited alongside, same era.
A guide to convolution arithmetic for deep learning
Dumoulin, V.; and Visin, F. 2016 · 2016
Cited alongside, same era.
An attentional model for speech translation without transcription
Duong, L.; Anastasopoulos, A.; Chiang, D.; Bird, S.; and Cohn, T. 2016 · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Oord, A. v. d.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; and Kavukcuoglu, K. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
End-to-End Speech Translation with the Transformer
Vila, L. C.; Escolano, C.; Fonollosa, J. A.; and Costa-jussà, M. R. 2018 · 2018
Later among the works it cites.
The Zero Resource Speech Challenge 2019: TTS Without T
Dunbar, E.; Algayres, R.; Karadayi, J.; Bernard, M.; Benjumea, J.; Cao, X.-N.; Miskic, L.; Dugrain, C.; Ondel, L.; Black, A. W.; and et al. 2019 · 2019
Later among the works it cites.
Unsupervised acoustic unit discovery for speech synthesis using discrete latent-variable neural networks
Eloff, R.; Nortje, A.; van Niekerk, B.; Govender, A.; Nortje, L.; Pretorius, A.; van Biljon, E.; van der Westhuizen, E.; van Staden, L.; and Kamper, H. 2019 · 2019
Later among the works it cites.
Direct speech-to-speech translation with a sequence-to-sequence model
Jia, Y.; Weiss, R. J.; Biadsy, F.; Macherey, W.; Johnson, M.; Chen, Z.; and Wu, Y. 2019 · 2019
Later among the works it cites.
Almost Unsupervised Text to Speech and Automatic Speech Recognition
Ren, Y.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; and Liu, T.-Y. 2019 · 2019
Later among the works it cites.
Attention-Passing Models for Robust and Data-Efficient End-to-End Speech Translation
Sperber, M.; Neubig, G.; Niehues, J.; and Waibel, A. 2019 · 2019
Later among the works it cites.
VQVAE Unsupervised Unit Discovery and Multi-scale Code2Spec Inverter for Zerospeech Challenge 2019
Tjandra, A.; Sisman, B.; Zhang, M.; Sakti, S.; Li, H.; and Nakamura, S. 2019 · 2019
Later among the works it cites.
Towards zero-shot learning for automatic phonemic transcription
Li, X.; Dalmia, S.; Mortensen, D. R.; Li, J.; Black, A. W.; and Metze, F. 2020 · 2020
Closest in time.
Speech technology for unwritten languages
Scharenborg, O.; Besacier, L.; Black, A.; Hasegawa-Johnson, M.; Metze, F.; Neubig, G.; Stüker, S.; Godard, P.; Müller, M.; Ondel, L.; et al. 2020 · 2020
Closest in time.
Unsupervised speech representation learning using wavenet autoencoders
Chorowski, J.; Weiss, R. J.; Bengio, S.; and van den Oord, A. 2019 · 2053
Closest in time.