Fetching the paper…
Reading the bibliography…
We present an attention-based sequence-to-sequence neural network which can directly translate speech from one language into speech in another language, without relying on an intermediate text representation.
D. Griffin and J. Lim, “Signal estimation from modified short-time Fourier transform,”
1984
Earlier work this paper cites.
A. Lavie, A. Waibel, L. Levin, M. Finke, D. Gates, M. Gavalda, T. Zeppenfeld, and P. Zhan, “JANUS-III: Speech-to-speech translation in multiple languages,” in
1997
Earlier work this paper cites.
E. Vidal, “Finite-state speech-to-speech translation,” in
1997
Earlier work this paper cites.
H. Ney, “Speech translation: Coupling of recognition and translation,” in
1999
Earlier work this paper cites.
W. Wahlster,
2000
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: A method for automatic evaluation of machine translation,” in
2002
Earlier work this paper cites.
F. Casacuberta, H. Ney, F. J. Och, E. Vidal, J. M. Vilar
2004
Earlier work this paper cites.
E. Matusov, S. Kanthak, and H. Ney, “On the integration of speech recognition and statistical machine translation,” in
2005
Earlier work this paper cites.
S. Nakamura, K. Markov, H. Nakaiwa, G.-i. Kikui, H. Kawai, T. Jitsuhiro, J.-S. Zhang, H. Yamamoto, E. Sumita, and S. Yamamoto, “The ATR multilingual speech-to-speech translation system,”
2006
Earlier work this paper cites.
P. Aguero, J. Adell, and A. Bonafonte, “Prosody generation for speech-to-speech translation,” in
2006
Earlier work this paper cites.
M. Kurimo, W. Byrne, J. Dines, P. N. Garner, M. Gibson, Y. Guan, T. Hirsimäki, R. Karhila, S. King, H. Liang
2010
Earlier work this paper cites.
A. F. Machado and M. Queiroz, “Voice conversion: A critical survey,” in
2010
Earlier work this paper cites.
M. Wester, J. Dines, M. Gibson, H. Liang
2010
Earlier work this paper cites.
M. Post, G. Kumar, A. Lopez, D. Karakos, C. Callison-Burch
2013
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
International Telecommunication Union, “ITU-T F.745: Functional requirements for network-based speech-to-speech translation services,” 2016
2016
Cited alongside, same era.
A. Bérard, O. Pietquin, C. Servan, and L. Besacier, “Listen and translate: A proof of concept for end-to-end speech-to-text translation,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Q. T. Do, S. Sakti, and S. Nakamura, “Toward expressive speech translation: a unified sequence-to-sequence LSTMs approach for translating words and emphasis,” in
2017
Cited alongside, same era.
R. J. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z. Chen, “Sequence-to-sequence models can directly translate foreign speech,” in
2017
C.-C. Chiu, T. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. Weiss, K. Rao
2018
Later among the works it cites.
N. Shazeer and M. Stern, “Adafactor: Adaptive learning rates with sublinear memory cost,” in
2018
Later among the works it cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. v. d. Oord, S. Dieleman
2018
Later among the works it cites.
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche
2018
Later among the works it cites.
Y. Wang, D. Stanton, Y. Zhang, R. Skerry-Ryan, E. Battenberg, J. Shor
2018
Later among the works it cites.
Y. Chen, Y. Assael, B. Shillingford, D. Budden, S. Reed, H. Zen, Q. Wang, L. C. Cobo, A. Trask, B. Laurie
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le
2017
Cited alongside, same era.
D. Krueger, T. Maharaj, J. Kramár, M. Pezeshki, N. Ballas, N. R. Ke, A. Goyal, Y. Bengio
2017
Cited alongside, same era.
T. Kano, S. Takamichi, S. Sakti, G. Neubig, T. Toda, and S. Nakamura, “An end-to-end model for cross-lingual transformation of paralinguistic information,”
2018
Cited alongside, same era.
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf, “Fitting new speakers based on a short untranscribed sample,” in
2018
Cited alongside, same era.
S. O. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” in
2018
Cited alongside, same era.
2019
Closest in time.
Y. Jia, M. Johnson, W. Macherey, R. J. Weiss, Y. Cao, C.-C. Chiu, N. Ari
2019
Closest in time.
J. Zhang, Z. Ling, L.-J. Liu, Y. Jiang, and L.-R. Dai, “Sequence-to-sequence acoustic modeling for voice conversion,”
2019
Closest in time.
F. Biadsy, R. J. Weiss, P. J. Moreno, D. Kanevsky, and Y. Jia, “Parrotron: An end-to-end speech-to-speech conversion model and its applications to hearing-impaired speech and speech separation,” in
2019
Closest in time.
M. Guo, A. Haque, and P. Verma, “End-to-end spoken language translation,”
2019
Closest in time.
A. Zhang, Q. Wang, Z. Zhu, J. Paisley, and C. Wang, “Fully supervised speaker diarization,” in
2019
Closest in time.
J. Shen, P. Nguyen, Y. Wu, Z. Chen
2019
Closest in time.
2019
Closest in time.
Y. Lee and T. Kim, “Robust and fine-grained prosody control of end-to-end speech synthesis,” in
2019
Closest in time.
W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen
2019
Closest in time.