Fetching the paper…
Reading the bibliography…
We present Translatotron 2, a neural direct speech-to-speech translation model that can be trained end-to-end.
Signal estimation from modified short-time Fourier transform
Griffin, D. and Lim, J · 1984
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
JANUS-III: Speech-to-speech translation in multiple languages
Lavie, A., Waibel, A., Levin, L., Finke, M., Gates, D., Gavalda, M., Zeppenfeld, T., and Zhan, P · 1997
Earlier work this paper cites.
Verbmobil: Foundations of speech-to-speech translation
Wahlster, W · 2000
Earlier work this paper cites.
The ATR multilingual speech-to-speech translation system
Nakamura, S., Markov, K., Nakaiwa, H., Kikui, G., Kawai, H., Jitsuhiro, T., Zhang, J.-S., Yamamoto, H., Sumita, E., and Yamamoto, S · 2006
Earlier work this paper cites.
Privacy-preserving speaker verification and identification using Gaussian mixture models
Pathak, M. A. and Raj, B · 2012
Earlier work this paper cites.
Improved speech-to-text translation with the Fisher and Callhome Spanish–English speech translation corpus
Post, M., Kumar, G., Lopez, A., Karakos, D., Callison-Burch, C., and Khudanpur, S · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
LibriSpeech: an ASR corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
ITU-T F.745: Functional requirements for network-based speech-to-speech translation services, 2016
ITU · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Zoneout: Regularizing RNNs by randomly preserving hidden activations
Krueger, D., Maharaj, T., Kramár, J., Pezeshki, M., Ballas, N., Ke, N. R., Goyal, A., Bengio, Y., Courville, A., and Pal, C · 2017
Earlier work this paper cites.
Neural discrete representation learning
Oord, A. v. d., Vinyals, O., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Sequence-to-sequence models can directly translate foreign speech
Weiss, R. J., Chorowski, J., Jaitly, N., Wu, Y., and Chen, Z · 2017
Earlier work this paper cites.
Speech recognition for medical conversations
Chiu, C.-C., Tripathi, A., Chou, K., Co, C., Jaitly, N., Jaunzeikare, D., Kannan, A., Nguyen, P., Sak, H., Sankar, A., Tansuwan, J., Wan, N., Wu, Y., and Zhang, X · 2018
Earlier work this paper cites.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Jia, Y., Zhang, Y., Weiss, R. J., Wang, Q., Shen, J., Ren, F., Chen, Z., Nguyen, P., Pang, R., Moreno, I. L., and Wu, Y · 2018
Earlier work this paper cites.
Efficient neural audio synthesis
Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., Oord, A. v. d., Dieleman, S., and Kavukcuoglu, K · 2018
Earlier work this paper cites.
Parallel WaveNet: Fast high-fidelity speech synthesis
Oord, A., Li, Y., Babuschkin, I., Simonyan, K., Vinyals, O., Kavukcuoglu, K., Driessche, G., Lockhart, E., Cobo, L., Stimberg, F., et al · 2018
Earlier work this paper cites.
Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions
Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerrv-Ryan, R., Saurous, R. A., Agiomyrgiannakis, Y., and Wu, Y · 2018
Earlier work this paper cites.
Generalized end-to-end loss for speaker verification
Wan, L., Wang, Q., Papir, A., and Moreno, I. L · 2018
Cited alongside, same era.
Cross-lingual, multi-speaker text-to-speech synthesis using neural speaker embedding
Chen, M., Chen, M., Liang, S., Ma, J., Chen, L., Wang, S., and Xiao, J · 2019
Cited alongside, same era.
One-to-many multilingual end-to-end speech translation
Di Gangi, M. A., Negri, M., and Turchi, M · 2019
Cited alongside, same era.
Robust sequence-to-sequence acoustic modeling with stepwise monotonic attention for neural TTS
He, M., Deng, Y., and He, L · 2019
Cited alongside, same era.
Recognizing long-form speech using streaming end-to-end models
Narayanan, A., Prabhavalkar, R., Chiu, C.-C., Rybach, D., Sainath, T. N., and Strohman, T · 2019
Cited alongside, same era.
SpecAugment: A simple data augmentation method for automatic speech recognition
Libri-light: A benchmark for ASR with limited or no supervision
Kahn, J., Rivière, M., Zheng, W., Kharitonov, E., Xu, Q., Mazaré, P.-E., Karadayi, J., Liptchinsky, V., Collobert, R., Fuegen, C., Likhomanenko, T., Synnaeve, G., Joulin, A., Mohamed, A., and Dupoux, E · 2020
Later among the works it cites.
SkinAugment: Auto-encoding speaker conversions for automatic speech translation
McCarthy, A. D., Puzon, L., and Pino, J · 2020
Later among the works it cites.
Improved noisy student training for automatic speech recognition
Park, D. S., Zhang, Y., Jia, Y., Han, W., Chiu, C.-C., Li, B., Wu, Y., and Le, Q. V · 2020
Later among the works it cites.
Non-autoregressive neural text-to-speech
Peng, K., Ping, W., Song, Z., and Zhao, K · 2020
Later among the works it cites.
Shen, J., Jia, Y., Chrzanowski, M., Zhang, Y., Elias, I., Zen, H., and Wu, Y · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Park, D. S., Chan, W., Zhang, Y., Chiu, C.-C., Zoph, B., Cubuk, E. D., and Le, Q. V · 2019
Cited alongside, same era.
FastSpeech: Fast, robust and controllable text to speech
Ren, Y., Ruan, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2019
Cited alongside, same era.
Lingvo: A modular and scalable framework for sequence-to-sequence modeling
Shen, J., Nguyen, P., Wu, Y., Chen, Z., et al · 2019
Cited alongside, same era.
Speech-to-speech translation between untranscribed unknown languages
Tjandra, A., Sakti, S., and Nakamura, S · 2019
Cited alongside, same era.
ASVspoof 2019: Future horizons in spoofed and fake audio detection
Todisco, M., Wang, X., Vestman, V., Sahidullah, M., Delgado, H., Nautsch, A., Yamagishi, J., Evans, N., Kinnunen, T., and Lee, K. A · 2019
Cited alongside, same era.
Speech synthesis evaluation – state-of-the-art assessment and suggestion for a novel research program
Wagner, P., Beskow, J., Betz, S., Edlund, J., Gustafson, J., Eje Henter, G., Le Maguer, S., Malisz, Z., Székely, É., Tånnander, C., et al · 2019
Cited alongside, same era.
LibriTTS: A corpus derived from LibriSpeech for text-to-speech
Zen, H., Dang, V., Clark, R., Zhang, Y., Weiss, R. J., Jia, Y., Chen, Z., and Wu, Y · 2019
Cited alongside, same era.
ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Wang, X., Yamagishi, J., Todisco, M., Delgado, H., Nautsch, A., Evans, N., Sahidullah, M., Vestman, V., Kinnunen, T., Lee, K. A., et al · 2020
Later among the works it cites.
Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion
Yi, Z., Huang, W.-C., Tian, X., Yamagishi, J., Das, R. K., Kinnunen, T., Ling, Z., and Toda, T · 2020
Later among the works it cites.
Findings of the IWSLT 2021 evaluation campaign
Anastasopoulos, A., Bojar, O., Bremerman, J., Cattoni, R., Elbayad, M., Federico, M., Ma, X., Nakamura, S., Negri, M., Niehues, J., et al · 2021
Closest in time.
Recent developments on espnet toolkit boosted by conformer
Guo, P., Boyer, F., Chang, X., Hayashi, T., Higuchi, Y., Inaguma, H., Kamo, N., Li, C., Garcia-Romero, D., Shi, J., et al · 2021
Closest in time.
TTS-by-TTS: TTS-driven data augmentation for fast and high-quality speech synthesis
Hwang, M.-J., Yamamoto, R., Song, E., and Kim, J.-M · 2021
Closest in time.
PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS
Jia, Y., Zen, H., Shen, J., Zhang, Y., and Wu, Y · 2021
Closest in time.
Transformer-based direct speech-to-speech translation with transcoder
Kano, T., Sakti, S., and Nakamura, S · 2021
Closest in time.
Direct simultaneous speech to speech translation
Ma, X., Gong, H., Liu, D., Lee, A., Tang, Y., Chen, P.-J., Hsu, W.-N., Heafield, K., Koehn, P., and Pino, J · 2021
Closest in time.
FastSpeech 2: Fast and high-quality end-to-end text-to-speech
Ren, Y., Hu, C., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2021
Closest in time.
Assessing evaluation metrics for speech-to-speech translation
Salesky, E., Mäder, J., and Klinger, S · 2021
Closest in time.
Disentangled speaker and language representations using mutual information minimization and domain adaptation for cross-lingual TTS
Xin, D., Komatsu, T., Takamichi, S., and Saruwatari, H · 2021
Closest in time.
UWSpeech: Speech to speech translation for unwritten languages
Zhang, C., Tan, X., Ren, Y., Qin, T., Zhang, K., and Liu, T.-Y · 2021
Closest in time.
CVSS corpus and massively multilingual speech-to-speech translation
Jia, Y., Tadmor Ramanovich, M., Wang, Q., and Zen, H · 2022
Closest in time.
Direct speech-to-speech translation with discrete units
Lee, A., Chen, P.-J., Wang, C., Gu, J., Ma, X., Polyak, A., Adi, Y., He, Q., Tang, Y., Pino, J., et al · 2022
Closest in time.