Fetching the paper…
Reading the bibliography…
This paper introduces SpeeChain, an open-source Pytorch-based toolkit designed to develop the machine speech chain for large-scale use.
“Dropout: a simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Listening while speaking: Speech chain by deep learning,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2017
Earlier work this paper cites.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Earlier work this paper cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Earlier work this paper cites.
“Machine speech chain with one-shot speaker adaptation,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2018
Earlier work this paper cites.
“Training neural speech recognition systems with synthetic speech augmentation,”
Jason Li, Ravi Gadde, Boris Ginsburg, and Vitaly Lavrukhin, · 2018
Earlier work this paper cites.
“Espnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, et al., · 2018
Earlier work this paper cites.
“Sequence-to-sequence speech recognition with time-depth separable convolutions,”
Awni Hannun, Ann Lee, Qiantong Xu, and Ronan Collobert, · 2019
Cited alongside, same era.
“Neural speech synthesis with transformer network,”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu, · 2019
Cited alongside, same era.
“Almost unsupervised text to speech and automatic speech recognition,”
Yi Ren, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Cited alongside, same era.
“Speech recognition with augmented synthesized speech,”
Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia, Pedro Moreno, Yonghui Wu, and Zelin Wu, · 2019
Cited alongside, same era.
“Multi-speaker sequence-to-sequence speech synthesis for data augmentation in acoustic-to-word speech recognition,”
Sei Ueno, Masato Mimura, Shinsuke Sakai, and Tatsuya Kawahara, · 2019
Cited alongside, same era.
“Generating synthetic audio data for attention-based speech recognition systems,”
Nick Rossenbach, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2020
Later among the works it cites.
“Espnet-tts: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit,”
Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, and Xu Tan, · 2020
Later among the works it cites.
“Self-training for end-to-end speech recognition,”
Jacob Kahn, Ann Lee, and Awni Hannun, · 2020
Later among the works it cites.
“Comparing the benefit of synthetic training data for various automatic speech recognition architectures,”
Nick Rossenbach, Mohammad Zeineldeen, Benedikt Hilmes, Ralf Schlüter, and Hermann Ney, · 2021
Later among the works it cites.
“Exploring machine speech chain for domain adaptation and few-shot speaker adaptation,”
Fengpeng Yue, Yan Deng, Lei He, and Tom Ko, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli, · 2019
Cited alongside, same era.
“Libritts: A corpus derived from librispeech for text-to-speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Cited alongside, same era.
“Multispeech: Multi-speaker text to speech with transformer,”
Mingjian Chen, Xu Tan, Yi Ren, Jin Xu, Hao Sun, Sheng Zhao, Tao Qin, and Tie-Yan Liu, · 2020
Cited alongside, same era.
“Machine speech chain,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2020
Cited alongside, same era.
“Improving speech recognition using consistent predictions on synthesized speech,”
Gary Wang, Andrew Rosenberg, Zhehuai Chen, Yu Zhang, Bhuvana Ramabhadran, Yonghui Wu, and Pedro Moreno, · 2020
Cited alongside, same era.
“Eat: Enhanced asr-tts for self-supervised speech recognition,”
Murali Karthick Baskar, Lukáš Burget, Shinji Watanabe, Ramon Fernandez Astudillo, et al., · 2021
Later among the works it cites.
Changhan Wang, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Ann Lee, Peng-Jen Chen, Jiatao Gu, and Juan Pino, · 2021
Later among the works it cites.
“Speechbrain: A general-purpose speech toolkit,”
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, et al., · 2021
Later among the works it cites.
“Paddlespeech: An easy-to-use all-in-one speech toolkit,”
Hui Zhang, Tian Yuan, Junkun Chen, Xintong Li, Renjie Zheng, Yuxin Huang, Xiaojie Chen, Enlei Gong, Zeyu Chen, Xiaoguang Hu, et al., · 2022
Later among the works it cites.
“Momentum pseudo-labeling: Semi-supervised asr with continuously improving pseudo-labels,”
Yosuke Higuchi, Niko Moritz, Jonathan Le Roux, and Takaaki Hori, · 2022
Later among the works it cites.