Fetching the paper…
Reading the bibliography…
Inspired by a human speech chain mechanism, a machine speech chain framework based on deep learning was recently proposed for the semi-supervised development of automatic speech recognition (ASR) and text-to-speech synthesis TTS) systems.
M. Badian, E. Appel, D. Palm, W. Rupp, W. Sittig, and K. Taeuber, “Standardized mental stress in healthy volunteers induced by delayed auditory feedback (DAF),”
1979
Earlier work this paper cites.
D. B. Paul and J. M. Baker, “The design for the Wall Street Journal-based CSR corpus,” in
1992
Earlier work this paper cites.
P. B. Denes and E. N. Pinson,
1993
Earlier work this paper cites.
J. S. Perkell, M. Matthies, H. Lane, F. H. Guenther, R. Wilhelms-Tricarico, J. Wozniak, and P. Guiod, “Speech motor control: Acoustic goals, saturation effects, auditory feedback and internal models,”
1997
Earlier work this paper cites.
M. Saraclar, M. Riley, E. Bocchieri, and V. Goffin, “Towards automatic closed captioning : Low latency real time broadcast news transcription,” in
2002
Earlier work this paper cites.
C. Fügen, A. H. Waibel, and M. Kolss, “Simultaneous translation of lectures and speeches,”
2007
Earlier work this paper cites.
H. Sak, M. Saraclar, and T. Gungor, “On-the-fly lattice rescoring for real-time automatic speech recognition,” in
2010
Earlier work this paper cites.
2012
Earlier work this paper cites.
T. Baumann and D. Schlangen, “Evaluating prosodic processing for incremental speech synthesis,” in
2012
Earlier work this paper cites.
A. Graves, “Supervised sequence labelling,” in
2012
Earlier work this paper cites.
T. Baumann, “Decision tree usage for incremental parametric speech synthesis,” in
2014
Earlier work this paper cites.
T. Mieno, G. Neubig, S. Sakti, T. Toda, and S. Nakamura, “Speed or accuracy? a study in evaluation of simultaneous speech translation,” in
2015
Earlier work this paper cites.
M. Pouget, T. Hueber, G. Bailly, and T. Baumann, “HMM training strategy for incremental speech synthesis,” in
2015
Cited alongside, same era.
2015
Cited alongside, same era.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in
2016
Cited alongside, same era.
N. Jaitly, Q. V. Le, O. Vinyals, I. Sutskever, D. Sussillo, and S. Bengio, “An online sequence-to-sequence model using partial conditioning,” in
2016
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio
2017
Cited alongside, same era.
T. Yanagita, S. Sakti, and S. Nakamura, “Incremental TTS for Japanese language,” in
2018
Later among the works it cites.
A. Tjandra, S. Sakti, and S. Nakamura, “Multi-scale alignment and contextual history for attention mechanism in sequence-to-sequence model,” in
2018
Later among the works it cites.
T. N. Sainath, C.-C. Chiu, R. Prabhavalkar, A. Kannan, Y. Wu, P. Nguyen, and Z. Chen, “Improving the performance of online neural transducer models,” in
2018
Later among the works it cites.
N.-Q. Pham, T.-S. Nguyen, J. Niehues, M. Müller, and A. Waibel, “Very deep self-attention networks for end-to-end speech recognition,” in
2019
Later among the works it cites.
M. Chen, M. Chen, S. Liang, J. Ma, L. Chen, S. Wang, and J. Xiao, “Cross-lingual, multi-speaker text-to-speech synthesis using neural speaker embedding,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
——, “Listening while speaking: Speech chain by deep learning,” in
2017
Cited alongside, same era.
C. Lavania and J. Bilmes, “Reducing total latency in online real-time inference and decoding via combined context window and model smoothing latencies,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
C.-C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, E. Gonina
2018
Cited alongside, same era.
——, “Machine speech chain with one-shot speaker adaptation,” in
2018
Cited alongside, same era.
V. Peddinti, Y. Wang, D. Povey, and S. Khudanpur, “Low latency acoustic modeling using temporal convolution and LSTMs,”
2018
Cited alongside, same era.
——, “End-to-end feedback loss in speech chain framework via straight-through estimator,” in
2019
Later among the works it cites.
S. Nakamura, K. Sudoh, and S. Sakti, “Towards machine speech-to-speech translation,”
2019
Later among the works it cites.
S. Novitasari, A. Tjandra, S. Sakti, and S. Nakamura, “Sequence-to-sequence learning via attention transfer for incremental speech recognition,” in
2019
Later among the works it cites.
——, “Neural iTTS: Toward synthesizing speech in real-time with end-to-end neural text-to-speech framework,” in
2019
Later among the works it cites.
M. Ma, B. Zheng, K. Liu, R. Zheng, H. Liu, K. Peng, K. Church, and L. Huang, “Incremental text-to-speech synthesis with prefix-to-prefix framework,” 2019
2019
Later among the works it cites.
A. Tjandra, S. Sakti, and S. Nakamura, “Machine speech chain,”
2020
Closest in time.