Fetching the paper…
Reading the bibliography…
Despite the close relationship between speech perception and production, research in automatic speech recognition (ASR) and text-to-speech synthesis (TTS) has progressed more or less independently without exerting much mutual influence on each other.
K. H. Davis, R. Biddulph, and S. Balashek, “Automatic recognition of spoken digits,” Acoustic Society of America , pp. 627–642, 1952
1952
Earlier work this paper cites.
T. K. Vintsyuk, “Speech discrimination by dynamic programming,” Kibernetika , pp. 81–88, 1968
1968
Earlier work this paper cites.
F. Jelinek, “Continuous speech recognition by statistical methods,” IEEE , vol. 64, pp. 532–536, 1976
1976
Earlier work this paper cites.
J. P. Olive, “Rule synthesis of speech from dyadic units,” in Proceedings of ICASSP , 1977, pp. 568–570
1977
Earlier work this paper cites.
H. Sakoe and S. Chiba, “Dynamic programming algorithm quantization for spoken word recognition,” IEEE Transaction on Acoustics, Speech and Signal Processing , vol. ASSP-26, no. 1, pp. 43–49, 1978
1978
Earlier work this paper cites.
D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
Y. Sagisaka, “Speech synthesis by rule using an optimal selection of non-uniform synthesis units,” in Proceedings of ICASSP , 1988
1988
Earlier work this paper cites.
J. G. Wilpon, L. R. Rabiner, C. H. Lee, and E. R. Goldman, “Automatic recognition of keywords in unconstrained speech using hidden Markov models,” IEEE Transaction on Acoustics, Speech and Signal Processing , vol. 38, no. 11, pp. 1870–1878, 1990
1990
Earlier work this paper cites.
Y. Sagisaka, N. Kaiki, N. Iwahashi, and K. Mimura, “Atr υ \upsilon -talk speech,” in Proceedings of ICSLP , 1992, pp. 483–486
1992
Earlier work this paper cites.
P. Denes and E. Pinson, The Speech Chain , ser. Anchor books. Worth Publishers, 1993. [Online]. Available: https://books.google.co.jp/books?id=ZMTm3nlDfroC
1993
Earlier work this paper cites.
K. Tokuda, T. Kobayashi, and S. Imai, “Speech parameter generation from HMM using dynamic features,” in Proceedings of ICASSP , 1995, pp. 660––663
1995
Earlier work this paper cites.
A. Hunt and A. Black, “Unit selection in a concatenative speech synthesis system using a large speech database,” in Proceedings of ICASSP , 1996, pp. 373–376
1996
Earlier work this paper cites.
T. Yoshimura, K. Tokuda, T. Masuko, T. Kobayashi, and T. Kitamura, “Simultaneous modeling of spectrum, pitch and duration in HMM-based speech synthesis,” in Proceedings of Eurospeech , 1999, pp. 2347––2350
1999
Cited alongside, same era.
G. Kikui, E. Sumita, T. Takezawa, and S. Yamamoto, “Creating corpora for speech-to-speech translation,” in Eighth European Conference on Speech Communication and Technology , 2003
2003
Cited alongside, same era.
S. Sakti, M. Paul, A. Finch, X. Hu, J. Ni, N. Kimura, S. Matsuda, C. Hori, Y. Ashikari, H. Kawai, H. Kashioka, E. Sumita, and S. Nakamura, “Distributed speech translation technologies for multiparty multilingual communication,” ACM Trans. Speech Lang. Process. , vol. 9, no. 2, pp. 4:1–4:27, Aug. 2012. [Online]. Available: http://doi.acm.org/10.1145/2287710.2287712
2012
Cited alongside, same era.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in Acoustics, speech and signal processing (icassp), 2013 ieee international conference on . IEEE, 2013, pp. 6645–6649
T. N. Sainath, R. J. Weiss, A. W. Senior, K. W. Wilson, and O. Vinyals, “Learning the speech front-end with raw waveform cldnns.” in Interspeech , vol. 2015, 2015
2015
Later among the works it cites.
2015
Later among the works it cites.
2015
Later among the works it cites.
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
H. Zen, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2013, pp. 7962–7966
2013
Cited alongside, same era.
2013
Cited alongside, same era.
2014
Cited alongside, same era.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence-to-Sequence learning with neural networks,” in Advances in neural information processing systems , 2014, pp. 3104–3112
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2014
Cited alongside, same era.
D. Palaz, M. M. Doss, and R. Collobert, “Convolutional neural networks-based continuous speech recognition using raw speech signal,” in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on . IEEE, 2015, pp. 4295–4299
2015
Cited alongside, same era.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on . IEEE, 2016, pp. 4960–4964
2016
Later among the works it cites.
D. He, Y. Xia, T. Qin, L. Wang, N. Yu, T. Liu, and W.-Y. Ma, “Dual learning for machine translation,” in Advances in Neural Information Processing Systems , 2016, pp. 820–828
2016
Later among the works it cites.
2016
Later among the works it cites.
2017
Closest in time.
2017
Closest in time.
B. McFee, M. McVicar, O. Nieto, S. Balke, C. Thome, D. Liang, E. Battenberg, J. Moore, R. Bittner, R. Yamamoto, and et al., “librosa 0.5.0,” Feb 2017
2017
Closest in time.
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015 , 2015, pp. 2048–2057
2057
Closest in time.