Fetching the paper…
Reading the bibliography…
This paper investigates the use of word surprisal, a measure of the predictability of a word in a given context, as a feature to aid speech synthesis prosody.
R. P. Rao and D. H. Ballard, “Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects,” Nature neuroscience , vol. 2, no. 1, pp. 79–87, 1999
1999
Earlier work this paper cites.
S. Pan and K. McKeown, “Word informativeness and automatic pitch accent modeling,” 1999
1999
Earlier work this paper cites.
J. Hale, “A probabilistic earley parser as a psycholinguistic model,” in Proceedings of the second meeting of the North American Chapter of the Association for Computational Linguistics on Language technologies . Association for Computational Linguistics, 2001, pp. 1–8
2001
Earlier work this paper cites.
D. C. Knill and A. Pouget, “The bayesian brain: the role of uncertainty in neural coding and computation,” TRENDS in Neurosciences , vol. 27, no. 12, pp. 712–719, 2004
2004
Earlier work this paper cites.
D. Watson, J. E. Arnold, and M. K. Tanenhaus, “Acoustic prominence and reference accessibility in language production,” Target , vol. 73, no. 74, p. 75, 2006
2006
Earlier work this paper cites.
V. K. R. Sridhar, A. Nenkova, S. Narayanan, and D. Jurafsky, “Detecting prominence in conversational speech: pitch accent, givenness and focus,” in Proceedings of Speech Prosody , vol. 453. Citeseer, 2008, p. 456
2008
Earlier work this paper cites.
C. Féry and S. Ishihara, “How focus and givenness shape prosody,” Information structure: Theoretical, typological, and experimental perspectives , pp. 36–63, 2010
2010
Earlier work this paper cites.
S. Kakouros and O. Räsänen, “Automatic detection of sentence prominence in speech using predictability of word-level acoustic features,” in Sixteenth Annual Conference of the International Speech Communication Association , 2015
2015
Earlier work this paper cites.
J. Cole, “Prosody in context: A review,” Language, Cognition and Neuroscience , vol. 30, no. 1-2, pp. 1–31, 2015
2015
Earlier work this paper cites.
S. Kakouros and O. Räsänen, “Perception of sentence stress in speech correlates with the temporal unpredictability of prosodic features,” Cognitive science , vol. 40, no. 7, pp. 1739–1774, 2016
2016
Earlier work this paper cites.
S. Kakouros, J. Pelemans, L. Verwimp, P. Wambacq, and O. Räsänen, “Analyzing the contribution of top-down lexical and bottom-up acoustic cues in the detection of sentence prominence,” Proceedings Interspeech 2016 , vol. 8, pp. 1074–1078, 2016
2016
Cited alongside, same era.
A. Zarcone, M. Van Schijndel, J. Vogels, and V. Demberg, “Salience and attention in surprisal-based accounts of language processing,” Frontiers in psychology , vol. 7, p. 844, 2016
2016
Cited alongside, same era.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Cited alongside, same era.
A. Suni, J. Šimko, D. Aalto, and M. Vainio, “Hierarchical representation and estimation of prosody using continuous wavelet transform,” Computer Speech & Language , vol. 45, pp. 123–136, 2017
2017
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Later among the works it cites.
T. Kenter, M. K. Sharma, and R. Clark, “Improving prosody of rnn-based english text-to-speech synthesis by incorporating a bert model,” 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Kakouros, N. Salminen, and O. Räsänen, “Making predictable unpredictable with style–behavioral and electrophysiological evidence for the critical role of prosodic expectations in the perception of prominence in speech,” Neuropsychologia , vol. 109, pp. 181–199, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Goodkind and K. Bicknell, “Predictive power of word surprisal for reading times is a linear function of language model quality,” in Proceedings of the 8th workshop on cognitive modeling and computational linguistics (CMCL 2018) , 2018, pp. 10–18
2018
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
[Online]. Available: https://platform.openai.com/docs/models/gpt-3-5
Cited in the paper.
[Online]. Available: https://platform.openai.com/docs/models/gpt-4
Cited in the paper.
F. Kügler and S. Calhoun, “Prosodic encoding of information structure: A typological perspective,” 2020
2020
Later among the works it cites.
A. Łańcucki, “Fastpitch: Parallel text-to-speech with pitch prediction,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6588–6592
2021
Later among the works it cites.
J. Šimko, A. Adigwe, A. Suni, and M. Vainio, “A hierarchical predictive processing approach to modelling prosody,” Proceedings of Speech Prosody 2022 , 2022
2022
Later among the works it cites.
R. Badlani, A. Łańcucki, K. J. Shih, R. Valle, W. Ping, and B. Catanzaro, “One tts alignment to rule them all,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6092–6096
2022
Later among the works it cites.
2023
Closest in time.
B.-D. Oh and W. Schuler, “Why does surprisal from larger transformer-based language models provide a poorer fit to human reading times?” Transactions of the Association for Computational Linguistics , vol. 11, pp. 336–350, 2023
2023
Closest in time.