Fetching the paper…
Reading the bibliography…
Transformer has obtained promising results on cognitive speech signal processing field, which is of interest in various applications ranging from emotion to neurocognitive disorder analysis.
H. Robbins and S. Monro, “A stochastic approximation method,” The annals of mathematical statistics , pp. 400–407, 1951
1951
Earlier work this paper cites.
K. E. Goodglass H, “The boston diagnostic aphasia examination,” Lea & Febinger, Philadelphia , 1983
1983
Earlier work this paper cites.
J. T. Becker, F. Boiler, O. L. Lopez, J. Saxton, and K. L. McGonigle, “The natural history of alzheimer’s disease: Description of study cohort and accuracy of diagnosis,” Archives of Neurology , vol. 51, no. 6, pp. 585–594, 1994
1994
Earlier work this paper cites.
K. Tokuda, T. Kobayashi, and S. Imai, “Speech parameter generation from HMM using dynamic features,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 1995, pp. 660–663
1995
Earlier work this paper cites.
W. Reichl and W. Chou, “Robust decision tree state tying for continuous speech recognition,” IEEE Transactions on Speech and Audio Processing , vol. 8, no. 5, pp. 555–566, 2000
2000
Earlier work this paper cites.
B. Schuller, G. Rigoll, and M. Lang, “Hidden Markov model-based speech emotion recognition,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2003, pp. II–1
2003
Earlier work this paper cites.
B. Moore, L. Tyler, and W. Marslen-Wilson, “Introduction. the perception of speech: from sound to meaning,” Philosophical transactions of the Royal Society of London. Series B, Biological sciences , vol. 363, no. 1493, pp. 917–921, Mar. 2008
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” Language Resources and Evaluation , vol. 42, no. 4, pp. 335–359, 2008
2008
Earlier work this paper cites.
J. Yuan and M. Y. Liberman, “Speaker identification on the scotus corpus,” Journal of the Acoustical Society of America , vol. 123, pp. 3878–3878, 2008
2008
Earlier work this paper cites.
A. Stuhlsatz, C. Meyer, F. Eyben, T. Zielke, G. Meier, and B. Schuller, “Deep neural networks for acoustic emotion recognition: Raising the benchmarks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2011, pp. 5688–5691
2011
Earlier work this paper cites.
K. Poon-Feng, D.-Y. Huang, M. Dong, and H. Li, “Acoustic emotion recognition based on fusion of multiple feature-dependent deep boltzmann machines,” in The 9th International Symposium on Chinese Spoken Language Processing , 2014, pp. 584–588
2014
Earlier work this paper cites.
J. Gratch, R. Artstein, G. Lucas, G. Stratou, S. Scherer, A. Nazarian, R. Wood, J. Boberg, D. DeVault, S. Marsella, D. Traum, S. Rizzo, and L.-P. Morency, “The distress analysis interview corpus of human and computer interviews,” in Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14) , May 2014, pp. 3123–3128
2014
Earlier work this paper cites.
J. Lee and I. Tashev, “High-level feature representation using recurrent neural network for speech emotion recognition,” in Proc. Interspeech , 2015, pp. 1537–1540
2015
Cited alongside, same era.
L. Yang, D. Jiang, L. He, E. Pei, M. C. Oveneke, and H. Sahli, “Decision tree based depression classification from audio video and language information,” in Proceedings of the 6th International Workshop on Audio/Visual Emotion Challenge , 2016, pp. 89–96
2016
Cited alongside, same era.
2016
Cited alongside, same era.
M. Valstar, J. Gratch, B. Schuller, F. Ringeval, D. Lalanne, M. Torres Torres, S. Scherer, G. Stratou, R. Cowie, and M. Pantic, “Avec 2016: Depression, mood, and emotion recognition workshop and challenge,” in Proceedings of the 6th International Workshop on Audio/Visual Emotion Challenge , 2016, pp. 3–10
2016
M. R. Makiuchi, T. Warnita, N. Inoue, K. Shinoda, M. Yoshimura, M. Kitazawa, K. Funaki, Y. Eguchi, and T. Kishimoto, “Speech paralinguistic approach for detecting dementia using gated convolutional neural network,” IEICE transactions on Information and Systems , vol. 104, no. 11, pp. 1930–1940, 2021
2021
Later among the works it cites.
H. Solieman and E. A. Pustozerov, “The detection of depression using multimodal models based on text and voice quality features,” in IEEE Conference of Russian Young Researchers in Electrical and Electronic Engineering , 2021, pp. 1843–1848
2021
Later among the works it cites.
S. T. Rajamani, K. T. Rajamani, A. Mallol-Ragolta, S. Liu, and B. Schuller, “A novel attention-based gated recurrent unit and its efficacy in speech emotion recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 6294–6298
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
S. Mirsamadi, E. Barsoum, and C. Zhang, “Automatic speech emotion recognition using recurrent neural networks with local attention,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 2227–2231
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , 2017, pp. 5998–6008
2017
Cited alongside, same era.
2019
Cited alongside, same era.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition.” in Proc. Interspeech , 2019, pp. 3465–3469
2019
Cited alongside, same era.
J. Liang, R. Li, and Q. Jin, “Semi-supervised multi-modal emotion recognition with cross-modal distribution matching,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 2852–2861
2020
Cited alongside, same era.
L. Guo, L. Wang, C. Xu, J. Dang, E. S. Chng, and H. Li, “Representation learning with spectro-temporal-channel attention for speech emotion recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 6304–6308
2021
Cited alongside, same era.
N. Seneviratne and C. Espy-Wilson, “Generalized dilated CNN models for depression detection using inverted vocal tract variables,” in Proc. Interspeech , 2021, pp. 4513–4517
2021
Cited alongside, same era.
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
X. Wang, M. Wang, W. Qi, W. Su, X. Wang, and H. Zhou, “A novel end-to-end speech emotion recognition network with stacked transformer layers,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 6289–6293
2021
Later among the works it cites.
Z. Lian, B. Liu, and J. Tao, “Ctnet: Conversational transformer network for emotion recognition,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 985–1000, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Yin, Y. Gu, L. Yao, Y. Zhou, X. Liang, and H. Zhang, “Progressive co-teaching for ambiguous speech emotion recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 6264–6268
2021
Later among the works it cites.
F. Bertini, D. Allevi, G. Lutero, L. Calzà, and D. Montesi, “An automatic alzheimer’s disease classifier based on spontaneous spoken english,” Computer Speech & Language , vol. 72, p. 101298, 2022
2022
Closest in time.