Fetching the paper…
Reading the bibliography…
Progress in speech processing has been facilitated by shared datasets and benchmarks.
“The ATIS spoken language systems pilot corpus,”
C. T. Hemphill, J. J. Godfrey, and G. R. Doddington, · 1990
Earlier work this paper cites.
“OntoNotes: the 90% solution,”
E. Hovy, M. Marcus, M. Palmer, L. Ramshaw, and R. Weischedel, · 2006
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“IEMOCAP: Interactive emotional dyadic motion capture database,”
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, · 2008
Earlier work this paper cites.
“The NXT-format Switchboard corpus: a rich resource for investigating the syntax, semantics, pragmatics and prosody of dialogue,”
S. Calhoun, J. Carletta, J. M. Brenier, N. Mayo, D. Jurafsky, M. Steedman, and D. Beaver, · 2010
Earlier work this paper cites.
“A joint model for entity analysis: Coreference, typing, and linking,”
G. Durrett and D. Klein, · 2014
Earlier work this paper cites.
“A Practical Guide to Sentiment Annotation: Challenges and Solutions,”
S. Mohammad, · 2016
Earlier work this paper cites.
“VoxCeleb: A large-scale speaker identification dataset,”
A. Nagraniy, J. S. Chungy, and A. Zisserman, · 2017
Earlier work this paper cites.
“Unsupervised learning approach to feature analysis for automatic speech emotion recognition,”
S. E. Eskimez, Z. Duan, and W. Heinzelman, · 2018
Earlier work this paper cites.
“Representation Learning with Contrastive Predictive Coding,”
A. van den Oord, Y. Li, and O. Vinyals, · 2018
Earlier work this paper cites.
A. Coucke, A. Saade, A. Ball, T. Bluche, A. Caulier, D. Leroy, C. Doumouro, T. Gisselbrecht, F. Caltagirone, T. Lavril, et al., · 2018
Earlier work this paper cites.
“GLUE: A multi-task benchmark and analysis platform for natural language understanding,”
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, · 2018
Earlier work this paper cites.
“Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph,”
A. Zadeh, P. P. Liang, J. Vanbriesen, S. Poria, E. Tong, E. Cambria, M. Chen, and L. P. Morency, · 2018
Earlier work this paper cites.
“TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation,”
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Estève, · 2018
Earlier work this paper cites.
“An unsupervised autoregressive model for speech representation learning,”
Y. A. Chung, W. N. Hsu, H. Tang, and J. Glass, · 2019
Earlier work this paper cites.
“Learning problem-agnostic speech representations from multiple self-supervised tasks,”
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, and Y. Bengio, · 2019
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, · 2019
Cited alongside, same era.
“RoBERTa: A robustly optimized BERT pretraining approach,”
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, · 2019
Cited alongside, same era.
“wav2vec: Unsupervised pre-training for speech recognition,”
S. Schneider, A. Baevski, R. Collobert, and M. Auli, · 2019
Cited alongside, same era.
“Improving Transformer-based Speech Recognition Using Unsupervised Pre-training,”
D. Jiang, X. Lei, W. Li, N. Luo, Y. Hu, W. Zou, and X. Li, · 2019
Cited alongside, same era.
“Large-scale unsupervised pre-training for end-to-end spoken language understanding,”
P. Wang, L. Wei, Y. Cao, J. Xie, and Z. Nie, · 2020
Later among the works it cites.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,”
A. T. Liu, S. W. Yang, P. H. Chi, P. C. Hsu, and H. Y. Lee, · 2020
Later among the works it cites.
“A large scale speech sentiment corpus,”
E. Y. Chen, Z. Lu, H. Xu, L. Cao, Y. Zhang, and J. Fan, · 2020
Later among the works it cites.
“The MSP-Conversation corpus.,”
L. Martinez-Lucas, M. Abdelwahab, and C. Busso, · 2020
Later among the works it cites.
“End-to-end named entity recognition from English speech,”
H. Yadav, S. Ghosh, Y. Yu, and R. R. Shah, · 2020
Later among the works it cites.
“Speech Sentiment Analysis via Pre-Trained Features from End-to-End ASR Models,”
Z. Lu, L. Cao, Y. Zhang, C. C. Chiu, and J. Fan, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Learning speaker representations with mutual information,”
M. Ravanelli and Y. Bengio, · 2019
Cited alongside, same era.
“Speech model pre-training for end-to-end spoken language understanding,”
L. Lugosch, M. Ravanelli, P. Ignoto, V. S. Tomar, and Y. Bengio, · 2019
Cited alongside, same era.
“Audio de-identification: A new entity recognition task,”
I. Cohn, I. Laish, G. Beryozkin, G. Li, I. Shafran, I. Szpektor, T. Hartman, A. Hassidim, and Y. Matias, · 2019
Cited alongside, same era.
“End-To-End Named Entity and Semantic Concept Extraction from Speech,”
S. Ghannay, A. Caubriere, Y. Esteve, N. Camelin, E. Simonnet, A. Laurent, and E. Morin, · 2019
Cited alongside, same era.
“DeBERTa: Decoding-enhanced BERT with Disentangled Attention,”
P. He, X. Liu, J. Gao, and W. Chen, · 2020
Cited alongside, same era.
“Deep contextualized acoustic representations for semi-supervised speech recognition,”
S. Ling, Y. Liu, J. Salazar, and K. Kirchhoff, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, · 2020
Cited alongside, same era.
Later among the works it cites.
“Transformers: State-of-the-art natural language processing,”
T. Wolf, J. Chaumond, L. Debut, V. Sanh, C. Delangue, A. Moi, P. Cistac, M. Funtowicz, J. Davison, S. Shleifer, and Others, · 2020
Later among the works it cites.
“SUPERB: Speech processing universal performance benchmark,”
S.-w. Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, et al., · 2021
Closest in time.
“ASR-GLUE: A New Multi-task Benchmark for ASR-Robust Natural Language Understanding,”
L. Feng, J. Yu, D. Cai, S. Liu, H. Zheng, and Y. Wang, · 2021
Closest in time.
“HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,”
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, · 2021
Closest in time.
“Performance-Efficiency Trade-offs in Unsupervised Pre-training for Speech Recognition,”
F. Wu, K. Kim, J. Pan, K. Han, K. Q. Weinberger, and Y. Artzi, · 2021
Closest in time.
“Contrastive unsupervised learning for speech emotion recognition,”
M. Li, B. Yang, J. Levy, A. Stolcke, V. Rozgic, S. Matsoukas, C. Papayiannis, D. Bone, and C. Wang, · 2021
Closest in time.
“Audio Albert: A Lite Bert for Self-Supervised Learning of Audio Representation,”
P. H. Chi, P. H. Chung, T. H. Wu, C. C. Hsieh, Y. H. Chen, S. W. Li, and H. Y. Lee, · 2021
Closest in time.
“Semi-supervised spoken language understanding via self-supervised speech and language model pretraining,”
C. I. Lai, Y. S. Chuang, H. Y. Lee, S. W. Li, and J. Glass, · 2021
Closest in time.
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, · 2021
Closest in time.