Fetching the paper…
Reading the bibliography…
The foundation model paradigm leverages a shared foundation model to achieve state-of-the-art (SOTA) performance for various tasks, requiring minimal downstream-specific modeling and data annotation.
L.-J. Liu, Z.-H. Ling, Y. Jiang, M. Zhou, and L.-R. Dai, “Wavenet vocoder with limited training data for voice conversion.” in Interspeech , 2018, pp. 1983–1987
1987
Earlier work this paper cites.
D. S. Pallet, W. M. Fisher, and J. G. Fiscus, “Tools for the analysis of benchmark speech recognition tests,” in International Conference on Acoustics, Speech, and Signal Processing . IEEE, 1990, pp. 97–100
1990
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “Darpa timit acoustic-phonetic continous speech corpus cd-rom. nist speech disc 1-1.1,” NASA STI/Recon technical report n , vol. 93, p. 27403, 1993
1993
Earlier work this paper cites.
S. Bengio and J. Mariéthoz, “The expected performance curve: a new assessment measure for person authentication,” in Proc. The Speaker and Language Recognition Workshop (Odyssey 2004) , 2004, pp. 279–284
2004
Earlier work this paper cites.
P. Koehn, “Statistical significance tests for machine translation evaluation,” in Proceedings of the 2004 conference on empirical methods in natural language processing , 2004, pp. 388–395
2004
Earlier work this paper cites.
J. W. Du Bois, W. L. Chafe, C. Meyer, S. A. Thompson, and N. Martey, “Santa Barbara corpus of spoken American English,” CD-ROM. Philadelphia: Linguistic Data Consortium , 2000 – 2005
2005
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
C. Busso et al. , “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, no. 4, pp. 335–359, 2008
2008
Earlier work this paper cites.
T. Giorgino, “Computing and visualizing dynamic time warping alignments in R: The dtw package,” Journal of Statistical Software , vol. 31, no. 7, pp. 1–24, 2009
2009
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and S. King, “The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,” in 2013 international conference oriental COCOSDA held jointly with 2013 conference on Asian spoken language research and evaluation (O-COCOSDA/CASLRE) . IEEE, 2013, pp. 1–4
2013
Earlier work this paper cites.
L. J. Rodríguez-Fuentes, A. Varona, M. Penagarikano, G. Bordel, and M. Diez, “Gtts-ehu systems for quesst at mediaeval 2014.” in Proceedings of the MediaEval 2014 Workshop , 2014
2014
Earlier work this paper cites.
L. J. Rodriguez-Fuentes, A. Varona, M. Penagarikano, G. Bordel, and M. Diez, “Gtts-ehu systems for quesst at mediaeval 2014,” in MediaEval , 2014
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in ICASSP , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
X. Anguera, L. Rodriguez-Fuentes, A. Buzo, F. Metze, I. Szöke, and M. Penagarikano, “Quesst2014: Evaluating query-by-example speech search in a zero-resource setting with real-life queries,” in ICASSP , 2015, pp. 5833–5837
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 708–712
2015
Earlier work this paper cites.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from non-parallel corpora using variational auto-encoder,” in 2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA) . IEEE, 2016, pp. 1–6
2016
Earlier work this paper cites.
——, “Voice conversion from unaligned corpora using variational autoencoding wasserstein generative adversarial networks,” Interspeech 2017 , 2017
2017
Earlier work this paper cites.
P. Warden, “Speech commands: A public dataset for single-word speech recognition.” Dataset available online , 2017
2017
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 241–245
2017
Earlier work this paper cites.
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 10, pp. 1901–1913, 2017
2017
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 5220–5224
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in EMNLP , 2018, pp. 353–355
2018
Earlier work this paper cites.
M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in NAACL , 2018, pp. 2227–2237
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Sarma, P. Ghahremani, D. Povey, N. K. Goel, K. K. Sarma, and N. Dehak, “Emotion identification from raw speech signals using dnns.” in Interspeech , 2018, pp. 3097–3101
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in ICASSP , 2018, pp. 5329–5333
2018
Earlier work this paper cites.
F. Wang, J. Cheng, W. Liu, and H. Liu, “Additive margin softmax for face verification,” IEEE Signal Processing Letters , vol. 25, no. 7, pp. 926–930, 2018
2018
Earlier work this paper cites.
M. Post, “A call for clarity in reporting BLEU scores,” in Proceedings of the Third Conference on Machine Translation: Research Papers . Belgium, Brussels: Association for Computational Linguistics, Oct. 2018, pp. 186–191
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS Synthesis by Conditioning WaveNet on MEL Spectrogram Predictions,” in Proc. ICASSP , 2018, pp. 4779–4783
2018
Earlier work this paper cites.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 10, pp. 1702–1726, 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An Unsupervised Autoregressive Model for Speech Representation Learning,” in Interspeech , 2019, pp. 146–150
2019
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition.” in Interspeech , 2019
2019
Earlier work this paper cites.
I. Tenney, D. Das, and E. Pavlick, “Bert rediscovers the classical nlp pipeline,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 4593–4601
2019
Earlier work this paper cites.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “Autovc: Zero-shot voice style transfer with only autoencoder loss,” in International Conference on Machine Learning . PMLR, 2019, pp. 5210–5219
2019
Earlier work this paper cites.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in Proceedings of the 2019 Conference of NACCL (Demonstrations) , 2019, pp. 48–53
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” in Interspeech , 2019, pp. 2613–2617
2019
Cited alongside, same era.
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. Le Roux, “Wham!: Extending speech separation to noisy environments,” Proc. Interspeech 2019 , pp. 1368–1372, 2019
2019
Cited alongside, same era.
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, “End-to-end neural speaker diarization with permutation-free objectives,” in Interspeech , 2019, pp. 4300–4304
2019
Cited alongside, same era.
L. Lugosch, M. Ravanelli, P. Ignoto, V. S. Tomar, and Y. Bengio, “Speech model pre-training for end-to-end spoken language understanding,” in Interspeech , 2019, pp. 814–818
2019
Cited alongside, same era.
S. wen Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee, “SUPERB: Speech Processing Universal PERformance Benchmark,” in Proc. Interspeech 2021 , 2021, pp. 1194–1198
2021
Later among the works it cites.
A. Pasad, J.-C. Chou, and K. Livescu, “Layer-wise analysis of a self-supervised speech representation model,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2021, pp. 914–921
2021
Later among the works it cites.
2021
Later among the works it cites.
R. Vygon and N. Mikhaylovskiy, “Learning efficient representations for keyword spotting with triplet loss,” in Speech and Computer: 23rd International Conference, SPECOM 2021, St. Petersburg, Russia, September 27–30, 2021, Proceedings 23 . Springer, 2021, pp. 773–785
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Tomashenko et al. , “Recent advances in end-to-end spoken language understanding,” in International Conference on Statistical Language and Speech Processing , 2019, pp. 44–55
2019
Cited alongside, same era.
S. Kornblith, J. Shlens, and Q. V. Le, “Do better imagenet models transfer better?” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 2661–2671
2019
Cited alongside, same era.
Q. Lin, R. Yin, M. Li, H. Bredin, and C. Barras, “Lstm based similarity measurement with spectral clustering for speaker diarization,” in Annual Conference of the International Speech Communication Association , 2019
2019
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in Proceedings of the 37th ICML , ser. Proceedings of Machine Learning Research, H. D. III and A. Singh, Eds., vol. 119. PMLR, 13–18 Jul 2020, pp. 1597–1607
2020
Cited alongside, same era.
A. T. Liu, S.-w. Yang, P.-H. Chi, P.-c. Hsu, and H.-y. Lee, “Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,” ICASSP , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Baevski, S. Schneider, and M. Auli, “vq-wav2vec: Self-supervised learning of discrete speech representations,” in ICLR , 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in NeurIPS , 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
Y. Qian, X. Bianv, Y. Shi, N. Kanda, L. Shen, Z. Xiao, and M. Zeng, “Speech-language pre-training for end-to-end spoken language understanding,” in ICASSP . IEEE, 2021, pp. 7458–7462
2021
Later among the works it cites.
C.-I. Lai, Y.-S. Chuang, H.-Y. Lee, S.-W. Li, and J. Glass, “Semi-supervised spoken language understanding via self-supervised speech and language model pretraining,” in ICASSP , 2021, pp. 7468–7472
2021
Later among the works it cites.
X. e. a. Li, “Multilingual speech translation from efficient finetuning of pretrained models,” in Proceedings of the 59th ACL , C. Zong, F. Xia, W. Li, and R. Navigli, Eds., Aug. 2021, pp. 827–838
2021
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Later among the works it cites.
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “Data2vec: A general framework for self-supervised learning in speech, vision and language,” in International Conference on Machine Learning . PMLR, 2022, pp. 1298–1312
2022
Later among the works it cites.
S. Shon, A. Pasad, F. Wu, P. Brusco, Y. Artzi, K. Livescu, and K. J. Han, “Slue: New benchmark tasks for spoken language understanding evaluation on natural speech,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7927–7931
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Conneau, A. Bapna, Y. Zhang, M. Ma, P. von Platen, A. Lozhkov, C. Cherry, Y. Jia, C. Rivera, M. Kale, D. van Esch, V. Axelrod, S. Khanuja, J. Clark, O. Firat, M. Auli, S. Ruder, J. Riesa, and M. Johnson, “XTREME-S: Evaluating Cross-lingual Speech Representations,” in Proc. Interspeech 2022 , 2022, pp. 3248–3252
2022
Later among the works it cites.
H.-S. Tsai, H.-J. Chang, W.-C. Huang, Z. Huang, K. Lakhotia, S.-w. Yang, S. Dong, A. Liu, C.-I. Lai, J. Shi et al. , “SUPERB-SG: Enhanced speech processing universal performance benchmark for semantic and generative capabilities,” in Proceedings of the 60th ACL , 2022, pp. 8479–8492
2022
Later among the works it cites.
J. Shor, A. Jansen, W. Han, D. Park, and Y. Zhang, “Universal paralinguistic speech representations using self-supervised conformers,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 3169–3173
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
W.-C. Huang, S.-W. Yang, T. Hayashi, H.-Y. Lee, S. Watanabe, and T. Toda, “S3prl-vc: Open-source voice conversion framework with self-supervised speech representations,” in ICASSP , 2022, pp. 6552–6556
2022
Later among the works it cites.
Y.-Y. Yang, M. Hira, Z. Ni, A. Astafurov, C. Chen, C. Puhrsch, D. Pollack, D. Genzel, D. Greenberg, E. Z. Yang et al. , “Torchaudio: Building blocks for audio and speech processing,” in ICASSP , 2022, pp. 6982–6986
2022
Later among the works it cites.
H.-J. Chang, S.-w. Yang, and H.-y. Lee, “Distilhubert: Speech representation learning by layer-wise distillation of hidden-unit bert,” in ICASSP 2022 . IEEE, 2022, pp. 7087–7091
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Chen, Y. Wu, C. Wang, Z. Chen, Z. Chen, S. Liu, J. Wu, Y. Qian, F. Wei, J. Li et al. , “Unispeech-sat: Universal speech representation learning with speaker aware pre-training,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6152–6156
2022
Later among the works it cites.
Z. Huang, S. Watanabe, S.-w. Yang, P. García, and S. Khudanpur, “Investigating self-supervised learning for speech enhancement and separation,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6837–6841
2022
Later among the works it cites.
K. Qian, Y. Zhang, H. Gao, J. Ni, C.-I. Lai, D. Cox, M. Hasegawa-Johnson, and S. Chang, “Contentvec: An improved self-supervised speech representation by disentangling speakers,” in International Conference on Machine Learning . PMLR, 2022, pp. 18 003–18 017
2022
Later among the works it cites.
M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Nouby et al. , “Dinov2: Learning robust visual features without supervision,” Transactions on Machine Learning Research , 2023
2023
Later among the works it cites.
A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “Fleurs: Few-shot learning evaluation of universal representations of speech,” in 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2023, pp. 798–805
2023
Later among the works it cites.
A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “Fleurs: Few-shot learning evaluation of universal representations of speech,” in 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2023, pp. 798–805
2023
Later among the works it cites.
T.-h. Feng, A. Dong, C.-F. Yeh, S.-w. Yang, T.-Q. Lin, J. Shi, K.-W. Chang, Z. Huang, H. Wu, X. Chang et al. , “SUPERB@SLT 2022: Challenge on generalization and efficiency of self-supervised speech representation learning,” in 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2023, pp. 1096–1103
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International Conference on Machine Learning . PMLR, 2023, pp. 28 492–28 518
2023
Later among the works it cites.
2023
Later among the works it cites.
V. S. Lodagala, S. Ghosh, and S. Umesh, “data2vec-aqc: Search for the right teaching assistant in the teacher-student training setup,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
——, “Ccc-wav2vec 2.0: Clustering aided cross contrastive self-supervised learning of speech representations,” in 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2023, pp. 1–8
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Pasad, B. Shi, and K. Livescu, “Comparative layer-wise analysis of self-supervised speech models,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
T. Parcollet, H. Nguyen, S. Evain, M. Z. Boito, A. Pupier, S. Mdhaffar, H. Le, S. Alisamir, N. Tomashenko, M. Dinarelli et al. , “Lebenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of french speech,” Computer Speech & Language , p. 101622, 2024
2024
Closest in time.
Z. Huang, Y. Shao, S.-X. Zhang, and D. Yu, “Unix-encoder: A universal x-channel speech encoder for ad-hoc microphone array speech processing,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 11 991–11 995
2024
Closest in time.
C.-y. Huang, K.-H. Lu, S.-H. Wang, C.-Y. Hsiao, C.-Y. Kuan, H. Wu, S. Arora, K.-W. Chang, J. Shi, Y. Peng et al. , “Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 12 136–12 140
2024
Closest in time.