Fetching the paper…
Reading the bibliography…
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
M. Witbrock and A. G. Hauptmann, “Speech recognition and information retrieval: Experiments in retrieving spoken documents,” in Proceedings of the DARPA speech recognition workshop , vol. 97, 1997
1997
Earlier work this paper cites.
S.-y. Ishikawa, T. Ikeda, K. Miki, F. Adachi, R. Isotani, K.-I. Iso, and A. Okumura, “Speech-activated text retrieval system for multimodal cellular phones,” in 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 1. IEEE, 2004, pp. I–453
2004
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
M. J. F. Gales, K. M. Knill, A. Ragni, and S. P. Rath, “Speech recognition and keyword spotting for low-resource languages: Babel project research at cued,” in n Spoken Language Technologies for Under-Resourced Languages , 2014
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in Proc. of ICASSP . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition.” in INTERSPEECH . ISCA, 2015, pp. 3586–3589. [Online]. Available: http://dblp.uni-trier.de/db/conf/interspeech/interspeech2015.html#KoPPK15
2015
Earlier work this paper cites.
L.-s. Lee, J. Glass, H.-y. Lee, and C.-a. Chan, “Spoken content retrieval—beyond cascading speech recognition with text retrieval,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 9, pp. 1389–1420, 2015
2015
Earlier work this paper cites.
R. K. Mohapatra, T. K. Mishra, S. Panda, and B. Majhi, “Ohcs: A database for handwritten atomic odia character recognition,” in 2015 Fifth National Conference on Computer Vision, Pattern Recognition, Image Processing and Graphics (NCVPRIPG) . IEEE, 2015, pp. 1–4
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. of NIPS , 2017
2017
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 5220–5224
2017
Earlier work this paper cites.
A. W. Black, “CMU wilderness multilingual speech dataset,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5971–5975
2019
Earlier work this paper cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” Proc. Interspeech 2019 , pp. 2613–2617, 2019
2019
Earlier work this paper cites.
Y. Yang, G. H. Ábrego, S. Yuan, M. Guo, Q. Shen, D. Cer, Y.-H. Sung, B. Strope, and R. Kurzweil, “Improving multilingual sentence embedding using bi-directional dual encoder with additive margin softmax,” in IJCAI , 2019
2019
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” Proc. of Interspeech , 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. of NeurIPS , 2020
2020
Cited alongside, same era.
Q. Xu, A. Baevski, T. Likhomanenko, P. Tomasello, A. Conneau, R. Collobert, G. Synnaeve, and M. Auli, “Self-training and pre-training are complementary for speech recognition,” in Proc. of ICASSP , 2020
2020
Cited alongside, same era.
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, “VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in Proc. of ACL , 2021
2021
Later among the works it cites.
R. Cattoni, M. A. Di Gangi, L. Bentivogli, M. Negri, and M. Turchi, “Must-c: A multilingual corpus for end-to-end speech translation,” Computer Speech & Language , vol. 66, p. 101155, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A large-scale multilingual dataset for speech research,” in Proc. of Interspeech , 2020
2020
Cited alongside, same era.
C. Wang, A. Wu, and J. Pino, “CoVoST 2 and massively multilingual speech-to-text translation,” arXiv , 2020
2020
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common Voice: A massively-multilingual speech corpus,” Proc. of LREC , 2020
2020
Cited alongside, same era.
J. Valk and T. Alumäe, “VoxLingua107: a dataset for spoken language recognition,” in Proc. of SLT , 2020
2020
Cited alongside, same era.
J. Iranzo-Sánchez, J. A. Silvestre-Cerda, J. Jorge, N. Roselló, A. Giménez, A. Sanchis, J. Civera, and A. Juan, “Europarl-ST: A multilingual corpus for speech translation of parliamentary debates,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 8229–8233
2020
Cited alongside, same era.
2020
Cited alongside, same era.
X. Li, S. Dalmia, J. Li, M. Lee, P. Littell, J. Yao, A. Anastasopoulos, D. R. Mortensen, G. Neubig, A. W. Black et al. , “Universal phone recognition with a multilingual allophone system,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 8249–8253
2020
Cited alongside, same era.
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, “Unsupervised cross-lingual representation learning for speech recognition,” in Proc. of Interspeech , 2021
2021
Cited alongside, same era.
Later among the works it cites.
A. Fan, S. Bhosale, H. Schwenk, Z. Ma, A. El-Kishky, S. Goyal, M. Baines, O. Celebi, G. Wenzek, V. Chaudhary et al. , “Beyond english-centric multilingual machine translation,” Journal of Machine Learning Research , vol. 22, no. 107, pp. 1–48, 2021
2021
Later among the works it cites.
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y. Wu, “W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,” 2021
2021
Later among the works it cites.
Y. Zhang, D. S. Park, W. Han, J. Qin, A. Gulati, J. Shor, A. Jansen, Y. Xu, Y. Huang, S. Wang, Z. Zhou, B. Li, M. Ma, W. Chan, J. Yu, Y. Wang, L. Cao, K. C. Sim, B. Ramabhadran, T. N. Sainath, F. Beaufays, Z. Chen, Q. V. Le, C.-C. Chiu, R. Pang, and Y. Wu, “BigSSL: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition,” arXiv , 2021
2021
Later among the works it cites.
A. Jain, M. Guo, K. Srinivasan, T. Chen, S. Kudugunta, C. Jia, Y. Yang, and J. Baldridge, “MURAL: Multimodal, multitask representations across languages,” in Findings of the Association for Computational Linguistics: EMNLP 2021 . Punta Cana, Dominican Republic: Association for Computational Linguistics, Nov. 2021, pp. 3449–3463. [Online]. Available: https://aclanthology.org/2021.findings-emnlp.293
2021
Later among the works it cites.
P.-A. Duquenne, H. Gong, and H. Schwenk, “Multimodal and multilingual embeddings for large-scale speech mining,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
2022
Closest in time.
2022
Closest in time.
F. Feng, Y. Yang, D. Cer, N. Arivazhagan, and W. Wang, “Language-agnostic BERT sentence embedding,” in Proceedings of ACL . Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 878–891. [Online]. Available: https://aclanthology.org/2022.acl-long.62
2022
Closest in time.
D. van Esch, T. Lucassen, S. Ruder, I. Caswell, and C. E. Rivera, “Writing system and speaker metadata for 2,800+ language varieties,” in Proceedings of LREC , 2022
2022
Closest in time.