Fetching the paper…
Reading the bibliography…
We present Maestro, a self-supervised training method to unify representations learnt from speech and text modalities.
R. Kuhn and R. De Mori, “A cache-based natural language model for speech recognition,” IEEE transactions on pattern analysis and machine intelligence , vol. 12, no. 6, pp. 570–583, 1990
1990
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Esteve, “TED-LIUM: an automatic speech recognition dedicated corpus.” in Proc. LREC , 2012, pp. 125–129
2012
Earlier work this paper cites.
M. J. Gales, K. M. Knill, A. Ragni, and S. P. Rath, “Speech recognition and keyword spotting for low-resource languages: Babel project research at cued,” in Fourth International workshop on spoken language technologies for under-resourced languages (SLTU-2014) , 2014, pp. 16–23
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in Proc. ICASSP , 2015
2015
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” Proc. Interspeech 2019 , pp. 2613–2617, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira, “Perceiver: General perception with iterative attention,” in International Conference on Machine Learning . PMLR, 2021, pp. 4651–4664
2021
Later among the works it cites.
S. Ueno, M. Mimura, S. Sakai, and T. Kawahara, “Data augmentation for asr using TTS via a discrete representation,” ASRU , 2021
2021
Later among the works it cites.
I. Elias, H. Zen, J. Shen, Y. Zhang, Y. Jia, R. J. Weiss, and Y. Wu, “Parallel Tacotron: Non-autoregressive and controllable TTS,” in Proc. ICASSP , 2021, pp. 5709–5713
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Wang, A. Rosenberg, Z. Chen, Y. Zhang, B. Ramabhadran, Y. Wu, and P. Moreno, “Improving speech recognition using consistent predictions on synthesized speech,” in ICASSP , 2020
2020
Cited alongside, same era.
J. Kahn, M. Rivière, W. Zheng et al. , “Libri-Light: A benchmark for asr with limited or no supervision,” in Proc. ICASSP , 2020, pp. 7669–7673
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Cited alongside, same era.
W.-N. Hsu, Y.-H. H. Tsai, B. Bolte, R. Salakhutdinov, and A. Mohamed, “HuBERT: How much can a bad teacher benefit ASR pre-training?” in Proc. ICASSP , 2021, pp. 6533–6537
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Valk and T. Alumäe, “VoxLingua107: A dataset for spoken language recognition,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 652–658
2021
Later among the works it cites.
Y. Tang, C. Tran, X. Li, P.-J. Chen, N. Goyal, V. Chaudhary, J. Gu, and A. Fan, “Multilingual translation from denoising pre-training,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , 2021, pp. 3450–3466
2021
Later among the works it cites.
2022
Closest in time.
2022
Closest in time.
Z. Chen, Y. Zhang, A. Rosenberg, B. Ramabhadran, P. Moreno, and G. Wang, “Tts4pretrain 2.0: Advancing the use of text and speech in ASR pretraining with consistency and contrastive losses,” in Proc. ICASSP , 2022
2022
Closest in time.