Fetching the paper…
Reading the bibliography…
A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach.
2005
Earlier work this paper cites.
D. P. Kingma, “Adam: A method for stochastic optimization,” Proceedings of ICLR , 2015
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2016, pp. 4960–4964
2016
Earlier work this paper cites.
M. Post, “A call for clarity in reporting bleu scores,” arXiv preprint arXiv:1804.08771 , 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
B. Zhang and R. Sennrich, “Root mean square layer normalization,” Proceedings of NeurIPS , 2019
2019
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research , 2020
2020
Earlier work this paper cites.
Y. Liu, J. Gu, N. Goyal, X. Li, S. Edunov, M. Ghazvininejad, M. Lewis, and L. Zettlemoyer, “Multilingual denoising pre-training for neural machine translation,” in Transactions of the Association for Computational Linguistics . MIT Press, 2020, pp. 726–742
2020
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . ACL, 2020, pp. 7871–7880
2020
Earlier work this paper cites.
G. Ramírez-Sánchez, J. Zaragoza-Bernabeu, M. Bañón, and S. Ortiz-Rojas, “Bifixer and bicleaner: two open-source tools to clean your parallel data.” in Proceedings of the 22nd Annual Conference of the European Association for Machine Translation . European Association for Machine Translation, 2020, pp. 291–298
2020
Earlier work this paper cites.
R. Zheng, J. Chen, M. Ma, and L. Huang, “Fused acoustic and text encoding for multimodal bilingual pretraining and speech translation,” in International Conference on Machine Learning . PMLR, 2021, pp. 12 736–12 746
2021
Earlier work this paper cites.
P. Żelasko, D. Povey, J. Trmal, and S. Khudanpur, “Lhotse: a speech data representation library for the modern deep learning ecosystem,” Proceedings of NeurIPS Workshop on Data-Centric AI , 2021
2021
Cited alongside, same era.
J.-B. Alayrac et al. , “Flamingo: a visual language model for few-shot learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 23 716–23 736, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “Fleurs: Few-shot learning evaluation of universal representations of speech,” in 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2023, pp. 798–805
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Goyal, C. Gao, V. Chaudhary, P.-J. Chen, G. Wenzek, D. Ju, S. Krishnan, M. Ranzato, F. Guzmán, and A. Fan, “The flores-101 evaluation benchmark for low-resource and multilingual machine translation,” Transactions of the Association for Computational Linguistics , vol. 10, pp. 522–538, 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
J. Bai et al. , “Qwen technical report,” arXiv preprint arXiv:2309.16609 , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
M. Post, T. Gowda, R. Grundkiewicz, H. Khayrallah, R. Jain, and M. Junczys-Dowmunt, “SOTASTREAM: A streaming approach to machine translation training,” in Proceedings of the 3rd Workshop for Natural Language Processing Open Source Software (NLP-OSS 2023) , L. Tan, D. Milajevs, G. Chauhan, J. Gwinnup, and E. Rippeth, Eds. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 110–119. [Online]. Available: https://aclanthology.org/2023.nlposs-1.13
2023
Cited alongside, same era.
“NVIDIA Megatron NMT en-to-any 500M.” [Online]. Available: https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/models/megatronnmt_en_any_500m
Cited in the paper.
“NVIDIA Megatron NMT any-to-en 500M.” [Online]. Available: https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/models/megatronnmt_any_en_500m
Cited in the paper.
2023
Later among the works it cites.
Z. Chen, H. Huang, A. Andrusenko, O. Hrinchuk, K. C. Puvvada, J. Li, S. Ghosh, J. Balam, and B. Ginsburg, “Salm: Speech-augmented language model with in-context learning for speech recognition and translation,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 13 521–13 525
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
K. C. Puvvada, P. Żelasko, H. Huang, O. Hrinchuk, N. R. Koluguri, K. Dhawan, S. Majumdar, E. Rastorgueva, Z. Chen, V. Lavrukhin, J. Balam, and B. Ginsburg, “Less is more: Accurate speech recognition & translation without web-scale data,” in Proceedings of Interspeech , 2024
2024
Closest in time.