Fetching the paper…
Reading the bibliography…
Recent advances in large language models (LLMs) have stepped forward the development of multilingual speech and machine translation by its reduced representation errors and incorporated external knowledge.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 1908
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber. 2019 · 1912
Earlier work this paper cites.
Bitiimt: A bilingual text-infilling method for interactive machine translation
Yanling Xiao, Lemao Liu, et al. 2022 · 1969
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Covost 2 and massively multilingual speech-to-text translation
Changhan Wang, Anne Wu, and Juan Pino. 2020 · 2007
Earlier work this paper cites.
Findings of the 2016 conference on machine translation (wmt16)
Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, et al. 2016 · 2016
Earlier work this paper cites.
Must-c: a multilingual speech translation corpus
Mattia A Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2019 · 2017
Earlier work this paper cites.
chrf++: words helping character n-grams
Maja Popović. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
Findings of the 2019 conference on machine translation
Loıc Barrault, Ondrej Bojar, Marta R Costa-Jussa, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, et al. 2019 · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Earlier work this paper cites.
Paracrawl: Web-scale acquisition of parallel corpora
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, et al. 2020 · 2020
Earlier work this paper cites.
Findings of the 2020 conference on machine translation (wmt20)
Barrault Loïc, Biesialska Magdalena, Bojar Ondřej, Federmann Christian, Graham Yvette, Grundkiewicz Roman, Haddow Barry, Huck Matthias, et al. 2020 · 2020
Earlier work this paper cites.
JParaCrawl: A large scale web-based English-Japanese parallel corpus
Makoto Morishita, Jun Suzuki, and Masaaki Nagata. 2020 · 2020
Earlier work this paper cites.
Xls-r: Self-supervised cross-lingual speech representation learning at scale
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, et al. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Fastcorrect: Fast error correction with edit alignment for automatic speech recognition
Yichong Leng, Xu Tan, Linchen Zhu, Jin Xu, Renqian Luo, Linquan Liu, Tao Qin, Xiangyang Li, Edward Lin, and Tie-Yan Liu. 2021 · 2021
Cited alongside, same era.
The ustc-nelslip systems for simultaneous speech translation task at iwslt 2021
Dan Liu, Mengge Du, Xiaoxi Li, Yuchen Hu, and Lirong Dai. 2021 · 2021
Cited alongside, same era.
Streaming transformer asr with blockwise synchronous beam search
Emiru Tsunoo, Yosuke Kashiwagi, and Shinji Watanabe. 2021 · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2021 · 2021
Cited alongside, same era.
Multi-task language modeling for improving speech recognition of rare words
Chao-Han Huck Yang, Linda Liu, et al. 2021a · 2021
Seamless: Multilingual expressive and streaming speech translation
Loïc Barrault, Yu-An Chung, Mariano Coria Meglioli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, John Hoffman, et al. 2023b · 2023
Later among the works it cites.
Audiolm: a language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al. 2023 · 2023
Later among the works it cites.
Metric-oriented speech enhancement using diffusion probabilistic model
Chen Chen, Yuchen Hu, Weiwei Weng, and Eng Siong Chng. 2023a · 2023
Later among the works it cites.
Improving multilingual and code-switching asr using large language model generated text
Ke Hu, Tara N Sainath, et al. 2023a · 2023
Later among the works it cites.
Comsl: A composite speech-language model for end-to-end speech-to-text translation
Chenyang Le, Yao Qian, Long Zhou, Shujie Liu, Michael Zeng, and Xuedong Huang. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Noise-robust speech recognition with 10 minutes unparalleled in-domain data
Chen Chen, Nana Hou, Yuchen Hu, Shashank Shirol, and Eng Siong Chng. 2022a · 2022
Cited alongside, same era.
Self-critical sequence training for automatic speech recognition
Chen Chen, Yuchen Hu, Nana Hou, Xiaofeng Qi, Heqing Zou, and Eng Siong Chng. 2022b · 2022
Cited alongside, same era.
Fleurs: Few-shot learning evaluation of universal representations of speech
Alexis Conneau, Min Ma, et al. 2023 · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022 · 2022
Cited alongside, same era.
The flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, et al. 2022 · 2022
Cited alongside, same era.
Interactive feature fusion for end-to-end noise-robust speech recognition
Yuchen Hu, Nana Hou, Chen Chen, and Eng Siong Chng. 2022 · 2022
Cited alongside, same era.
Can language models learn from explanations in context?
Andrew K Lampinen, Ishita Dasgupta, Stephanie CY Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L McClelland, Jane X Wang, and Felix Hill. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
New trends in machine translation using large language models: Case examples with chatgpt
Chenyang Lyu, Jitao Xu, and Longyue Wang. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023 · 2023
Later among the works it cites.
Whispering llama: A cross-modal generative error correction framework for speech recognition
Srijith Radhakrishnan, Chao-Han Yang, Sumeer Khan, et al. 2023 · 2023
Later among the works it cites.
Audiopalm: A large language model that can speak and listen
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, et al. 2023 · 2023
Later among the works it cites.
Can whisper perform speech-based in-context learning
Siyin Wang, Chao-Han Huck Yang, Ji Wu, and Chao Zhang. 2023 · 2023
Later among the works it cites.
Generative speech recognition error correction with large language models and task-activating prompting
Chao-Han Huck Yang, Yile Gu, Yi-Chieh Liu, Shalini Ghosh, Ivan Bulyko, and Andreas Stolcke. 2023a · 2023
Later among the works it cites.
Low-rank adaptation of large language model rescoring for parameter-efficient speech recognition
Yu Yu, Chao-Han Huck Yang, Jari Kolehmainen, et al. 2023 · 2023
Later among the works it cites.
Robust data2vec: Noise-robust speech representation learning for asr by combining regression and improved contrastive learning
Qiu-Shi Zhu, Long Zhou, Jie Zhang, Shu-Jie Liu, Yu-Chen Hu, and Li-Rong Dai. 2023 · 2023
Later among the works it cites.
Chen Chen, Ruizhe Li, Yuchen Hu, Sabato Marco Siniscalchi, Pin-Yu Chen, Ensiong Chng, and Chao-Han Huck Yang. 2024 · 2024
Closest in time.
Large language models are efficient learners of noise-robust speech recognition
Yuchen Hu, Chen Chen, Chao-Han Huck Yang, Ruizhe Li, Chao Zhang, Pin-Yu Chen, and EnSiong Chng. 2024 · 2024
Closest in time.
Multichannel av-wav2vec2: A framework for learning multichannel multi-modal speech representation
Qiushi Zhu, Jie Zhang, Yu Gu, Yuchen Hu, and Lirong Dai. 2024 · 2024
Closest in time.