Fetching the paper…
Reading the bibliography…
Recent advancements in large language models (LLMs) have remarkably enhanced performances on a variety of tasks in multiple languages.
Korquad1. 0: Korean qa dataset for machine reading comprehension
Seungyoung Lim, Myungji Kim, and Jooyoul Lee. 2019 · 1909
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Glu variants improve transformer
Noam Shazeer. 2020 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Orthogonal language and task adapters in zero-shot cross-lingual transfer
Marko Vidoni, Ivan Vulić, and Goran Glavaš. 2020 · 2012
Earlier work this paper cites.
Montreal neural machine translation systems for WMT’15
Sébastien Jean, Orhan Firat, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Earlier work this paper cites.
Correcting length bias in neural machine translation
Kenton Murray and David Chiang. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Earlier work this paper cites.
Improving pre-trained multilingual model with vocabulary expansion
Hai Wang, Dian Yu, Kai Sun, Jianshu Chen, and Dong Yu. 2019 · 2019
Earlier work this paper cites.
Load what you need: Smaller versions of multilingual bert
Amine Abdaoui, Camille Pradel, and Grégoire Sigel. 2020 · 2020
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Earlier work this paper cites.
Parsing with multilingual BERT, a small corpus, and a small treebank
Ethan C. Chau, Lucy H. Lin, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Mad-x: An adapter-based framework for multi-task cross-lingual transfer
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020 · 2020
Cited alongside, same era.
UDapter: Language adaptation for truly Universal Dependency parsing
Ahmet Üstün, Arianna Bisazza, Gosse Bouma, and Gertjan van Noord. 2020 · 2020
Cited alongside, same era.
MAD-G: Multilingual adapter generation for efficient cross-lingual transfer
Alan Ansell, Edoardo Maria Ponti, Jonas Pfeiffer, Sebastian Ruder, Goran Glavaš, Ivan Vulić, and Anna Korhonen. 2021 · 2021
Cited alongside, same era.
Efficient inference for multilingual neural machine translation
Alexandre Berard, Dain Lee, Stephane Clinchant, Kweonwoo Jung, and Vassilina Nikoulina. 2021 · 2021
Cited alongside, same era.
Xl-sum: Large-scale multilingual abstractive summarization for 44 languages
Tahmid Hasan, Abhik Bhattacharjee, Md Saiful Islam, Kazi Samin, Yuan-Fang Li, Yong-Bin Kang, M Sohel Rahman, and Rifat Shahriyar. 2021 · 2021
Cited alongside, same era.
Model card and evaluations for claude models
Antropic. 2023 · 2023
Later among the works it cites.
The belebele benchmark: a parallel reading comprehension dataset in 122 language variants
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2023 · 2023
Later among the works it cites.
Large language model inference with lexical shortlisting
Nikolay Bogoychev, Pinzhen Chen, Barry Haddow, and Alexandra Birch. 2023 · 2023
Later among the works it cites.
Efficient and effective text encoding for chinese llama and alpaca
Yiming Cui, Ziqing Yang, and Xin Yao. 2023 · 2023
Later among the works it cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Avocado: Strategy for adapting vocabulary to downstream domain
Jimin Hong, Taehee Kim, Hyesu Lim, and Jaegul Choo. 2021 · 2021
Cited alongside, same era.
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Cited alongside, same era.
The devil is in the details: On the pitfalls of vocabulary selection in neural machine translation
Tobias Domhan, Eva Hasler, Ke Tran, Sony Trenous, Bill Byrne, and Felix Hieber. 2022 · 2022
Cited alongside, same era.
The flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022 · 2022
Cited alongside, same era.
JGLUE: Japanese general language understanding evaluation
Kentaro Kurihara, Daisuke Kawahara, and Tomohide Shibata. 2022 · 2022
Cited alongside, same era.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022 · 2022
Cited alongside, same era.
Tri Dao. 2023 · 2023
Later among the works it cites.
FOCUS: Effective embedding initialization for monolingual specialization of multilingual models
Konstantin Dobler and Gerard de Melo. 2023 · 2023
Later among the works it cites.
Gpts are gpts: An early look at the labor market impact potential of large language models
Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock. 2023 · 2023
Later among the works it cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Later among the works it cites.
Copy is all you need
Tian Lan, Deng Cai, Yan Wang, Heyan Huang, and Xian-Ling Mao. 2023 · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023 · 2023
Later among the works it cites.
Chipnemo: Domain-adapted llms for chip design
Mingjie Liu, Teo Ene, Robert Kirby, Chris Cheng, Nathaniel Pinckney, Rongjian Liang, Jonah Alben, Himyanshu Anand, Sanmitra Banerjee, Ismet Bayraktaroglu, et al. 2023 · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI. 2023 · 2023
Later among the works it cites.
Language model tokenizers introduce unfairness between languages
Aleksandar Petrov, Emanuele La Malfa, Philip HS Torr, and Adel Bibi. 2023 · 2023
Later among the works it cites.
High-throughput generative inference of large language models with a single gpu
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E Gonzalez, et al. 2023 · 2023
Later among the works it cites.
Evaluating the social impact of generative ai systems in systems and society
Irene Solaiman, Zeerak Talat, William Agnew, Lama Ahmad, Dylan Baker, Su Lin Blodgett, Hal Daumé III, Jesse Dodge, Ellie Evans, Sara Hooker, et al. 2023 · 2023
Later among the works it cites.
An efficient multilingual language model compression through vocabulary trimming
Asahi Ushio, Yi Zhou, and Jose Camacho-Collados. 2023 · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Later among the works it cites.