Fetching the paper…
Reading the bibliography…
Recent large language models (LLM) exhibit sub-optimal performance on low-resource languages, as the training data of these models is usually dominated by English and other high-resource languages.
P. Gage, “A new algorithm for data compression,” The C Users Journal Archive , vol. 12, pp. 23–38, 1994. [Online]. Available: https://api.semanticscholar.org/CorpusID:59804030
1994
Earlier work this paper cites.
A. Broder, “On the resemblance and containment of documents,” in Proceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No.97TB100171) , 1997, pp. 21–29
1997
Earlier work this paper cites.
R. M. French, “Catastrophic forgetting in connectionist networks,” Trends in cognitive sciences , vol. 3, no. 4, pp. 128–135, 1999
1999
Earlier work this paper cites.
V. Vincze, D. Szauter, A. Almási, G. Móra, Z. Alexin, and J. Csirik, “Hungarian dependency treebank,” in Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10) . Valletta, Malta: European Language Resources Association (ELRA), May 2010. [Online]. Available: http://www.lrec-conf.org/proceedings/lrec2010/pdf/465_Paper.pdf
2010
Earlier work this paper cites.
M. Cettolo, C. Girardi, and M. Federico, “WIT3: Web inventory of transcribed and translated talks,” in Proceedings of the 16th Annual conference of the European Association for Machine Translation . Trento, Italy: European Association for Machine Translation, May 28–30 "2012", pp. 261–268. [Online]. Available: https://www.aclweb.org/anthology/2012.eamt-1.60
2012
Earlier work this paper cites.
N. Silveira, T. Dozat, M.-C. de Marneffe, S. Bowman, M. Connor, J. Bauer, and C. D. Manning, “A gold standard dependency corpus for English,” in Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC-2014) , 2014
2014
Earlier work this paper cites.
J. Nivre, M.-C. de Marneffe, F. Ginter, Y. Goldberg, J. Hajič, C. D. Manning, R. McDonald, S. Petrov, S. Pyysalo, N. Silveira, R. Tsarfaty, and D. Zeman, “Universal Dependencies v1: A multilingual treebank collection,” in Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16) . Portorož, Slovenia: European Language Resources Association (ELRA), May 2016, pp. 1659–1666. [Online]. Available: https://aclanthology.org/L16-1262
2016
Earlier work this paper cites.
H. Riza, M. Purwoadi, Gunarso, T. Uliniansyah, A. A. Ti, S. M. Aljunied, L. C. Mai, V. T. Thang, N. P. Thai, V. Chea, R. Sun, S. Sam, S. Seng, K. M. Soe, K. T. Nwet, M. Utiyama, and C. Ding, “Introduction of the asian language treebank,” 2016 Conference of The Oriental Chapter of International Committee for Coordination and Standardization of Speech Databases and Assessment Techniques (O-COCOSDA) , pp. 1–6, 2016. [Online]. Available: https://api.semanticscholar.org/CorpusID:45848332
2016
Earlier work this paper cites.
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández, “The lambada dataset: Word prediction requiring a broad discourse context,” 2016
2016
Earlier work this paper cites.
D. Zeman, M. Popel, M. Straka, J. Hajič, J. Nivre, F. Ginter, J. Luotolahti, S. Pyysalo, S. Petrov, M. Potthast, F. Tyers, E. Badmaeva, M. Gokirmak, A. Nedoluzhko, S. Cinková, J. Hajič jr., J. Hlaváčová, V. Kettnerová, Z. Urešová, J. Kanerva, S. Ojala, A. Missilä, C. D. Manning, S. Schuster, S. Reddy, D. Taji, N. Habash, H. Leung, M.-C. de Marneffe, M. Sanguinetti, M. Simi, H. Kanayama, V. de Paiva, K. Droganova, H. Martínez Alonso, Ç. Çöltekin, U. Sulubacak, H. Uszkoreit, V. Macketanz, A. Burchardt, K. Harris, K. Marheinecke, G. Rehm, T. Kayadelen, M. Attia, A. Elkahky, Z. Yu, E. Pitler, S. Lertpradit, M. Mandl, J. Kirchner, H. F. Alcalde, J. Strnadová, E. Banerjee, R. Manurung, A. Stella, A. Shimada, S. Kwak, G. Mendonça, T. Lando, R. Nitisaroj, and J. Li, “CoNLL 2017 shared task: Multilingual parsing from raw text to Universal Dependencies,” in Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies . Vancouver, Canada: Association for Computational Linguistics, Aug. 2017, pp. 1–19. [Online]. Available: https://aclanthology.org/K17-3001
2017
Earlier work this paper cites.
A. Conneau, G. Lample, R. Rinott, A. Williams, S. R. Bowman, H. Schwenk, and V. Stoyanov, “Xnli: Evaluating cross-lingual sentence representations,” in Conference on Empirical Methods in Natural Language Processing , 2018. [Online]. Available: https://api.semanticscholar.org/CorpusID:52271711
2018
Earlier work this paper cites.
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal, “Can a suit of armor conduct electricity? a new dataset for open book question answering,” 2018
2018
Earlier work this paper cites.
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord, “Think you have solved question answering? try arc, the ai2 reasoning challenge,” 2018
2018
Earlier work this paper cites.
D. Khashabi, S. Chaturvedi, M. Roth, S. Upadhyay, and D. Roth, “Looking beyond the surface:a challenge set for reading comprehension over multiple sentences,” in Proceedings of North American Chapter of the Association for Computational Linguistics (NAACL) , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Narayan, S. B. Cohen, and M. Lapata, “Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization,” 2018
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
J. Ács. (2019, February) Exploring bert’s vocabulary. [Online]. Available: https://juditacs.github.io/2019/02/19/bert-tokenization-stats.html
2019
Earlier work this paper cites.
G. Wenzek, M.-A. Lachaux, A. Conneau, V. Chaudhary, F. Guzmán, A. Joulin, and E. Grave, “Ccnet: Extracting high quality monolingual datasets from web crawl data,” 2019
2019
Earlier work this paper cites.
H. Nomoto, “Interpersonal meaning annotation for asian language corpora: The case of tufs asian language parallel corpus (talpco),” 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID:209441604
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Suriyawongkul, E. Chuangsuwanich, P. Chormai, and C. Polpanumas, “Pythainlp/wisesight-sentiment: First release,” Sep. 2019. [Online]. Available: https://doi.org/10.5281/zenodo.3457447
2019
Earlier work this paper cites.
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “Hellaswag: Can a machine really finish your sentence?” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019
2019
Earlier work this paper cites.
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova, “Boolq: Exploring the surprising difficulty of natural yes/no questions,” 2019
2019
Earlier work this paper cites.
Y. Bisk, R. Zellers, R. L. Bras, J. Gao, and Y. Choi, “Piqa: Reasoning about physical commonsense in natural language,” 2019
2019
Earlier work this paper cites.
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi, “Winogrande: An adversarial winograd schema challenge at scale,” 2019
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, M. Kelcey, J. Devlin, K. Lee, K. N. Toutanova, L. Jones, M.-W. Chang, A. Dai, J. Uszkoreit, Q. Le, and S. Petrov, “Natural questions: a benchmark for question answering research,” Transactions of the Association of Computational Linguistics , 2019
2019
Cited alongside, same era.
A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov, “Unsupervised cross-lingual representation learning at scale,” 2020
2020
Cited alongside, same era.
J. Phang, I. Calixto, P. M. Htut, Y. Pruksachatkun, H. Liu, C. Vania, K. Kann, and S. R. Bowman, “English intermediate-task training improves zero-shot cross-lingual transfer too,” 2020
O. Shliazhko, A. Fenogenova, M. Tikhonova, V. Mikhailov, A. Kozlova, and T. Shavrina, “mgpt: Few-shot learners go multilingual,” 2022
2022
Later among the works it cites.
X. V. Lin, T. Mihaylov, M. Artetxe, T. Wang, S. Chen, D. Simig, M. Ott, N. Goyal, S. Bhosale, J. Du, R. Pasunuru, S. Shleifer, P. S. Koura, V. Chaudhary, B. O’Horo, J. Wang, L. Zettlemoyer, Z. Kozareva, M. Diab, V. Stoyanov, and X. Li, “Few-shot learning with multilingual language models,” 2022
2022
Later among the works it cites.
T. Vu, A. Barua, B. Lester, D. Cer, M. Iyyer, and N. Constant, “Overcoming catastrophic forgetting in zero-shot cross-lingual generation,” 2022
2022
Later among the works it cites.
H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal, and C. A. Raffel, “Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 1950–1965, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” 2020
2020
Cited alongside, same era.
J. Pfeiffer, I. Vulić, I. Gurevych, and S. Ruder, “MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Online: Association for Computational Linguistics, Nov. "2020", pp. 7654–7673. [Online]. Available: https://aclanthology.org/2020.emnlp-main.617
2020
Cited alongside, same era.
D. M. Nemeskey, “Natural language processing methods for language modeling,” Ph.D. dissertation, Eötvös Loránd University, 2020. [Online]. Available: https://hlt.bme.hu/media/pdf/nemeskey_thesis.pdf
2020
Cited alongside, same era.
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy, “The pile: An 800gb dataset of diverse text for language modeling,” 2020
2020
Cited alongside, same era.
B. Buschbeck-Wolf and M. Exel, “A parallel evaluation data set of software documentation with document structure annotation,” in Workshop on Asian Translation , 2020. [Online]. Available: https://api.semanticscholar.org/CorpusID:221095497
2020
Cited alongside, same era.
2020
Cited alongside, same era.
N. Chumpolsathien, “Using knowledge distillation from keyword extraction to improve the informativeness of neural cross-lingual summarization,” Master’s thesis, Beijing Institute of Technology, 2020
2020
Cited alongside, same era.
Y. Nie, A. Williams, E. Dinan, M. Bansal, J. Weston, and D. Kiela, “Adversarial nli: A new benchmark for natural language understanding,” 2020
2020
Cited alongside, same era.
J. Abadji, P. O. Suarez, L. Romary, and B. Sagot, “Towards a cleaner document-oriented multilingual crawled corpus,” 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
N. Ligeti-Nagy, G. Ferenczi, E. Héja, K. Jelencsik-Mátyus, L. J. Laki, N. Vadász, Z. G. Yang, and T. Vadász, “Hulu: magyar nyelvű benchmark adatbázis kiépítése a neurális nyelvmodellek kiértékelése céljából,” in XVIII. Magyar Számítógépes Nyelvészeti Konferencia , 2022, pp. 431–446
2022
Later among the works it cites.
B. Workshop, “Bloom: A 176b-parameter open-access multilingual language model,” 2023
2023
Closest in time.
N. Muennighoff, T. Wang, L. Sutawika, A. Roberts, S. Biderman, T. L. Scao, M. S. Bari, S. Shen, Z.-X. Yong, H. Schoelkopf, X. Tang, D. Radev, A. F. Aji, K. Almubarak, S. Albanie, Z. Alyafeai, A. Webson, E. Raff, and C. Raffel, “Crosslingual generalization through multitask finetuning,” 2023
2023
Closest in time.
Z.-X. Yong, H. Schoelkopf, N. Muennighoff, A. F. Aji, D. I. Adelani, K. Almubarak, M. S. Bari, L. Sutawika, J. Kasai, A. Baruwa, G. I. Winata, S. Biderman, E. Raff, D. Radev, and V. Nikoulina, “Bloom+1: Adding language support to bloom for zero-shot prompting,” 2023
2023
Closest in time.
J. Ye, X. Tao, and L. Kong, “Language versatilists vs. specialists: An empirical revisiting on multilingual transfer ability,” 2023
2023
Closest in time.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023
2023
Closest in time.
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom, “Llama 2: Open foundation and fine-tuned chat models,” 2023
2023
Closest in time.
F. Stollenwerk, “Training and evaluation of a multilingual tokenizer for gpt-sw3,” 2023
2023
Closest in time.
S. Cahyawijaya, H. Lovenia, T. Yu, W. Chung, and P. Fung, “Instruct-align: Teaching novel languages with to llms through alignment-based cross-lingual instruction,” 2023
2023
Closest in time.
R. Pires, H. Abonizio, T. S. Almeida, and R. Nogueira, “Sabiá: Portuguese large language models,” 2023
2023
Closest in time.
2023
Closest in time.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” 2023
2023
Closest in time.
2023
Closest in time.
H. Nguyen, “The oig dataset,” Mar 2023. [Online]. Available: https://laion.ai/blog/oig-dataset/
2023
Closest in time.
S. Iyer, X. V. Lin, R. Pasunuru, T. Mihaylov, D. Simig, P. Yu, K. Shuster, T. Wang, Q. Liu, P. S. Koura, X. Li, B. O’Horo, G. Pereyra, J. Wang, C. Dewan, A. Celikyilmaz, L. Zettlemoyer, and V. Stoyanov, “Opt-iml: Scaling language model instruction meta learning through the lens of generalization,” 2023
2023
Closest in time.
C. Mou, C. Ha, K. Enevoldsen, and P. Liu, “Chenghaomou/text-dedup: Reference snapshot,” Sep. 2023. [Online]. Available: https://doi.org/10.5281/zenodo.8364980
2023
Closest in time.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction-following llama model,” https://github.com/tatsu-lab/stanford_alpaca , 2023
2023
Closest in time.
M. Conover, M. Hayes, A. Mathur, J. Xie, J. Wan, S. Shah, A. Ghodsi, P. Wendell, M. Zaharia, and R. Xin. (2023) Free dolly: Introducing the world’s first truly open instruction-tuned llm. [Online]. Available: https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.