Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are typically multilingual due to pretraining on diverse multilingual corpora.
Language control and lexical competition in bilinguals: an event-related fmri study
Abutalebi, J., Annoni, J.-M., Zimine, I., Pegna, A. J., Seghier, M. L., Lee-Jahnke, H., Lazeyras, F., Cappa, S. F., and Khateb, A · 2008
Earlier work this paper cites.
Bilingualism: consequences for mind and brain
Bialystok, E., Craik, F. I., and Luk, G · 2012
Earlier work this paper cites.
The cognitive benefits of being bilingual
Marian, V. and Shook, A · 2012
Earlier work this paper cites.
Improvements to bm25 and language models examined
Trotman, A., Puurula, A., and Burgess, B · 2014
Earlier work this paper cites.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
Think you have solved question answering? try ARC, the AI2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2018
Earlier work this paper cites.
Simple, scalable adaptation for neural machine translation
Bapna, A. and Firat, O · 2019
Earlier work this paper cites.
Introduction to Natural Language Processing
Eisenstein, J · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
Lample, G. and Conneau, A · 2019
Earlier work this paper cites.
How multilingual is multilingual BERT?
Pires, T., Schlinger, E., and Garrette, D · 2019
Earlier work this paper cites.
Code-switching for enhancing nmt with pre-specified translation
Song, K., Zhang, Y., Yu, H., Luo, W., Wang, K., and Zhang, M · 2019
Earlier work this paper cites.
Beto, Bentz, Becas: The surprising cross-lingual effectiveness of BERT
Wu, S. and Dredze, M · 2019
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Artetxe, M., Ruder, S., and Yogatama, D · 2020
Earlier work this paper cites.
Compositionality decomposed: How do neural networks generalise?
Hupkes, D., Dankers, V., Mul, M., and Bruni, E · 2020
Earlier work this paper cites.
X-factr: Multilingual factual knowledge retrieval from pretrained language models
Jiang, Z., Anastasopoulos, A., Araki, J., Ding, H., and Neubig, G · 2020
Earlier work this paper cites.
COGS: a compositional generalization challenge based on semantic interpretation
Kim, N. and Linzen, T · 2020
Earlier work this paper cites.
COMET: A neural framework for MT evaluation
Rei, R., Stewart, C., Farinha, A. C., and Lavie, A · 2020
Earlier work this paper cites.
Csp: code-switching pre-training for neural machine translation
Yang, Z., Hu, B., Han, A., Huang, S., and Ju, Q · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Multilingual lama: Investigating knowledge in multilingual pretrained language models
Kassner, N., Dufter, P., and Schütze, H · 2021
Earlier work this paper cites.
Zmbart: An unsupervised cross-lingual transfer framework for language generation
Maurya, K. K., Desarkar, M. S., Kano, Y., and Deepshikha, K · 2021
Cited alongside, same era.
Ccmatrix: Mining billions of high-quality parallel sentences on the web
Schwenk, H., Wenzek, G., Edunov, S., Grave, É., Joulin, A., and Fan, A · 2021
Cited alongside, same era.
Cross-lingual language model pretraining for retrieval
Yu, P., Fei, H., and Li, P · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zhao, Z., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Cited alongside, same era.
Towards a cleaner document-oriented multilingual crawled corpus
Abadji, J., Suarez, P. O., Romary, L., and Sagot, B · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
Extrapolating large language models to non-English by aligning languages
Zhu, W., Lv, Y., Dong, Q., Yuan, F., Xu, J., Huang, S., Kong, L., Chen, J., and Li, L · 2023
Later among the works it cites.
Aya 23: Open weight releases to further multilingual progress
Aryabumi, V., Dang, J., Talupuru, D., Dash, S., Cairuz, D., Lin, H., Venkitesh, B., Smith, M., Marchisio, K., Ruder, S., et al · 2024
Closest in time.
LLM2vec: Large language models are secretly powerful text encoders
BehnamGhader, P., Adlakha, V., Mosbach, M., Bahdanau, D., Chapados, N., and Reddy, S · 2024
Closest in time.
Cross-lingual editing in multilingual language models
Beniwal, H., Singh, M., et al · 2024
Closest in time.
The reversal curse: LLMs trained on “A is B” fail to learn “B is A”
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., et al · 2022
Cited alongside, same era.
The flores-101 evaluation benchmark for low-resource and multilingual machine translation
Goyal, N., Gao, C., Chaudhary, V., Chen, P.-J., Wenzek, G., Ju, D., Krishnan, S., Ranzato, M., Guzmán, F., and Fan, A · 2022
Cited alongside, same era.
Enhancing cross-lingual natural language inference by soft prompting with multilingual verbalizer
Li, S., Hu, X., Liu, A., Yang, Y., Ma, F., Yu, P. S., and Wen, L · 2022
Cited alongside, same era.
Overcoming catastrophic forgetting in zero-shot cross-lingual generation
Vu, T., Barua, A., Lester, B., Cer, D., Iyyer, M., and Constant, N · 2022
Cited alongside, same era.
Compositional generalization in unsupervised compositional representation learning: A study on disentanglement and emergent language
Xu, Z., Niethammer, M., and Raffel, C. A · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Studying large language model generalization with influence functions
Grosse, R., Bae, J., Anil, C., Elhage, N., Tamkin, A., Tajdini, A., Steiner, B., Li, D., Durmus, E., Perez, E., et al · 2023
Cited alongside, same era.
Premise order matters in reasoning with large language models
Chen, X., Chi, R. A., Wang, X., and Zhou, D · 2024
Closest in time.
Empowering large language models on robotic manipulation with affordance prompting
Cheng, G., Zhang, C., Cai, W., Zhao, L., Sun, C., and Bian, J · 2024
Closest in time.
Physically grounded vision-language models for robotic manipulation
Gao, J., Sarkar, B., Xia, F., Xiao, T., Wu, J., Ichter, B., Majumdar, A., and Sadigh, D · 2024
Closest in time.
Zamba: A compact 7b ssm hybrid model
Glorioso, P., Anthony, Q., Tokpanov, Y., Whittington, J., Pilault, J., Ibrahim, A., and Millidge, B · 2024
Closest in time.
Exploring human-like translation strategy with large language models
He, Z., Liang, T., Jiao, W., Zhang, Z., Yang, Y., Wang, R., Tu, Z., Shi, S., and Wang, X · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Tofu: A task of fictitious unlearning for llms
Maini, P., Feng, Z., Schwarzschild, A., Lipton, Z. C., and Kolter, J. Z · 2024
Closest in time.
Introducing Meta LLaMA 3: The most capable openly available LLM to date
Meta · 2024
Closest in time.
Mistral large
Mistral · 2024
Closest in time.
Large language models sensitivity to the order of options in multiple-choice questions
Pezeshkpour, P. and Hruschka, E · 2024
Closest in time.
Unintended impacts of LLM alignment on global representation
Ryan, M. J., Held, W., and Yang, D · 2024
Closest in time.
Reuse your rewards: Reward model transfer for zero-shot cross-lingual alignment
Wu, Z., Balashankar, A., Kim, Y., Eisenstein, J., and Beirami, A · 2024
Closest in time.
A paradigm shift in machine translation: Boosting translation performance of large language models
Xu, H., Kim, Y. J., Sharaf, A., and Awadalla, H. H · 2024
Closest in time.
Skill-Mix: A flexible and expandable family of evaluations for AI models
Yu, D., Kaur, S., Gupta, A., Brown-Cohen, J., Goyal, A., and Arora, S · 2024
Closest in time.
AdaMergeX: Cross-lingual transfer with large language models via adaptive adapter merging
Zhao, Y., Zhang, W., Wang, H., Kawaguchi, K., and Bing, L · 2024
Closest in time.
Large language models are not robust multiple choice selectors
Zheng, C., Zhou, H., Meng, F., Zhou, J., and Huang, M · 2024
Closest in time.
Multilingual machine translation with large language models: Empirical results and analysis
Zhu, W., Liu, H., Dong, Q., Xu, J., Huang, S., Kong, L., Chen, J., and Li, L · 2024
Closest in time.