Fetching the paper…
Reading the bibliography…
While recent large language models (LLMs) demonstrate remarkable abilities in responding to queries in diverse languages, their ability to handle long multilingual contexts is unexplored.
On the cross-lingual transferability of monolingual representations
Artetxe, M.; Ruder, S.; and Yogatama, D. 2019 · 1910
Earlier work this paper cites.
How Context Affects Language Models’ Factual Predictions
Petroni, F.; Lewis, P.; Piktus, A.; Rocktäschel, T.; Wu, Y.; Miller, A. H.; and Riedel, S. 2020 · 2005
Earlier work this paper cites.
MKQA: A Linguistically Diverse Benchmark for Multilingual Open Domain Question Answering
Longpre, S.; Lu, Y.; and Daiber, J. 2021 · 2007
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use Context
Khandelwal, U.; He, H.; Qi, P.; and Jurafsky, D. 2018 · 2018
Earlier work this paper cites.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J.; Le, Q.; and Salakhutdinov, R. 2019 · 2019
Earlier work this paper cites.
MLQA: Evaluating Cross-lingual Extractive Question Answering
Lewis, P.; Oguz, B.; Rinott, R.; Riedel, S.; and Schwenk, H. 2020 · 2020
Earlier work this paper cites.
Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation
Reimers, N.; and Gurevych, I. 2020 · 2020
Earlier work this paper cites.
Towards End-to-End Multilingual Question Answering
Loginova, E.; Varanasi, S.; and Neumann, G. 2021 · 2021
Earlier work this paper cites.
mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset
Bonifacio, L.; Jeronymo, V.; Abonizio, H. Q.; Campiotti, I.; Fadaee, M.; Lotufo, R.; and Nogueira, R. 2022 · 2022
Earlier work this paper cites.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Dao, T.; Fu, D. Y.; Ermon, S.; Rudra, A.; and Ré, C. 2022 · 2022
Earlier work this paper cites.
CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities
Lee, M.; Liang, P.; and Yang, Q. 2022 · 2022
Cited alongside, same era.
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Press, O.; Smith, N. A.; and Lewis, M. 2022 · 2022
Cited alongside, same era.
LaMDA: Language Models for Dialog Applications
Thoppilan, R.; Freitas, D. D.; Hall, J.; Shazeer, N.; Kulshreshtha, A.; Cheng, H.-T.; Jin, A.; Bos, T.; Baker, L.; Du, Y.; Li, Y.; Lee, H.; Zheng, H. S.; Ghafouri, A.; Menegali, M.; Huang, Y.; Krikun, M.; Lepikhin, D.; Qin, J.; Chen, D.; Xu, Y.; Chen, Z.; Roberts, A.; Bosma, M.; Zhao, V.; Zhou, Y.; Chang, C.-C.; Krivokon, I.; Rusch, W.; Pickett, M.; Srinivasan, P.; Man, L.; Meier-Hellstern, K.; Morris, M. R.; Doshi, T.; Santos, R. D.; Duke, T.; Soraker, J.; Zevenbergen, B.; Prabhakaran, V.; Diaz, M.; Hutchinson, B.; Olson, K.; Molina, A.; Hoffman-John, E.; Lee, J.; Aroyo, L.; Rajakumar, R.; Butryna, A.; Lamm, M.; Kuzmina, V.; Fenton, J.; Cohen, A.; Bernstein, R.; Kurzweil, R.; Aguera-Arcas, B.; Cui, C.; Croak, M.; Chi, E.; and Le, Q. 2022 · 2022
Cited alongside, same era.
Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; and Martin, L. 2023 · 2023
Later among the works it cites.
Llama 3 Model Card
AI@Meta. 2024 · 2024
Closest in time.
Aya 23: Open Weight Releases to Further Multilingual Progress
Aryabumi, V.; Dang, J.; Talupuru, D.; Dash, S.; Cairuz, D.; Lin, H.; Venkitesh, B.; Smith, M.; Campos, J. A.; Tan, Y. C.; Marchisio, K.; Bartolo, M.; Ruder, S.; Locatelli, A.; Kreutzer, J.; Frosst, N.; Gomez, A.; Blunsom, P.; Fadaee, M.; Üstün, A.; and Hooker, S. 2024 · 2024
Closest in time.
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Bai, Y.; Lv, X.; Zhang, J.; Lyu, H.; Tang, J.; Huang, Z.; Du, Z.; Liu, X.; Zeng, A.; Hou, L.; Dong, Y.; Tang, J.; and Li, J. 2024 · 2024
Closest in time.
Dubey, A.; and Jauhri, A. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, X.; Thakur, N.; Ogundepo, O.; Kamalloo, E.; Alfonso-Hermelo, D.; Li, X.; Liu, Q.; Rezagholizadeh, M.; and Lin, J. 2022 · 2022
Cited alongside, same era.
L-Eval: Instituting Standardized Evaluation for Long Context Language Models
An, C.; Gong, S.; Zhong, M.; Li, M.; Zhang, J.; Kong, L.; and Qiu, X. 2023 · 2023
Cited alongside, same era.
Analyzing and Reducing the Performance Gap in Cross-Lingual Transfer with Fine-tuning Slow and Fast
Guo, Y.; Liang, Y.; Zhao, D.; Liu, B.; and Duan, N. 2023 · 2023
Cited alongside, same era.
Efficient Long-Text Understanding with Short-Text Models
Ivgi, M.; Shaham, U.; and Berant, J. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; de Las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; Lavaud, L. R.; Lachaux, M.-A.; Stock, P.; Scao, T. L.; Lavril, T.; Wang, T.; Lacroix, T.; and Sayed, W. E. 2023 · 2023
Cited alongside, same era.
Lost in the Middle: How Language Models Use Long Contexts
Liu, N. F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P. 2023 · 2023
Cited alongside, same era.
The NLP Task Effectiveness of Long-Range Transformers
Qin, G.; Feng, Y.; and Van Durme, B. 2023 · 2023
Cited alongside, same era.
ZeroSCROLLS: A Zero-Shot Benchmark for Long Text Understanding
Shaham, U.; Ivgi, M.; Efrat, A.; Berant, J.; and Levy, O. 2023b · 2023
Cited alongside, same era.
RULER: What’s the Real Context Size of Your Long-Context Language Models?
Hsieh, C.-P.; Sun, S.; Kriman, S.; Acharya, S.; Rekesh, D.; Jia, F.; Zhang, Y.; and Ginsburg, B. 2024a
Cited in the paper.
Closest in time.
Multilingual Pretraining and Instruction Tuning Improve Cross-Lingual Knowledge Alignment, But Only Shallowly
Gao, C.; Hu, H.; Hu, P.; Chen, J.; Li, J.; and Huang, S. 2024 · 2024
Closest in time.
Long-context LLMs Struggle with Long In-context Learning
Li, T.; Zhang, G.; Do, Q. D.; Yue, X.; and Chen, W. 2024 · 2024
Closest in time.
Multilingual Instruction Tuning With Just a Pinch of Multilinguality
Shaham, U.; Herzig, J.; Aharoni, R.; Szpektor, I.; Tsarfaty, R.; and Eyal, M. 2024 · 2024
Closest in time.
Wang, H.; Shi, H.; Tan, S.; Qin, W.; Wang, W.; Zhang, T.; Nambi, A.; Ganu, T.; and Wang, H. 2024 · 2024
Closest in time.
∞ \infty Bench: Extending Long Context Evaluation Beyond 100K Tokens
Zhang, X.; Chen, Y.; Hu, S.; Xu, Z.; Chen, J.; Hao, M. K.; Han, X.; Thai, Z. L.; Wang, S.; Liu, Z.; and Sun, M. 2024 · 2024
Closest in time.
How do Large Language Models Handle Multilingualism?
Zhao, Y.; Zhang, W.; Chen, G.; Kawaguchi, K.; and Bing, L. 2024 · 2024
Closest in time.