Fetching the paper…
Reading the bibliography…
Existing multilingual long-context benchmarks, often based on the popular needle-in-a-haystack test, primarily evaluate a model's ability to locate specific information buried within irrelevant texts.
Towards ai-complete question answering: A set of prerequisite toy tasks
Weston, Jason, Antoine Bordes, Sumit Chopra, Alexander M. Rush, Bart van Merriënboer, Armand Joulin, and Tomas Mikolov. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, Pranav, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Language models are few-shot learners
Brown, Tom B., Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
The state and fate of linguistic diversity and inclusion in the NLP world
Joshi, Pratik, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Earlier work this paper cites.
MLQA: Evaluating cross-lingual extractive question answering
Lewis, Patrick, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020 · 2020
Earlier work this paper cites.
Xcopa: A multilingual dataset for causal commonsense reasoning
Ponti, Edoardo M., Olga Majewska Goran Glava, Qianchu Liuand Ivan Vuliand, and Anna Korhonen. 2020 · 2020
Earlier work this paper cites.
XL-WiC: A multilingual benchmark for evaluating semantic contextualization
Raganato, Alessandro, Tommaso Pasini, Jose Camacho-Collados, and Mohammad Taher Pilehvar. 2020 · 2020
Earlier work this paper cites.
Making monolingual sentence embeddings multilingual using knowledge distillation
Reimers, Nils and Iryna Gurevych. 2020 · 2020
Earlier work this paper cites.
Automatic machine translation evaluation using source language inputs and cross-lingual language model
Takahashi, Kosuke, Katsuhito Sudoh, and Satoshi Nakamura. 2020 · 2020
Earlier work this paper cites.
mmarco: A multilingual version of the ms marco passage ranking dataset
Bonifacio, Luiz, Vitor Jeronymo, Hugo Queiroz Abonizio, Israel Campiotti, Marzieh Fadaee, Roberto Lotufo, and Rodrigo Nogueira. 2022 · 2022
Earlier work this paper cites.
Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities
Lee, Mina, Percy Liang, and Qian Yang. 2022 · 2022
Earlier work this paper cites.
MEGA: Multilingual evaluation of generative AI
Ahuja, Kabir, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2023 · 2023
Earlier work this paper cites.
L-eval: Instituting standardized evaluation for long context language models
An, Chen, Shansan Gong, Ming Zhong, Mukai Li, Jun Zhang, Lingpeng Kong, and Xipeng Qiu. 2023 · 2023
Earlier work this paper cites.
Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting
Huang, Haoyang, Tianyi Tang, Dongdong Zhang, Wayne Xin Zhao, Ting Song, Yan Xia, and Furu Wei. 2023 · 2023
Earlier work this paper cites.
Needle in a haystack- pressure testing llms
Kamradt, Gregory. 2023 · 2023
Earlier work this paper cites.
Lost in the middle: How language models use long contexts
Liu, Nelson F., Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023 · 2023
Earlier work this paper cites.
Lost in translation: Large language models in non-english content analysis
Nicholas, Gabriel and Aliya Bhatia. 2023 · 2023
Earlier work this paper cites.
Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages
Qin, Libo, Qiguang Chen, Fuxuan Wei, Shijue Huang, and Wanxiang Che. 2023 · 2023
Cited alongside, same era.
Zeroscrolls: A zero-shot benchmark for long text understanding
Shaham, Uri, Maor Ivgi, Avia Efrat, Jonathan Berant, and Omer Levy. 2023 · 2023
Cited alongside, same era.
Multilingual llms are better cross-lingual in-context learners with alignment
Tanwar, Eshaan, Manish Borthakur, Subhabrata Dutta, and Tanmoy Chakraborty. 2023 · 2023
Cited alongside, same era.
Evaluating multilingual long-context models for retrieval and reasoning
Agrawal, Ameeta, Andy Dang, Sina Bagheri Nezhad, Rhitabrat Pokharel, and Russell Scheinberg. 2024 · 2024
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
Anthropic. 2024 · 2024
Cited alongside, same era.
Needlebench: Can llms do retrieval and reasoning in 1 million context window?
Li, Mo, Songyang Zhang, Yunxin Liu, and Kai Chen. 2024 · 2024
Later among the works it cites.
Llama 3.1 8B Instruct
Meta AI. 2024 · 2024
Later among the works it cites.
Gpt-4 technical report
OpenAI. 2024 · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, Machel, Nikolay Savinov, and Denis Teplyashin et al. 2024 · 2024
Later among the works it cites.
Code llama: Open foundation models for code
Rozière, Baptiste, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2024 · 2024
Later among the works it cites.
Multi-document financial question answering using llms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
LongBench: A bilingual, multitask benchmark for long context understanding
Bai, Yushi, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2024 · 2024
Cited alongside, same era.
Long-span question-answering: Automatic question generation and qa-system ranking via side-by-side evaluation
Bohnet, Bernd, Kevin Swersky, Rosanne Liu, Pranjal Awasthi, Azade Nova, Javier Snaider, Hanie Sedghi, Aaron T Parisi, Michael Collins, Angeliki Lazaridou, Orhan Firat, and Noah Fiedel. 2024 · 2024
Cited alongside, same era.
One mind, many tongues: A deep dive into language-agnostic knowledge neurons in large language models
Cao, Pengfei, Yuheng Chen, Zhuoran Jin, Yubo Chen, Kang Liu, and Jun Zhao. 2024 · 2024
Cited alongside, same era.
Optimised grouped-query attention mechanism for transformers
Chen, Yuang, Cheng Zhang, Xitong Gao, Robert D. Mullins, George A. Constantinides, and Yiren Zhao. 2024 · 2024
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Gao, Yunfan, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024 · 2024
Cited alongside, same era.
Is it really long context if all you need is retrieval? towards genuinely difficult long context nlp
Goldman, Omer, Alon Jacovi, Aviv Slobodkin, Aviya Maimon, Ido Dagan, and Reut Tsarfaty. 2024 · 2024
Cited alongside, same era.
Multilingual needle in a haystack: Investigating long-context behavior of multilingual large language models
Hengle, Amey, Prasoon Bajpai, Soham Dan, and Tanmoy Chakraborty. 2024 · 2024
Cited alongside, same era.
Shah, Shalin, Srikanth Ryali, and Ramasubbu Venkatesh. 2024 · 2024
Later among the works it cites.
Michelangelo: Long context evaluations beyond haystacks via latent structure queries
Vodrahalli, Kiran, Santiago Ontanon, Nilesh Tripuraneni, Kelvin Xu, Sanil Jain, Rakesh Shivanna, Jeffrey Hui, Nishanth Dikkala, Mehran Kazemi, Bahare Fatemi, Rohan Anil, Ethan Dyer, Siamak Shakeri, Roopali Vij, Harsh Mehta, Vinay Ramasesh, Quoc Le, Ed Chi, Yifeng Lu, Orhan Firat, Angeliki Lazaridou, Jean-Baptiste Lespiau, Nithya Attaluri, and Kate Olszewska. 2024 · 2024
Later among the works it cites.
Multimodal needle in a haystack: Benchmarking long-context capability of multimodal large language models
Wang, Hengyi, Haizhou Shi, Shiwei Tan, Weiyi Qin, Wenyuan Wang, Tunyu Zhang, Akshay Nambi, Tanuja Ganu, and Hao Wang. 2024 · 2024
Later among the works it cites.
∞ \infty bench: Extending long context evaluation beyond 100k tokens
Zhang, Xinrong, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Khai Hao, Xu Han, Zhen Leng Thai, Shuo Wang, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
Later among the works it cites.
How do large language models handle multilingualism?
Zhao, Yiran, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi, and Lidong Bing. 2024 · 2024
Later among the works it cites.
From local to global: A graph rag approach to query-focused summarization
Edge, Darren, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2025 · 2025
Closest in time.
Benchmax: A comprehensive multilingual evaluation suite for large language models
Huang, Xu, Wenhao Zhu, Hanxu Hu, Conghui He, Lei Li, Shujian Huang, and Fei Yuan. 2025 · 2025
Closest in time.
One ruler to measure them all: Benchmarking multilingual long-context language models
Kim, Yekyung, Jenna Russell, Marzena Karpinska, and Mohit Iyyer. 2025 · 2025
Closest in time.
Beyond english: The impact of prompt translation strategies across languages and tasks in multilingual llms
Mondshine, Itai, Tzuf Paz-Argaman, and Reut Tsarfaty. 2025 · 2025
Closest in time.
Evaluation metrics for search and recommendation systems
Monigatti, Leonie. 2025 · 2025
Closest in time.
Jina Reranker V2 Base Multilingual
Multilingual, Jina Reranker V2 Base. 2024 · 2025
Closest in time.