Fetching the paper…
Reading the bibliography…
We present ONERULER, a multilingual benchmark designed to evaluate long-context language models across 26 languages.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew · 1910
Earlier work this paper cites.
Know what you don‘t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Earlier work this paper cites.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury · 2020
Earlier work this paper cites.
Some languages are more equal than others: Probing deeper into the linguistic disparity in the NLP world
Surangika Ranathunga and Nisansa de Silva · 2022
Earlier work this paper cites.
Do all languages cost the same? tokenization in the era of commercial language models
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith, and Yulia Tsvetkov · 2023
Earlier work this paper cites.
Bamboo: A comprehensive benchmark for evaluating long text modeling capacities of large language models
Zican Dong, Tianyi Tang, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen · 2023
Earlier work this paper cites.
Needle in a haystack - pressure testing llms
Gregory Kamradt · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Haotong Zhang, and Ion Stoica · 2023
Earlier work this paper cites.
Zeroscrolls: A zero-shot benchmark for long text understanding
Uri Shaham, Maor Ivgi, Avia Efrat, Jonathan Berant, and Omer Levy · 2023
Earlier work this paper cites.
Evaluating multilingual long-context models for retrieval and reasoning
Ameeta Agrawal, Andy Dang, Sina Bagheri Nezhad, Rhitabrat Pokharel, and Russell Scheinberg · 2024
Earlier work this paper cites.
L-eval: Instituting standardized evaluation for long context language models
Chenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao, Mukai Li, Jun Zhang, Lingpeng Kong, and Xipeng Qiu · 2024
Earlier work this paper cites.
LongBench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li · 2024
Earlier work this paper cites.
Introducing goodai ltm benchmark
David Castillo, Joseph Davidson, Finlay Gray, José Solorzano, and Marek Rosa · 2024
Earlier work this paper cites.
How to train long-context language models (effectively)
Tianyu Gao, Alexander Wettig, Howard Yen, and Danqi Chen · 2024
Earlier work this paper cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, December 2024
Gemini Team · 2024
Cited alongside, same era.
Amey Hengle, Prasoon Bajpai, Soham Dan, and Tanmoy Chakraborty · 2024
Cited alongside, same era.
RULER: What’s the real context size of your long-context language models?
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, and Boris Ginsburg · 2024
Cited alongside, same era.
One thousand and one pairs: A “novel” challenge for long-context language models
Marzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal, and Mohit Iyyer · 2024
Cited alongside, same era.
FABLES: Evaluating faithfulness and content selection in book-length summarization
Yekyung Kim, Yapei Chang, Marzena Karpinska, Aparna Garimella, Varun Manjunatha, Kyle Lo, Tanya Goyal, and Mohit Iyyer · 2024
A benchmark for learning to translate a new language from one grammar book
Garrett Tanzer, Mirac Suzgun, Eline Visser, Dan Jurafsky, and Luke Melas-Kyriazi · 2024
Later among the works it cites.
Stress-testing long-context language models with lifelong ICL and task haystack
Xiaoyue Xu, Qinyuan Ye, and Xiang Ren · 2024
Later among the works it cites.
Lv-eval: A balanced long-context benchmark with 5 length levels up to 256k, 2024
Tao Yuan, Xuefei Ning, Dong Zhou, Zhijie Yang, Shiyao Li, Minghui Zhuang, Zheyue Tan, Zhuyu Yao, Dahua Lin, Boxun Li, Guohao Dai, Shengen Yan, and Yu Wang · 2024
Later among the works it cites.
∞ \infty Bench: Extending long context evaluation beyond 100K tokens
Xinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Hao, Xu Han, Zhen Thai, Shuo Wang, Zhiyuan Liu, and Maosong Sun · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BABILong: Testing the limits of LLMs with long context reasoning-in-a-haystack
Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Igorevich Sorokin, Artyom Sorokin, and Mikhail Burtsev · 2024
Cited alongside, same era.
Summary of a haystack: A challenge to long-context llms and rag systems
Philippe Laban, A. R. Fabbri, Caiming Xiong, and Chien-Sheng Wu · 2024
Cited alongside, same era.
Can long-context language models subsume retrieval, rag, sql, and more?
Jinhyuk Lee, Anthony Chen, Zhuyun Dai, Dheeru Dua, Devendra Singh Sachan, Michael Boratko, Yi Luan, Sébastien M. R. Arnold, Vincent Perot, Siddharth Dalmia, Hexiang Hu, Xudong Lin, Panupong Pasupat, Aida Amini, Jeremy R. Cole, Sebastian Riedel, Iftekhar Naim, Ming-Wei Chang, and Kelvin Guu · 2024
Cited alongside, same era.
Same task, more tokens: the impact of input length on the reasoning performance of large language models
Mosh Levy, Alon Jacoby, and Yoav Goldberg · 2024
Cited alongside, same era.
LooGLE: Can long-context language models understand long contexts?
Jiaqi Li, Mengmeng Wang, Zilong Zheng, and Muhan Zhang · 2024
Cited alongside, same era.
RepoQA: Evaluating long context code understanding
Jiawei Liu, Jia Le Tian, Vijay Daita, Yuxiang Wei, Yifeng Ding, Yuhan Katherine Wang, Jun Yang, and LINGMING ZHANG · 2024
Cited alongside, same era.
The Llama 3 Herd of Models, November 2024
Llama Team · 2024
Cited alongside, same era.
Benchmax: A comprehensive multilingual evaluation suite for large language models, 2025
Xu Huang, Wenhao Zhu, Hanxu Hu, Conghui He, Lei Li, Shujian Huang, and Fei Yuan · 2025
Closest in time.
Jamba: Hybrid transformer-mamba language models
Barak Lenz, Opher Lieber, Alan Arazi, Amir Bergman, Avshalom Manevich, Barak Peleg, Ben Aviram, Chen Almagor, Clara Fridman, Dan Padnos, Daniel Gissin, Daniel Jannai, Dor Muhlgay, Dor Zimberg, Edden M. Gerber, Elad Dolev, Eran Krakovsky, Erez Safahi, Erez Schwartz, Gal Cohen, Gal Shachaf, Haim Rozenblum, Hofit Bata, Ido Blass, Inbal Magar, Itay Dalmedigos, Jhonathan Osin, Julie Fadlon, Maria Rozman, Matan Danos, Michael Gokhman, Mor Zusman, Naama Gidron, Nir Ratner, Noam Gat, Noam Rozen, Oded Fried, Ohad Leshno, Omer Antverg, Omri Abend, Or Dagan, Orit Cohavi, Raz Alon, Ro’i Belson, Roi Cohen, Rom Gilad, Roman Glozman, Shahar Lev, Shai Shalev-Shwartz, Shaked Haim Meirom, Tal Delbari, Tal Ness, Tomer Asida, Tom Ben Gal, Tom Braude, Uriya Pumerantz, Josh Cohen, Yonatan Belinkov, Yuval Globerson, Yuval Peleg Levy, and Yoav Shoham · 2025
Closest in time.
Longreason: A synthetic long-context reasoning benchmark via context expansion, 2025
Zhan Ling, Kang Liu, Kai Yan, Yifan Yang, Weijian Lin, Ting-Han Fan, Lingfeng Shen, Zhengyin Du, and Jiecao Chen · 2025
Closest in time.
Openai o3-mini system card, January 2025
OpenAI · 2025
Closest in time.
Qwen2.5 Technical Report, January 2025
Qwen Team · 2025
Closest in time.
Counting-stars: A multi-evidence, position-aware, and scalable benchmark for evaluating long-context large language models
Mingyang Song, Mao Zheng, and Xuan Luo · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models, 2025
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen, Zhong, Hanjie Chen, and Xia Hu · 2025
Closest in time.
HELMET: How to evaluate long-context models effectively and thoroughly
Howard Yen, Tianyu Gao, Minmin Hou, Ke Ding, Daniel Fleischer, Peter Izsak, Moshe Wasserblat, and Danqi Chen · 2025
Closest in time.
Yang Zhou, Hongyi Liu, Zhuoming Chen, Yuandong Tian, and Beidi Chen · 2025
Closest in time.