Fetching the paper…
Reading the bibliography…
Recent studies have shown that Large Language Models (LLMs) struggle to accurately retrieve information and maintain reasoning capabilities when processing long-context inputs.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
O. Sener and S. Savarese · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2018
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
A. Ghorbani and J. Zou · 2019
Earlier work this paper cites.
Latent retrieval for weakly supervised open domain question answering
K. Lee, M.-W. Chang, and K. Toutanova · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
L-eval: Instituting standardized evaluation for long context language models
C. An, S. Gong, M. Zhong, M. Li, J. Zhang, L. Kong, and X. Qiu · 2023
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding
Y. Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, Z. Huang, Z. Du, X. Liu, A. Zeng, L. Hou, et al · 2023
Earlier work this paper cites.
Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance
Y. Fu, L. Ou, M. Chen, Y. Wan, H. Peng, and T. Khot · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou · 2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Cited alongside, same era.
H. Junqing, P. Kunhao, D. Xiaoqun, S. Zhuoyang, L. Yibo, L. Yuxin, W. Hao, S. Qianguo, Z. Songxin, X. Zejian, et al · 2023
Cited alongside, same era.
A survey on data selection for language models
A. Albalak, Y. Elazar, S. M. Xie, S. Longpre, N. Lambert, X. Wang, N. Muennighoff, B. Hou, L. Pan, H. Jeong, et al · 2024
Closest in time.
Make your llm fully utilize the context
S. An, Z. Ma, Z. Lin, N. Zheng, and J.-G. Lou · 2024
Closest in time.
Data engineering for scaling language models to 128k context
Y. Fu, R. Panda, X. Niu, X. Yue, H. Hajishirzi, Y. Kim, and H. Peng · 2024
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al · 2024
Closest in time.
Does fine-tuning llms on new knowledge encourage hallucinations?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Needle in a haystack - pressure testing llms
G. Kamradt · 2023
Cited alongside, same era.
Teaching arithmetic to small transformers
N. Lee, K. Sreenivasan, J. D. Lee, K. Lee, and D. Papailiopoulos · 2023
Cited alongside, same era.
Lost in the middle: How language models use long contexts
N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang · 2023
Cited alongside, same era.
Landmark attention: Random-access infinite context length for transformers
A. Mohtashami and M. Jaggi · 2023
Cited alongside, same era.
Attention sorting combats recency bias in long context language models
A. Peysakhovich and A. Lerer · 2023
Cited alongside, same era.
Found in the middle: Permutation self-consistency improves listwise ranking in large language models
R. Tang, X. Zhang, X. Ma, J. Lin, and F. Ture · 2023
Cited alongside, same era.
Retrieval meets long context large language models
P. Xu, W. Ping, X. Wu, L. McAfee, C. Zhu, Z. Liu, S. Subramanian, E. Bakhturina, M. Shoeybi, and B. Catanzaro · 2023
Cited alongside, same era.
Data-centric ai: Perspectives and challenges
D. Zha, Z. P. Bhat, K.-H. Lai, F. Yang, and X. Hu · 2023
Cited alongside, same era.
Z. Gekhman, G. Yona, R. Aharoni, M. Eyal, A. Feder, R. Reichart, and J. Herzig · 2024
Closest in time.
Llm maybe longlm: Selfextend llm context window without tuning
H. Jin, X. Han, J. Yang, Z. Jiang, Z. Liu, C.-Y. Chang, H. Chen, and X. Hu · 2024
Closest in time.
M. Levy, A. Jacoby, and Y. Goldberg · 2024
Closest in time.
Dataperf: Benchmarks for data-centric ai development
M. Mazumder, C. Banbury, X. Yao, B. Karlaš, W. Gaviria Rojas, S. Diamos, G. Diamos, L. He, A. Parrish, H. R. Kirk, et al · 2024
Closest in time.
Chatgpt, 2023
OpenAI · 2024
Closest in time.
Training with “paraphrasing the original text” improves long-context performance, 2024
Y. Yu · 2024
Closest in time.
Z. Zhang, R. Chen, S. Liu, Z. Yao, O. Ruwase, B. Chen, X. Wu, and Z. Wang · 2024
Closest in time.
Transformers can achieve length generalization but not robustly
Y. Zhou, U. Alon, X. Chen, X. Wang, R. Agarwal, and D. Zhou · 2024
Closest in time.