Fetching the paper…
Reading the bibliography…
We introduce a dynamic benchmarking system for conversational agents that evaluates their performance through a single, simulated, and lengthy user$\leftrightarrow$agent interaction.
‘improving ratings’: audit in the british university system
M. Strathern · 1997
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
T. Kočiský, J. Schwarz, P. Blunsom, C. Dyer, K. M. Hermann, G. Melis, and E. Grefenstette · 2017
Earlier work this paper cites.
A roadmap towards machine intelligence
T. Mikolov, A. Joulin, and M. Baroni · 2018
Earlier work this paper cites.
Evaluating theory of mind in question answering
A. Nematzadeh, K. Burns, E. Grant, A. Gopnik, and T. Griffiths · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
X. Ho, A.-K. Duong Nguyen, S. Sugawara, and A. Aizawa · 2020
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
J. W. Rae, A. Potapenko, S. M. Jayakumar, C. Hillier, and T. P. Lillicrap · 2020
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers
P. Dasigi, K. Lo, I. Beltagy, A. Cohan, N. A. Smith, and M. Gardner · 2021
Earlier work this paper cites.
Improving reproducibility in machine learning research (a report from the neurips 2019 reproducibility program)
J. Pineau, P. Vincent-Lamarre, K. Sinha, V. Larivière, A. Beygelzimer, F. d’Alché Buc, E. Fox, and H. Larochelle · 2021
Earlier work this paper cites.
ChapterBreak: A challenge dataset for long-range language models
S. Sun, K. Thai, and M. Iyyer · 2022
Earlier work this paper cites.
Y. Wu, M. N. Rabe, D. Hutchins, and C. Szegedy · 2022
Earlier work this paper cites.
Beyond goldfish memory: Long-term open-domain conversation
J. Xu, A. Szlam, and J. Weston · 2022
Earlier work this paper cites.
L-eval: Instituting standardized evaluation for long context language models
C. An, S. Gong, M. Zhong, X. Zhao, M. Li, J. Zhang, L. Kong, and X. Qiu · 2023
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding
Y. Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, Z. Huang, Z. Du, X. Liu, A. Zeng, L. Hou, Y. Dong, J. Tang, and J. Li · 2023
Cited alongside, same era.
Extending context window of large language models via positional interpolation
S. Chen, S. Wong, L. Chen, and Y. Tian · 2023
Cited alongside, same era.
Llf-bench: Benchmark for interactive learning from language feedback
C.-A. Cheng, A. Kolobov, D. Misra, A. Nie, and A. Swaminathan · 2023
Cited alongside, same era.
Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks
A. Jacovi, A. Caciularu, O. Goldman, and Y. Goldberg · 2023
Cited alongside, same era.
Evaluating language-model agents on realistic autonomous tasks
Needle In A Needlestack
T. Burns · 2024
Closest in time.
BAMBOO: A comprehensive benchmark for evaluating long text modeling capacities of large language models
Z. Dong, T. Tang, J. Li, W. X. Zhao, and J.-R. Wen · 2024
Closest in time.
Ruler: What’s the real context size of your long-context language models?
C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, and B. Ginsburg · 2024
Closest in time.
A survey on retrieval-augmented text generation for large language models
Y. Huang and J. Huang · 2024
Closest in time.
Needle In A Haystack - Pressure Testing LLMs
G. Kamradt · 2024
Closest in time.
Multiple Needles In A Haystack
L. Knight-Webb · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Kinniment, L. J. K. Sato, H. Du, B. Goodrich, M. Hasin, L. Chan, L. H. Miles, T. R. Lin, H. Wijk, J. Burget, et al · 2023
Cited alongside, same era.
Loogle: Can long-context language models understand long contexts?
J. Li, M. Wang, Z. Zheng, and M. Zhang · 2023
Cited alongside, same era.
Agentsims: An open-source sandbox for large language model evaluation
J. Lin, H. Zhao, A. Zhang, Y. Wu, H. Ping, and Q. Chen · 2023
Cited alongside, same era.
Memgpt: Towards llms as operating systems
C. Packer, V. Fang, S. G. Patil, K. Lin, S. Wooders, and J. E. Gonzalez · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2023
Cited alongside, same era.
Augmenting language models with long-term memory
W. Wang, L. Dong, H. Cheng, X. Liu, X. Yan, J. Gao, and F. Wei · 2023
Cited alongside, same era.
Rethinking benchmark and contamination for language models with rephrased samples
S. Yang, W.-L. Chiang, L. Zheng, J. E. Gonzalez, and I. Stoica · 2023
Cited alongside, same era.
Webarena: A realistic web environment for building autonomous agents
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, Y. Bisk, D. Fried, U. Alon, et al · 2023
Cited alongside, same era.
Closest in time.
Agentbench: Evaluating llms as agents
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, S. Zhang, X. Deng, A. Zeng, Z. Du, C. Zhang, S. Shen, T. Zhang, Y. Su, H. Sun, M. Huang, Y. Dong, and J. Tang · 2024
Closest in time.
Hello gpt-4o
OpenAI · 2024
Closest in time.
The Haystack Matters for NIAH Evals
L. Pekelis · 2024
Closest in time.
YaRN: Efficient context window extension of large language models
B. Peng, J. Quesnelle, H. Fan, and E. Shippole · 2024
Closest in time.
Memoryllm: Towards self-updatable large language models, 2024
Y. Wang, Y. Gao, X. Chen, H. Jiang, S. Li, J. Yang, Q. Yin, Z. Li, X. Li, B. Yin, J. Shang, and J. McAuley · 2024
Closest in time.
Berkeley function calling leaderboard
F. Yan, H. Mao, C. C.-J. Ji, T. Zhang, S. G. Patil, I. Stoica, and J. E. Gonzalez · 2024
Closest in time.
Memorybank: Enhancing large language models with long-term memory
W. Zhong, L. Guo, Q. Gao, H. Ye, and Y. Wang · 2024
Closest in time.