Fetching the paper…
Reading the bibliography…
Many existing evaluation benchmarks for Large Language Models (LLMs) quickly become outdated due to the emergence of new models and training data.
Measuring nominal scale agreement among many raters
Fleiss, J. L · 1971
Earlier work this paper cites.
Reflections of the environment in memory
Anderson, J. R. and Schooler, L. J · 1991
Earlier work this paper cites.
Okapi at trec-3
Robertson, S. E., Walker, S., Jones, S., Hancock-Beaulieu, M. M., Gatford, M., et al · 1995
Earlier work this paper cites.
A density-based algorithm for discovering clusters in large spatial databases with noise
Ester, M., Kriegel, H.-P., Sander, J., Xu, X., et al · 1996
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Common crawl news dataset
Nagel, S · 2016
Earlier work this paper cites.
Superforecasting: The art and science of prediction
Tetlock, P. E. and Gardner, D · 2016
Earlier work this paper cites.
Finding alternative translations in a large corpus of movie subtitle
Tiedemann, J · 2016
Earlier work this paper cites.
iSurvive: An interpretable, event-time prediction model for mHealth
Dempsey, W. H., Moreno, A., Scott, C. K., Dennis, M. L., Gustafson, D. H., Murphy, S. A., and Rehg, J. M · 2017
Earlier work this paper cites.
news-please: A generic news crawler and extractor
Hamborg, F., Meuschke, N., Breitinger, C., and Gipp, B · 2017
Earlier work this paper cites.
Modeling uncertainty in integrated assessment of climate change: A multimodel comparison
Gillingham, K., Nordhaus, W., Anthoff, D., Blanford, G., Bosetti, V., Christensen, P., McJeon, H., and Reilly, J · 2018
Earlier work this paper cites.
Homemade bookcorpus
Kobayashi, S · 2018
Earlier work this paper cites.
Openwebtext corpus
Gokaslan, A. and Cohen, V · 2019
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
Saxton, D., Grefenstette, E., Hill, F., and Kohli, P · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
A dataset for answering time-sensitive questions
Chen, W., Wang, X., Wang, W. Y., and Wang, W. Y · 2021
Cited alongside, same era.
ForecastQA: A question answering challenge for event forecasting with temporal text data
Jin, W., Khanna, R., Kim, S., Lee, D.-H., Morstatter, F., Galstyan, A., and Ren, X · 2021
Cited alongside, same era.
Mind the gap: Assessing temporal generalization in neural language models
Lazaridou, A., Kuncoro, A., Gribovskaya, E., Agrawal, D., Liska, A., Terzi, T., Gimenez, M., de Masson d'Autume, C., Kocisky, T., Ruder, S., Yogatama, D., Cao, K., Young, S., and Blunsom, P · 2021
Cited alongside, same era.
Temporal adaptation of BERT and performance on downstream document classification: Insights from social media
Röttger, P. and Pierrehumbert, J · 2021
Cited alongside, same era.
SituatedQA: Incorporating extra-linguistic contexts into QA
Zhang, M. and Choi, E · 2021
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Approaching human-level forecasting with language models
Halawi, D., Zhang, F., Yueh-Han, C., and Steinhardt, J · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Realtime qa: what’s the answer right now?
Kasai, J., Sakaguchi, K., Le Bras, R., Asai, A., Yu, X., Radev, D., Smith, N. A., Choi, Y., Inui, K., et al · 2024
Closest in time.
Task contamination: Language models may not be few-shot anymore
Li, C. and Flanigan, J · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Agarwal, O. and Nenkova, A · 2022
Cited alongside, same era.
Towards continual knowledge learning of language models
Jang, J., Ye, S., Yang, S., Shin, J., Han, J., Kim, G., Choi, S. J., and Seo, M · 2022
Cited alongside, same era.
Lifelong pretraining: Continually adapting language models to emerging corpora
Jin, X., Zhang, D., Zhu, H., Xiao, W., Li, S.-W., Wei, X., Arnold, A., and Ren, X · 2022
Cited alongside, same era.
Adapting a language model while preserving its general knowledge
Ke, Z., Shao, Y., Lin, H., Xu, H., Shu, L., and Liu, B · 2022
Cited alongside, same era.
Streamingqa: A benchmark for adaptation to new knowledge over time in question answering models
Liska, A., Kocisky, T., Gribovskaya, E., Terzi, T., Sezener, E., Agrawal, D., Cyprien De Masson, D., Scholtes, T., Zaheer, M., Young, S., et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2022
Cited alongside, same era.
Forecasting future world events with neural networks
Zou, A., Xiao, T., Jia, R., Kwon, J., Mazeika, M., Li, R., Song, D., Steinhardt, J., Evans, O., and Hendrycks, D · 2022
Cited alongside, same era.
Consent in crisis: The rapid decline of the ai data commons
Longpre, S., Mahari, R., Lee, A., Lund, C., Oderinwale, H., Brannon, W., Saxena, N., Obeng-Marnu, N., South, T., Hunter, C., et al · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., et al · 2024
Closest in time.
FreshLLMs: Refreshing large language models with search engine augmentation
Vu, T., Iyyer, M., Wang, X., Constant, N., Wei, J., Wei, J., Tar, C., Sung, Y.-H., Zhou, D., Le, Q., and Luong, T · 2024
Closest in time.
Benchmark data contamination of large language models: A survey
Xu, C., Guan, S., Greene, D., and Kechadi, M.-T · 2024
Closest in time.
Autocast++: Enhancing world event prediction with zero-shot ranking-based context retrieval
Yan, Q., Seraj, R., He, J., Meng, L., and Sylvain, T · 2024
Closest in time.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., et al · 2024
Closest in time.
Mirai: Evaluating llm agents for event forecasting
Ye, C., Hu, Z., Deng, Y., Huang, Z., Ma, M. D., Zhu, Y., and Wang, W · 2024
Closest in time.
Investigating continual pretraining in large language models: Insights and implications
Yıldız, Ç., Ravichandran, N. K., Punia, P., Bethge, M., and Ermis, B · 2024
Closest in time.
Analyzing temporal complex events with large language models? a benchmark towards temporal, long context understanding
Zhang, Z., Cao, Y., Ye, C., Ma, Y., Liao, L., and Chua, T.-S · 2024
Closest in time.
Forecastbench: A dynamic benchmark of ai forecasting capabilities
Karger, E., Bastani, H., Yueh-Han, C., Jacobs, Z., Halawi, D., Zhang, F., and Tetlock, P. E · 2025
Closest in time.
Inadequacies of large language model benchmarks in the era of generative artificial intelligence
McIntosh, T. R., Susnjak, T., Arachchilage, N., Liu, T., Xu, D., Watters, P., and Halgamuge, M. N · 2025
Closest in time.
Is your LLM outdated? a deep look at temporal generalization
Zhu, C., Chen, N., Gao, Y., Zhang, Y., Tiwari, P., and Wang, B · 2025
Closest in time.