Fetching the paper…
Reading the bibliography…
Existing mathematical reasoning benchmarks are predominantly English only or translation-based, which can introduce semantic drift and mask languagespecific reasoning errors.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2020) · 2009
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., and Schulman, J. (2021) · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Lee, K., Mazeika, M., and Steinhardt, J. (2021) · 2021
Earlier work this paper cites.
Challenges and applications of large language models
Kaddour, J. et al. (2023) · 2023
Earlier work this paper cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P. et al. (2023) · 2023
Earlier work this paper cites.
Humanit’s last exam: Benchmarking extreme reasoning in llms
Zheng, S., Wang, Y., Huang, K., and Shum, H. (2023) · 2023
Earlier work this paper cites.
Bayes, E., Azime, I. A., Alabi, J. O., Kgomo, J., Eloundou, T., Proehl, E., Chen, K., Khadir, I., Etori, N. A., Muhammad, S. H., Mpanza, C., Thete, I. P., Klakow, D., and Adelani, D. I. (2024) · 2024
Cited alongside, same era.
Command r+ model card
Cohere (2024) · 2024
Cited alongside, same era.
Generalization or memorization: Data contamination and trustworthy evaluation for large language models
Dong, Y., Jiang, X., Liu, H., Jin, Z., Gu, B., Yang, M., and Li, G. (2024) · 2024
Cited alongside, same era.
Fit for our purpose, not yours: Benchmark for a low-resource indigenous language
Duncan, S., Leoni, G., Steven, L., Mahelona, K., and Jones, P.-L. (2024) · 2024
Cited alongside, same era.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai
Glazer, E., Erdil, E., Besiroglu, T., Chicharro, D., Chen, E., Gunning, A., Olsson, C., Denain, J., Ho, A., Santos, E., et al. (2024) · 2024
Can generative ai solve geometry problems? strengths and weaknesses of llms for geometric reasoning in spanish
Parra, V., Sureda, P., Corica, A., Schiaffino, S., and Godoy, D. (2024) · 2024
Later among the works it cites.
Spanish and llm benchmarks: is mmlu lost in translation?
Plaza, I., Melero, N., del Pozo, C., Conde, J., Reviriego, P., Mayor-Rocher, M., and Grandury, M. (2024) · 2024
Later among the works it cites.
The prompt report: a systematic survey of prompt engineering techniques
Schulhoff, S., Ilie, M., Balepur, N., Kahadze, K., Liu, A., Si, C., et al. (2024) · 2024
Later among the works it cites.
Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evaluation
Singh, S., Romanou, A., Fourrier, C., Adelani, D. I., Ngui, J. G., Vila-Suero, D., et al. (2024) · 2024
Later among the works it cites.
Evaluating the social impact of generative ai systems in systems and society
Solaiman, I., Talat, Z., Agnew, W., Ahmad, L., Baker, D., Blodgett, S. L., et al. (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Worldbench: Quantifying geographic disparities in llm factual recall
Moayeri, M., Tabassi, E., and Feizi, S. (2024) · 2024
Cited alongside, same era.
Later among the works it cites.
Benchmark data contamination of large language models: A survey
Xu, C., Guan, S., Greene, D., and Kechadi, M. (2024) · 2024
Later among the works it cites.