Fetching the paper…
Reading the bibliography…
In this paper, we investigate the phenomena of "selection biases" in Large Language Models (LLMs), focusing on problems where models are tasked with choosing the optimal option from an ordered sequence.
MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Amini, A., Gabriel, S., Lin, S., Koncel-Kedziorski, R., Choi, Y., and Hajishirzi, H. (2019) · 1905
Earlier work this paper cites.
WINOGRANDE: An Adversarial Winograd Schema Challenge at Scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. (2019) · 1907
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. (2018) · 2018
Earlier work this paper cites.
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. (2018) · 2018
Earlier work this paper cites.
HellaSwag: Can a Machine Really Finish Your Sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 2019
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2021) · 2021
Earlier work this paper cites.
Calibrate Before Use: Improving Few-Shot Performance of Language Models
Zhao, T., Wallace, E., Feng, S., Klein, D., and Singh, S. (2021) · 2021
Earlier work this paper cites.
Introducing ChatGPT
OpenAI (2022) · 2022
Cited alongside, same era.
GPT-4 Technical Report
Achiam, R. (2023) · 2023
Cited alongside, same era.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Cited alongside, same era.
Open LLM Leaderboard
Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T. (2023) · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., et al. (2023) · 2023
Cited alongside, same era.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. (2023a) · 2023
Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization
Shen, C., Cheng, L., Nguyen, X.-P., You, Y., and Bing, L. (2023) · 2023
Later among the works it cites.
Can Large Language Models Be an Alternative to Human Evaluations?
Chiang, C.-H. and Lee, H.-Y. (2023) · 2023
Later among the works it cites.
ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
Chan, C.-M., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., and Liu, Z. (2023) · 2023
Later among the works it cites.
Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations
Si, C., Friedman, D., Joshi, N., Feng, S., Chen, D., and He, H. (2023) · 2023
Later among the works it cites.
Mitigating Label Biases for In-context Learning
Fei, Y., Hou, Y., Chen, Z., and Bosselut, A. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions
Pezeshkpour, P. and Hruschka, E. (2023) · 2023
Cited alongside, same era.
Gemini: A Family of Highly Capable Multimodal Models
Anil, R. (2023a)
Cited in the paper.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. (2023b)
Cited in the paper.
Large Language Models are not Fair Evaluators
Wang, P., Li, L., Chen, L., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., and Sui, Z. (2023a)
Cited in the paper.
Large Language Models Are Not Robust Multiple Choice Selectors
Zheng, C., Zhou, H., Meng, F., Zhou, J., and Huang, M. (2023b)
Cited in the paper.
Adversarial Demonstration Attacks on Large Language Models
Wang, J., Liu, Z.-Y., Park, K. H., Chen, M., and Xiao, C. (2023b)
Cited in the paper.
Zhu, K., Wang, J., Zhou, J., Wang, Z., Chen, H., Wang, Y., Yang, L., Ye, W., Gong, N. Z., Zhang, Y., and Xie, X. (2023) · 2023
Later among the works it cites.
Leveraging Large Language Models for Multiple Choice Question Answering
Robinson, J. and Wingate, D. (2023) · 2023
Later among the works it cites.