Fetching the paper…
Reading the bibliography…
In this paper, we uncover a systematic bias in the evaluation paradigm of adopting large language models~(LLMs), e.g., GPT-4, as a referee to score and compare the quality of responses generated by candidate models.
Behind the Scenes: An Exploration of Trigger Biases Problem in Few-Shot Event Classification
Wang, P.; Xun, R.; Liu, T.; Dai, D.; Chang, B.; and Sui, Z. 2021 · 1978
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
Interrater reliability: the kappa statistic
McHugh, M. L. 2012 · 2012
Earlier work this paper cites.
Do Supervised Distributional Methods Really Learn Lexical Inference Relations?
Levy, O.; Remus, S.; Biemann, C.; and Dagan, I. 2015 · 2015
Earlier work this paper cites.
Pay Attention to the Ending:Strong Neural Baselines for the ROC Story Cloze Task
Cai, Z.; Tu, L.; and Gimpel, K. 2017 · 2017
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
The Effect of Different Writing Tasks on Linguistic Style: A Case Study of the ROC Story Cloze Task
Schwartz, R.; Sap, M.; Konstas, I.; Zilles, L.; Choi, Y.; and Smith, N. A. 2017 · 2017
Earlier work this paper cites.
Annotation Artifacts in Natural Language Inference Data
Gururangan, S.; Swayamdipta, S.; Levy, O.; Schwartz, R.; Bowman, S.; and Smith, N. A. 2018 · 2018
Earlier work this paper cites.
Don’t Take the Premise for Granted: Mitigating Artifacts in Natural Language Inference
Belinkov, Y.; Poliak, A.; Shieber, S.; Van Durme, B.; and Rush, A. 2019 · 2019
Earlier work this paper cites.
Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
McCoy, T.; Pavlick, E.; and Linzen, T. 2019 · 2019
Earlier work this paper cites.
Compositional Questions Do Not Necessitate Multi-hop Reasoning
Min, S.; Wallace, E.; Singh, S.; Gardner, M.; Hajishirzi, H.; and Zettlemoyer, L. 2019 · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2020
Cited alongside, same era.
BERTScore: Evaluating Text Generation with BERT
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2020 · 2020
Cited alongside, same era.
BARTScore: Evaluating Generated Text as Text Generation
Yuan, W.; Neubig, G.; and Liu, P. 2021 · 2021
Cited alongside, same era.
PaLM: Scaling Language Modeling with Pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; Schuh, P.; Shi, K.; Tsvyashchenko, S.; Maynez, J.; Rao, A.; Barnes, P.; Tay, Y.; Shazeer, N. M.; Prabhakaran, V.; Reif, E.; Du, N.; Hutchinson, B. C.; Pope, R.; Bradbury, J.; Austin, J.; Isard, M.; Gur-Ari, G.; Yin, P.; Duke, T.; Levskaya, A.; Ghemawat, S.; Dev, S.; Michalewski, H.; García, X.; Misra, V.; Robinson, K.; Fedus, L.; Zhou, D.; Ippolito, D.; Luan, D.; Lim, H.; Zoph, B.; Spiridonov, A.; Sepassi, R.; Dohan, D.; Agrawal, S.; Omernick, M.; Dai, A. M.; Pillai, T. S.; Pellat, M.; Lewkowycz, A.; Moreira, E.; Child, R.; Polozov, O.; Lee, K.; Zhou, Z.; Wang, X.; Saeta, B.; Díaz, M.; Firat, O.; Catasta, M.; Wei, J.; Meier-Hellstern, K. S.; Eck, D.; Dean, J.; Petrov, S.; and Fiedel, N. 2022 · 2022
On the Blind Spots of Model-Based Evaluation Metrics for Text Generation
He, T.; Zhang, J.; Wang, T.; Kumar, S.; Cho, K.; Glass, J.; and Tsvetkov, Y. 2023 · 2023
Closest in time.
M3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
Li, L.; Yin, Y.; Li, S.; Chen, L.; Wang, P.; Ren, S.; Li, M.; Yang, Y.; Xu, J.; Sun, X.; et al. 2023 · 2023
Closest in time.
Lu, Q.; Qiu, B.; Ding, L.; Xie, L.; and Tao, D. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Peng, B.; Li, C.; He, P.; Galley, M.; and Gao, J. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A Survey for In-context Learning
Dong, Q.; Li, L.; Dai, D.; Zheng, C.; Wu, Z.; Chang, B.; Sun, X.; Xu, J.; and Sui, Z. 2022 · 2022
Cited alongside, same era.
Introducing ChatGPT
OpenAI. 2022 · 2022
Cited alongside, same era.
Learning Robust Representations for Continual Relation Extraction via Adversarial Class Augmentation
Wang, P.; Song, Y.; Liu, T.; Lin, B.; Cao, Y.; Li, S.; and Sui, Z. 2022 · 2022
Cited alongside, same era.
Eight things to know about large language models
Bowman, S. R. 2023 · 2023
Cited alongside, same era.
Human-in-the-Loop through Chain-of-Thought
Cai, Z.; Chang, B.; and Han, W. 2023 · 2023
Cited alongside, same era.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y.; Li, X.; Taori, R.; Zhang, T.; Gulrajani, I.; Ba, J.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Cited alongside, same era.
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Gao, P.; Han, J.; Zhang, R.; Lin, Z.; Geng, S.; Zhou, A.; Zhang, W.; Lu, P.; He, C.; Yue, X.; Li, H.; and Qiao, Y. J. 2023 · 2023
Cited alongside, same era.
HypoNLI: Exploring the Artificial Patterns of Hypothesis-only Bias in Natural Language Inference
Liu, T.; Xin, Z.; Chang, B.; and Sui, Z. 2020a
Cited in the paper.
Closest in time.
Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
Sun, Z.; Shen, Y.; Zhou, Q.; Zhang, H.; Chen, Z.; Cox, D. D.; Yang, Y.; and Gan, C. 2023 · 2023
Closest in time.
Turpin, M.; Michael, J.; Perez, E.; and Bowman, S. R. 2023 · 2023
Closest in time.
Enhancing Continual Relation Extraction via Classifier Decomposition
Xia, H.; Wang, P.; Liu, T.; Lin, B.; Cao, Y.; and Sui, Z. 2023 · 2023
Closest in time.
Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data
Xu, C.; Guo, D.; Duan, N.; and McAuley, J. 2023 · 2023
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023 · 2023
Closest in time.
LIMA: Less Is More for Alignment
Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; Yu, L.; Zhang, S.; Ghosh, G.; Lewis, M.; Zettlemoyer, L.; and Levy, O. 2023 · 2023
Closest in time.