Fetching the paper…
Reading the bibliography…
This paper describes the DSBA submissions to the Prompting Large Language Models as Explainable Metrics shared task, where systems were submitted to two tracks: small and large summarization tracks.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
COMET: commonsense transformers for automatic knowledge graph construction
Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Celikyilmaz, and Yejin Choi. 2019 · 1906
Earlier work this paper cites.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 1909
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir R. Radev. 2020 · 2007
Earlier work this paper cites.
Perception score: A learned metric for open-ended text generation evaluation
Jing Gu, Qingyang Wu, and Zhou Yu. 2021 · 2021
Earlier work this paper cites.
Openmeva: A benchmark for evaluating open-ended story generation metrics
Jian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu, Wenbiao Ding, Xiaoxi Mao, Changjie Fan, and Minlie Huang. 2021 · 2021
Earlier work this paper cites.
Effects of semantic plausibility, syntactic complexity and n-gram frequency on children’s sentence repetition
Kamila Polišenská, Shula Chiat, Jakub Szewczyk, and Katherine E Twomey. 2021 · 2021
Earlier work this paper cites.
Exploring syntactic and semantic features for authorship attribution
Haiyan Wu, Zhiqiang Zhang, and Qingfeng Wu. 2021 · 2021
Cited alongside, same era.
A survey for in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022 · 2022
Cited alongside, same era.
Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022 · 2022
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung yi Lee. 2023 · 2023
Cited alongside, same era.
The eval4nlp 2023 shared task on prompting large language models as explainable metrics
Christoph Leiter, Juri Opitz, Daniel Deutsch, Yang Gao, Rotem Dror, and Steffen Eger. 2023 · 2023
Closest in time.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Closest in time.
Orca: Progressive learning from complex explanation traces of gpt-4
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. 2023 · 2023
Closest in time.
Approximating human evaluation of social chatbots with prompting
Ekaterina Svikhnushina and Pearl Pu. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, André F. T. Martins, Graham Neubig, Ankush Garg, Jonathan H. Clark, Markus Freitag, and Orhan Firat. 2023 · 2023
Cited alongside, same era.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Cited alongside, same era.
In-context learning of large language models explained as kernel regression
Chi Han, Ziqi Wang, Han Zhao, and Heng Ji. 2023 · 2023
Cited alongside, same era.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Cited alongside, same era.
Platypus: Quick, cheap, and powerful refinement of llms
Ariel N. Lee, Cole J. Hunter, and Nataniel Ruiz. 2023 · 2023
Cited alongside, same era.
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Closest in time.
Larger language models do in-context learning differently
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al. 2023 · 2023
Closest in time.
Instructscore: Towards explainable text generation evaluation with automatic feedback
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Yang Wang, and Lei Li. 2023 · 2023
Closest in time.
Multimodal chain-of-thought reasoning in language models
Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2023 · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Closest in time.