Fetching the paper…
Reading the bibliography…
How can we construct an automated debate judge to evaluate an extensive, vibrant, multi-turn debate? This task is challenging, as judging a debate involves grappling with lengthy texts, intricate argument relationships, and multi-dimensional assessments.
Modeling argument strength in student essays
Isaac Persing and Vincent Ng. 2015 · 2015
Earlier work this paper cites.
Providing arguments in discussions on the basis of the prediction of human argumentative behavior
Ariel Rosenfeld and Sarit Kraus. 2016 · 2016
Earlier work this paper cites.
Conversational flow in Oxford-style debates
Justine Zhang, Ravi Kumar, Sujith Ravi, and Cristian Danescu-Niculescu-Mizil. 2016 · 2016
Earlier work this paper cites.
Towards debate automation: a recurrent model for predicting debate winners
Peter Potash and Anna Rumshisky. 2017 · 2017
Earlier work this paper cites.
Adapting healthy eating messages to personality
Rosemary Josekutty Thomas, Judith Masthoff, and Nir Oren. 2017 · 2017
Earlier work this paper cites.
Are you convinced? choosing the more convincing evidence with a Siamese network
Martin Gleize, Eyal Shnarch, Leshem Choshen, Lena Dankin, Guy Moshkowich, Ranit Aharonov, and Noam Slonim. 2019 · 2019
Earlier work this paper cites.
Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons
Margaret Li, Jason Weston, and Stephen Roller. 2019 · 2019
Earlier work this paper cites.
Is argumessage effective? a critical evaluation of the persuasive message generation system
Rosemary Josekutty Thomas, Judith Masthoff, and Nir Oren. 2019b · 2019
Earlier work this paper cites.
Exploiting personal characteristics of debaters for predicting persuasiveness
Khalid Al Khatib, Michael Völske, Shahbaz Syed, Nikolay Kolyada, and Benno Stein. 2020 · 2020
Earlier work this paper cites.
Look at the first sentence: Position bias in question answering
Miyoung Ko, Jinhyuk Lee, Hyunjae Kim, Gangwoo Kim, and Jaewoo Kang. 2020 · 2020
Earlier work this paper cites.
Vivesdebate: A new annotated multilingual corpus of argumentation in a debate tournament
Ramon Ruiz-Dolz, Montserrat Nofre, Mariona Taulé, Stella Heras, and Ana García-Fornes. 2021 · 2021
Cited alongside, same era.
An autonomous debating system
Noam Slonim, Yonatan Bilu, Carlos Alzate, Roy Bar-Haim, Ben Bogin, Francesca Bonin, Leshem Choshen, Edo Cohen-Karlik, Lena Dankin, Lilach Edelstein, Liat Ein-Dor, Roni Friedman-Melamed, Assaf Gavron, Ariel Gera, Martin Gleize, Shai Gretz, Dan Gutfreund, Alon Halfon, Daniel Hershcovich, Ron Hoory, Yufang Hou, Shay Hummel, Michal Jacovi, Charles Jochim, Yoav Kantor, Yoav Katz, David Konopnicki, Zvi Kons, Lili Kotlerman, Dalia Krieger, Dan Lahav, Tamar Lavee, Ran Levy, Naftali Liberman, Yosi Mass, Amir Menczel, Shachar Mirkin, Guy Moshkowich, Shila Ofek-Koifman, Matan Orbach, Ella Rabinovich, Ruty Rinott, Slava Shechtman, Dafna Sheinwald, Eyal Shnarch, Ilya Shnayderman, Aya Soffer, Artem Spector, Benjamin Sznajder, Assaf Toledo, Orith Toledo-Ronen, Elad Venezian, and Ranit Aharonov. 2021 · 2021
Cited alongside, same era.
Machine learning for utility prediction in argument-based computational persuasion
Ivan Donadello, Anthony Hunter, Stefano Teso, and Mauro Dragoni. 2022 · 2022
Cited alongside, same era.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2022
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Later among the works it cites.
Prd: Peer rank and discussion improve large language model based evaluations
Ruosen Li, Teerth Patel, and Xinya Du. 2023 · 2023
Later among the works it cites.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI. 2023 · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine Mcleavey, and Ilya Sutskever. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Benchmarking foundation models with language-model-as-an-examiner
Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, and Lei Hou. 2023 · 2023
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Cited alongside, same era.
Prompting large language models with the socratic method
Edward Y. Chang. 2023 · 2023
Cited alongside, same era.
Exploring the potential of large language models in computational argumentation
Guizhen Chen, Liying Cheng, Luu Anh Tuan, and Lidong Bing. 2023 · 2023
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Cited alongside, same era.
I wish to have an argument: Argumentative reasoning in large language models
Adrian de Wynter and Tommy Yuan. 2023 · 2023
Cited alongside, same era.
Can i influence you? development of a scale to measure perceived persuasiveness and two studies showing the use of the scale
Rosemary J. Thomas, Judith Masthoff, and Nir Oren. 2019a
Cited in the paper.
Is chatgpt a good nlg evaluator? a preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023a
Cited in the paper.
Automatic debate evaluation with argumentation semantics and natural language argument graph networks
Ramon Ruiz-Dolz, Stella Heras, and Ana Garcia. 2023 · 2023
Later among the works it cites.
World Universities Debating Championships Debating & Judging Manual
World Universities Debating Council. 2023 · 2023
Later among the works it cites.
Large language models are diverse role-players for summarization evaluation
Ning Wu, Ming Gong, Linjun Shou, Shining Liang, and Daxin Jiang. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
Judgelm: Fine-tuned large language models are scalable judges
Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2023 · 2023
Later among the works it cites.