Fetching the paper…
Reading the bibliography…
This paper describes the systems submitted by team6 for ChatEval, the DSTC 11 Track 4 competition.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 1910
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
MPNet: Masked and Permuted Pre-training for Language Understanding
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
Spearman rank correlation
Jerrold H Zar. 2005 · 2005
Earlier work this paper cites.
Statistics (international student edition)
David Freedman, Robert Pisani, and Roger Purves. 2007 · 2007
Earlier work this paper cites.
DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Earlier work this paper cites.
Chia-Wei Liu, Ryan Lowe, Iulian V. Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. 2017 · 2017
Earlier work this paper cites.
Towards an Automatic Turing Test: Learning to Evaluate Dialogue Responses
Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Cited alongside, same era.
Learning discourse-level diversity for neural dialog models using conditional variational autoencoders
Tiancheng Zhao, Ran Zhao, and Maxine Eskenazi. 2017 · 2017
Cited alongside, same era.
Personalizing Dialogue Agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018 · 2018
Cited alongside, same era.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Cited alongside, same era.
Introducing Retrospectives: ’Real Talk’ for your Past Papers
Ryan Lowe. 2019 · 2019
Cited alongside, same era.
Deep am-fm: Toolkit for automatic dialogue evaluation
Chen Zhang, Luis D’Haro, Rafael Banchs, Thomas Friedrichs, and Haizhou Li. 2020 · 2020
Later among the works it cites.
Perturbation checklists for evaluating nlg evaluation metrics
Ananya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M. Khapra. 2021 · 2021
Later among the works it cites.
A comprehensive assessment of dialog evaluation metrics
Yi-Ting Yeh, Maxine Eskenazi, and Shikib Mehri. 2021 · 2021
Later among the works it cites.
Gpt-neox-20b: An open-source autoregressive language model
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022 · 2022
Later among the works it cites.
Statistical Properties of the log-cosh Loss Function Used in Machine Learning
Resve A. Saleh and A. K. Md Ehsanes Saleh. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Re-Evaluating ADEM: A Deeper Look at Scoring Dialogue Responses
Ananya B. Sai, Mithun Das Gupta, Mitesh M. Khapra, and Mukundhan Srinivasan. 2019 · 2019
Cited alongside, same era.
What makes a good conversation? How controllable attributes affect human judgments
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Dialogue response ranking training with large-scale human feedback data
Xiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett, and Bill Dolan. 2020 · 2020
Cited alongside, same era.
Learning an unreferenced metric for online dialogue evaluation
Koustuv Sinha, Prasanna Parthasarathi, Jasmine Wang, Ryan Lowe, William L. Hamilton, and Joelle Pineau. 2020 · 2020
Cited alongside, same era.
Unsupervised Evaluation of Interactive Dialog with DialoGPT
Shikib Mehri and Maxine Eskenazi. 2020a
Cited in the paper.
USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation
Shikib Mehri and Maxine Eskenazi. 2020b
Cited in the paper.
Later among the works it cites.
Benchmarking generalization via in-context instructions on 1,600+ language tasks
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, A. Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Maitreya Patel, Kuntal Kumar Pal, M. Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Shailaja Keyur Sampat, Savan Doshi, Siddharth Deepak Mishra, Sujan C. Reddy, Sumanta Patro, Tanay Dixit, Xu dong Shen, Chitta Baral, Yejin Choi, Hannaneh Hajishirzi, Noah A. Smith, and Daniel Khashabi. 2022 · 2022
Later among the works it cites.
Large Language Models Are State-of-the-Art Evaluators of Translation Quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Closest in time.
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Closest in time.
Mario Rodríguez-Cantelar, Chen Zhang, Chengguang Tang, Ke Shi, Sarik Ghazarian, João Sedoc, Luis Fernando D’Haro, and Alexander Rudnicky. 2023 · 2023
Closest in time.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Closest in time.