Fetching the paper…
Reading the bibliography…
Numerous evaluation metrics have been developed for natural language generation tasks, but their effectiveness in evaluating stories is limited as they are not specifically tailored to assess intricate aspects of storytelling, such as fluency and interestingness.
A new measure of rank correlation
M. G. Kendall. 1938 · 1938
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Story cloze evaluator: Vector space representation evaluation by predicting what happens next
Nasrin Mostafazadeh, Lucy Vanderwende, Wen-tau Yih, Pushmeet Kohli, and James Allen. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
A knowledge-enhanced pretraining model for commonsense story generation
Jian Guan, Fei Huang, Zhihao Zhao, Xiaoyan Zhu, and Minlie Huang. 2020 · 2020
Earlier work this paper cites.
UNION: An Unreferenced Metric for Evaluating Open-ended Story Generation
Jian Guan and Minlie Huang. 2020 · 2020
Earlier work this paper cites.
How furiously can colorless green ideas sleep? sentence acceptability in context
Jey Han Lau, Carlos Armendariz, Shalom Lappin, Matthew Purver, and Chang Shu. 2020 · 2020
Earlier work this paper cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Earlier work this paper cites.
BERT-ATTACK: Adversarial attack against BERT using BERT
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020 · 2020
Earlier work this paper cites.
TextAttack: A framework for adversarial attacks, data augmentation, and adversarial training in NLP
John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post. 2020 · 2020
Cited alongside, same era.
MEGATRON-CNTRL: Controllable story generation with external knowledge using large-scale language models
Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Raul Puri, Pascale Fung, Anima Anandkumar, and Bryan Catanzaro. 2020 · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
Compression, transduction, and creation: A unified framework for evaluating natural language generation
Mingkai Deng, Bowen Tan, Zhengzhong Liu, Eric Xing, and Zhiting Hu. 2021 · 2021
Cited alongside, same era.
On the blind spots of model-based evaluation metrics for text generation
Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James R. Glass, and Yulia Tsvetkov. 2022 · 2022
Later among the works it cites.
DEMETR: Diagnosing evaluation metrics for translation
Marzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song, Ankita Gupta, and Mohit Iyyer. 2022 · 2022
Later among the works it cites.
T5score: Discriminative fine-tuning of generative evaluation metrics
Yiwei Qin, Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2022 · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Elizabeth-Jane Pavlick, Suzana Ili’c, Daniel Hesslow, Roman Castagn’e, Alexandra Sasha Luccioni, Franccois Yvon, and et al Matthias Gallé. 2022 · 2022
Later among the works it cites.
Re3: Generating longer stories with recursive reprompting and revision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Plot-guided adversarial example construction for evaluating open-domain story generation
Sarik Ghazarian, Zixi Liu, Akash S M, Ralph Weischedel, Aram Galstyan, and Nanyun Peng. 2021 · 2021
Cited alongside, same era.
The perils of using Mechanical Turk to evaluate open-ended text generation
Marzena Karpinska, Nader Akoury, and Mohit Iyyer. 2021 · 2021
Cited alongside, same era.
Perturbation CheckLists for evaluating NLG evaluation metrics
Ananya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M. Khapra. 2021 · 2021
Cited alongside, same era.
Progressive generation of long text with pretrained language models
Bowen Tan, Zichao Yang, Maruan Al-Shedivat, Eric Xing, and Zhiting Hu. 2021 · 2021
Cited alongside, same era.
Exploring story generation with multi-task objectives in variational autoencoders
Zhuohan Xie, Jey Han Lau, and Trevor Cohn. 2021 · 2021
Cited alongside, same era.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Cited alongside, same era.
StoryER: Automatic story evaluation via ranking, rating and reasoning
Hong Chen, Duc Vo, Hiroya Takamura, Yusuke Miyao, and Hideki Nakayama. 2022 · 2022
Cited alongside, same era.
Kevin Yang, Yuandong Tian, Nanyun Peng, and Dan Klein. 2022 · 2022
Later among the works it cites.
Persona-guided planning for controlling the protagonist’s persona in story generation
Zhexin Zhang, Jiaxin Wen, Jian Guan, and Minlie Huang. 2022b · 2022
Later among the works it cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Closest in time.
Can large language models be an alternative to human evaluations?
David Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Closest in time.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Closest in time.
G-eval: NLG evaluation using GPT-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Closest in time.
A survey of evaluation metrics used for NLG systems
Ananya B. Sai, Akash Kumar Mohankumar, and Mitesh M. Khapra. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Closest in time.
The next chapter: A study of large language models in storytelling
Zhuohan Xie, Trevor Cohn, and Jey Han Lau. 2023 · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023 · 2023
Closest in time.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2038
Closest in time.