Fetching the paper…
Reading the bibliography…
We evaluate recent Large Language Models (LLMs) on the challenging task of summarizing short stories, which can be lengthy, and include nuanced subtext or scrambled timelines.
Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel
Peter Kincaid, Robert P. Fishburne, Richard L. Rogers, and Brad S. Chissom. 1975 · 1975
Earlier work this paper cites.
Remembrance of things parsed: Story structure and recall
Jean M Mandler and Nancy S Johnson. 1977 · 1977
Earlier work this paper cites.
Narrative discourse: An essay in method , volume 3
Gérard Genette. 1980 · 1980
Earlier work this paper cites.
The rhetoric of fiction
Wayne C. Booth. 1983 · 1983
Earlier work this paper cites.
Beloved. 1987
Toni Morrison. 2004 · 1987
Earlier work this paper cites.
Assessing narrative comprehension in young children
Alison H. Paris and Scott G. Paris. 2003 · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Basic elements of narrative
David Herman. 2009 · 2009
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi. 2018 · 2018
Earlier work this paper cites.
A web-based collaborative annotation and consolidation tool
Tobias Daudert. 2020 · 2020
Earlier work this paper cites.
Exploring content selection in summarization of novel chapters
Faisal Ladhak, Bryan Li, Yaser Al-Onaizan, and Kathleen McKeown. 2020 · 2020
Earlier work this paper cites.
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Earlier work this paper cites.
Narrative theory for computational narrative understanding
Andrew Piper, Richard Jean So, and David Bamman. 2021 · 2021
Earlier work this paper cites.
Recursively summarizing books with human feedback
Jeff Wu, Long Ouyang, Daniel M. Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021 · 2021
Earlier work this paper cites.
Help me write a Poem - Instruction Tuning as a Vehicle for Collaborative Poetry Writing"
Tuhin Chakrabarty, Vishakh Padmakumar, and He He. 2022 · 2022
Earlier work this paper cites.
SummScreen: A dataset for abstractive screenplay summarization
Mingda Chen, Zewei Chu, Sam Wiseman, and Kevin Gimpel. 2022 · 2022
Cited alongside, same era.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022 · 2022
Cited alongside, same era.
FALTE: A toolkit for fine-grained annotation for long text evaluation
Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022a · 2022
Cited alongside, same era.
SNaC: Coherence error detection for narrative summarization
Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022c · 2022
Cited alongside, same era.
The Black side of the river: Race, language, and belonging in Washington, DC
Jessica A. Grieser. 2022 · 2022
Cited alongside, same era.
Creative writing with an AI-powered writing assistant: Perspectives from professional writers
OpenAI. 2023 · 2023
Later among the works it cites.
Summarization is (almost) dead
Xiao Pu, Mingqi Gao, and Xiaojun Wan. 2023 · 2023
Later among the works it cites.
Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors
Liyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, and Greg Durrett. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
AlignScore: Evaluating factual consistency with a unified alignment function
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daphne Ippolito, Ann Yuan, Andy Coenen, and Sehmon Burnam. 2022 · 2022
Cited alongside, same era.
BOOKSUM: A collection of datasets for long-form narrative summarization
Wojciech Kryscinski, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, and Dragomir Radev. 2022 · 2022
Cited alongside, same era.
SQuALITY: Building a long-document summarization dataset the hard way
Alex Wang, Richard Yuanzhe Pang, Angelica Chen, Jason Phang, and Samuel R. Bowman. 2022 · 2022
Cited alongside, same era.
Fantastic questions and where to find them: FairytaleQA – an authentic dataset for narrative comprehension
Ying Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Li, Nora Bradford, Branda Sun, Tran Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, and Mark Warschauer. 2022 · 2022
Cited alongside, same era.
Wordcraft: Story writing with large language models
Ann Yuan, Andy Coenen, Emily Reif, and Daphne Ippolito. 2022 · 2022
Cited alongside, same era.
Experimental Narratives: A Comparison of Human Crowdsourced Storytelling and AI Storytelling
Nina Begus. 2023 · 2023
Cited alongside, same era.
Tuhin Chakrabarty, Vishakh Padmakumar, Faeze Brahman, and Smaranda Muresan. 2023 · 2023
Cited alongside, same era.
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. 2023 · 2023
Later among the works it cites.
Mug: A general meeting understanding and generation benchmark
Qinglin Zhang, Chong Deng, Jiaqing Liu, Hai Yu, Qian Chen, Wen Wang, Zhijie Yan, Jinglin Liu, Yi Ren, and Zhou Zhao. 2023 · 2023
Later among the works it cites.
Fiction-writing mode: An effective control for human-machine collaborative writing
Wenjie Zhong, Jason Naradowsky, Hiroya Takamura, Ichiro Kobayashi, and Yusuke Miyao. 2023 · 2023
Later among the works it cites.
Art or artifice? Large language models and the false promise of creativity
Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien-Sheng Wu. 2024 · 2024
Closest in time.
Booookscore: A systematic exploration of book-length summarization in the era of LLMs
Yapei Chang, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2024 · 2024
Closest in time.
A comprehensive evaluation of large language models on benchmark biomedical text processing tasks
Israt Jahan, Md Tahmid Rahman Laskar, Chun Peng, and Jimmy Xiangji Huang. 2024 · 2024
Closest in time.
Fables: Evaluating faithfulness and content selection in book-length summarization
Yekyung Kim, Yapei Chang, Marzena Karpinska, Aparna Garimella, Varun Manjunatha, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2024 · 2024
Closest in time.
Benchmarking generation and evaluation capabilities of large language models for instruction controllable summarization
Yixin Liu, Alexander Fabbri, Jiawen Chen, Yilun Zhao, Simeng Han, Shafiq Joty, Pengfei Liu, Dragomir Radev, Chien-Sheng Wu, and Arman Cohan. 2024 · 2024
Closest in time.
Does writing with language models reduce content diversity?
Vishakh Padmakumar and He He. 2024 · 2024
Closest in time.
Hainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du, and Yulan He. 2024 · 2024
Closest in time.
Catherine Yeh, Gonzalo Ramos, Rachel Ng, Andy Huntington, and Richard Banks. 2024 · 2024
Closest in time.
Benchmarking Large Language Models for News Summarization
Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B. Hashimoto. 2024 · 2024
Closest in time.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2038
Closest in time.