Fetching the paper…
Reading the bibliography…
Summarizing book-length documents (>100K tokens) that exceed the context window size of large language models (LLMs) requires first breaking the input document into smaller chunks and then prompting an LLM to merge, update, and compress chunk-level summaries.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
BillSum: A corpus for automatic summarization of US legislation
Anastassia Kornilova and Vladimir Eidelman · 2019
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev · 2020
Earlier work this paper cites.
SUPERT: Towards new frontiers in unsupervised evaluation metrics for multi-document summarization
Yang Gao, Wei Zhao, and Steffen Eger · 2020
Earlier work this paper cites.
Inquisitive question generation for high level text comprehension
Wei-Jen Ko, Te-yuan Chen, Yiyan Huang, Greg Durrett, and Junyi Jessy Li · 2020
Earlier work this paper cites.
Fill in the blanc: Human-free quality estimation of document summaries, 2020
Oleg Vasilyev, Vedant Dharnidharka, and John Bohannon · 2020
Earlier work this paper cites.
Reducing quantity hallucinations in abstractive summarization
Zheng Zhao, Shay B. Cohen, and Bonnie Webber · 2020
Earlier work this paper cites.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey · 2021
Earlier work this paper cites.
Annotating and modeling fine-grained factuality in summarization
Tanya Goyal and Greg Durrett · 2021
Earlier work this paper cites.
Recursively summarizing books with human feedback, 2021
Jeff Wu, Long Ouyang, Daniel M. Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano · 2021
Earlier work this paper cites.
Is gpt-3 text indistinguishable from human text? scarecrow: A framework for scrutinizing machine text, 2022
Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith, and Yejin Choi · 2022
Cited alongside, same era.
BOOKSUM: A collection of datasets for long-form narrative summarization
Wojciech Kryscinski, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, and Dragomir Radev · 2022
Cited alongside, same era.
SQuALITY: Building a long-document summarization dataset the hard way
Alex Wang, Richard Yuanzhe Pang, Angelica Chen, Jason Phang, and Samuel R. Bowman · 2022
Cited alongside, same era.
Adapting pretrained text-to-text models for long text sequences, 2022
Wenhan Xiong, Anchit Gupta, Shubham Toshniwal, Yashar Mehdad, and Wen tau Yih · 2022
Cited alongside, same era.
From sparse to dense: Gpt-4 summarization with chain of density prompting, 2023
Griffin Adams, Alexander Fabbri, Faisal Ladhak, Eric Lehman, and Noémie Elhadad · 2023
Cited alongside, same era.
Longeval: Guidelines for human evaluation of faithfulness in long-form summarization
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo · 2023
Closest in time.
Revisiting the gold standard: Grounding summarization evaluation with robust human evaluation
Yixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir Radev · 2023
Closest in time.
FollowupQG: Towards information-seeking follow-up question generation
Yan Meng, Liangming Pan, Yixin Cao, and Min-Yen Kan · 2023
Closest in time.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation, 2023
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2023
Closest in time.
A question answering framework for decontextualizing user-facing snippets from scientific documents
Benjamin Newman, Luca Soldaini, Raymond Fok, Arman Cohan, and Kyle Lo · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Awesome: Gpu memory-constrained long document summarization using memory mechanism and global salient content, 2023
Shuyang Cao and Lu Wang · 2023
Cited alongside, same era.
Speak, memory: An archaeology of books known to chatgpt/gpt-4, 2023
Kent K. Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman · 2023
Cited alongside, same era.
Alpacafarm: A simulation framework for methods that learn from human feedback, 2023
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Cited alongside, same era.
The devil is in the errors: Leveraging large language models for fine-grained machine translation evaluation, 2023
Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, André F. T. Martins, Graham Neubig, Ankush Garg, Jonathan H. Clark, Markus Freitag, and Orhan Firat · 2023
Cited alongside, same era.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Cited alongside, same era.
Snac: Coherence error detection for narrative summarization, 2022a
Tanya Goyal, Junyi Jessy Li, and Greg Durrett
Cited in the paper.
News Summarization and Evaluation in the Era of GPT-3
Tanya Goyal, Junyi Jessy Li, and Greg Durrett
Cited in the paper.
Long document summarization with top-down and bottom-up inference
Bo Pang, Erik Nijkamp, Wojciech Kryściński, Silvio Savarese, Yingbo Zhou, and Caiming Xiong · 2023
Closest in time.
Is chatgpt a good nlg evaluator? a preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou · 2023
Closest in time.
Elaborative simplification as implicit questions under discussion, 2023
Yating Wu, William Sheffield, Kyle Mahowald, and Junyi Jessy Li · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Closest in time.
A discourse-aware attention model for abstractive summarization of long documents
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian · 2097
Closest in time.