Fetching the paper…
Reading the bibliography…
While long-context large language models (LLMs) can technically summarize book-length documents (>100K tokens), the length and complexity of the documents have so far prohibited evaluations of input-dependent aspects like faithfulness.
Okapi at trec-3
Stephen E. Robertson, Steve Walker, Susan Jones, Micheline M. Hancock-Beaulieu, Mike Gatford, et al · 1995
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Ani Nenkova and Rebecca Passonneau · 2004
Earlier work this paper cites.
Bayesian summarization at duc and a suggestion for extrinsic evaluation
Hal Daumé and D. Marcu · 2005
Earlier work this paper cites.
Non-expert evaluation of summarization systems is risky
Dan Gillick and Yang Liu · 2010
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Evaluating factuality in generation with dependency-level entailment
Tanya Goyal and Greg Durrett · 2020
Earlier work this paper cites.
Evaluating the factual consistency of abstractive text summarization
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher · 2020
Earlier work this paper cites.
Exploring content selection in summarization of novel chapters
Faisal Ladhak, Bryan Li, Yaser Al-Onaizan, and Kathleen McKeown · 2020
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald · 2020
Earlier work this paper cites.
Fact or fiction: Verifying scientific claims
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi · 2020
Earlier work this paper cites.
Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization
Mengyao Cao, Yue Dong, and Jackie Chi Kit Cheung · 2021
Earlier work this paper cites.
Annotating and modeling fine-grained factuality in summarization, 2021
Tanya Goyal and Greg Durrett · 2021
Earlier work this paper cites.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov · 2021
Earlier work this paper cites.
Get your vitamin C! robust fact verification with contrastive evidence
Tal Schuster, Adam Fisch, and Regina Barzilay · 2021
Cited alongside, same era.
Recursively summarizing books with human feedback, 2021
Jeff Wu, Long Ouyang, Daniel M. Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano · 2021
Cited alongside, same era.
Attributed question answering: Evaluation and modeling for attributed large language models
Bernd Bohnet, Vinh Q Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Massimiliano Ciaramita, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, et al · 2022
Cited alongside, same era.
Summscreen: A dataset for abstractive screenplay summarization, 2022
Mingda Chen, Zewei Chu, Sam Wiseman, and Kevin Gimpel · 2022
Cited alongside, same era.
BOOKSUM: A collection of datasets for long-form narrative summarization
Wojciech Kryscinski, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, and Dragomir Radev · 2022
Cited alongside, same era.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Fine-tuning language models for factuality
Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning, and Chelsea Finn · 2023
Later among the works it cites.
Annotating and detecting fine-grained factual errors for dialogue summarization
Rongxin Zhu, Jianzhong Qi, and Jey Han Lau · 2023
Later among the works it cites.
Model Card: Claude 3
Anthropic · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SQuALITY: Building a long-document summarization dataset the hard way
Alex Wang, Richard Yuanzhe Pang, Angelica Chen, Jason Phang, and Samuel R. Bowman · 2022
Cited alongside, same era.
Speak, memory: An archaeology of books known to ChatGPT/GPT-4
Kent Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman · 2023
Cited alongside, same era.
Enabling large language models to generate text with citations
Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen · 2023
Cited alongside, same era.
Wice: Real-world entailment for claims in wikipedia, 2023
Ryo Kamoi, Tanya Goyal, Juan Diego Rodriguez, and Greg Durrett · 2023
Cited alongside, same era.
Needle in a haystack
Greg Kamradt · 2023
Cited alongside, same era.
LongEval: Guidelines for human evaluation of faithfulness in long-form summarization
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo · 2023
Cited alongside, same era.
Unveiling the essence of poetry: Introducing a comprehensive dataset and benchmark for poem summarization
Ridwan Mahbub, Ifrad Khan, Samiha Anuva, Md Shihab Shahriar, Md Tahmid Rahman Laskar, and Sabbir Ahmed · 2023
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
Google Gemini Team · 2024
Closest in time.
Mixtral of experts, 2024
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Closest in time.
Faithfulness in abstractive summarization: Progress and challenges, 2024
Faisal Ladhak · 2024
Closest in time.
Same task, more tokens: the impact of input length on the reasoning performance of large language models, 2024
Mosh Levy, Alon Jacoby, and Yoav Goldberg · 2024
Closest in time.
Fine-grained hallucination detection and editing for language models
Abhika Mishra, Akari Asai, Vidhisha Balachandran, Yizhong Wang, Graham Neubig, Yulia Tsvetkov, and Hannaneh Hajishirzi · 2024
Closest in time.
Reading subtext: Evaluating large language models on short story summarization with writers, 2024
Melanie Subbiah, Sean Zhang, Lydia B. Chilton, and Kathleen McKeown · 2024
Closest in time.
Tofueval: Evaluating hallucinations of llms on topic-focused dialogue summarization
Liyan Tang, Igor Shalyminov, Amy Wing-mei Wong, Jon Burnsky, Jake W Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, et al · 2024
Closest in time.
Long-form factuality in large language models, 2024
Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, and Quoc V. Le · 2024
Closest in time.