Fetching the paper…
Reading the bibliography…
Plain language summarization with LLMs can be useful for improving textual accessibility of technical content.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Derivation of new readability formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for navy enlisted personnel
J Peter Kincaid, Robert P. Fishburne Jr., Richard L. Rogers, and Brad S. Chissom. 1975 · 1975
Earlier work this paper cites.
The well-built clinical question: a key to evidence-based decisions
W S Richardson, M C Wilson, J Nishikawa, and R S Hayward. 1995 · 1995
Earlier work this paper cites.
Evidence based medicine
David L Sackett, William MC Rosenberg, JA Muir Gray, R Brian Haynes, and W Scott Richardson. 1996 · 1996
Earlier work this paper cites.
Health literacy: report of the council on scientific affairs
Ad Hoc Committee on Health Literacy. 1999 · 1999
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Free-marginal multirater kappa (multirater k [free]): An alternative to fleiss’ fixed-marginal multirater kappa
Justus J Randolph. 2005 · 2005
Earlier work this paper cites.
Inter-coder agreement for computational linguistics
Ron Artstein and Massimo Poesio. 2008 · 2008
Earlier work this paper cites.
Seventy-five trials and eleven systematic reviews a day: how will we ever keep up?
Hilda Bastian, Paul Glasziou, and Iain Chalmers. 2010 · 2010
Earlier work this paper cites.
Fast and accurate prediction of sentence specificity
Junyi Jessy Li and Ani Nenkova. 2015 · 2015
Earlier work this paper cites.
Randomised controlled trials—the gold standard for effectiveness research
Eduardo Hariton and Joseph J Locascio. 2018 · 2018
Earlier work this paper cites.
Inferring which medical treatments work from reports of clinical trials
Eric Lehman, Jay DeYoung, Regina Barzilay, and Byron C. Wallace. 2019 · 2019
Earlier work this paper cites.
MOCHA: A dataset for training and evaluating generative reading comprehension metrics
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2020 · 2020
Earlier work this paper cites.
Evidence inference 2.0: More data, better models
Jay DeYoung, Eric Lehman, Benjamin Nye, Iain Marshall, and Byron C. Wallace. 2020 · 2020
Earlier work this paper cites.
Trialstreamer: Mapping and browsing medical evidence in real-time
Benjamin Nye, Ani Nenkova, Iain Marshall, and Byron C. Wallace. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Annotating and modeling fine-grained factuality in summarization
Tanya Goyal and Greg Durrett. 2021 · 2021
Cited alongside, same era.
State of the evidence: a survey of global disparities in clinical trials
Iain James Marshall, Veline L’Esperance, Rachel Marshall, James Thomas, Anna Noel-Storr, Frank Soboczenski, Benjamin Nye, Ani Nenkova, and Byron C Wallace. 2021 · 2021
Cited alongside, same era.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021 · 2021
Cited alongside, same era.
Questeval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Patrick Gallinari, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, and Alex Wang. 2021 · 2021
Cited alongside, same era.
Elaborative simplification: Content addition and explanation generation in text simplification
Neha Srikanth and Junyi Jessy Li. 2021 · 2021
Cited alongside, same era.
Thresh: A unified, customizable and deployable platform for fine-grained text evaluation
David Heineman, Yao Dou, and Wei Xu. 2023 · 2023
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Later among the works it cites.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Later among the works it cites.
Chatgpt as a factual inconsistency evaluator for text summarization
Zheheng Luo, Qianqian Xie, and Sophia Ananiadou. 2023 · 2023
Later among the works it cites.
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Digital health literacy as a predictor of awareness, engagement, and use of a national web-based personal health record: Population-based survey study
Christina Cheng, Emma Gearon, Melanie Hawkins, Crystal McPhee, Lisa Hanna, Roy Batterham, and Richard H Osborne. 2022 · 2022
Cited alongside, same era.
Evaluating factuality in text simplification
Ashwin Devaraj, William Sheffield, Byron Wallace, and Junyi Jessy Li. 2022 · 2022
Cited alongside, same era.
QAFactEval: Improved QA-based factual consistency evaluation for summarization
Alexander Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022 · 2022
Cited alongside, same era.
News summarization and evaluation in the era of GPT-3
Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022 · 2022
Cited alongside, same era.
TRUE: Re-evaluating factual consistency evaluation
Or Honovich, Roee Aharoni, Jonathan Herzig, Hagai Taitelbaum, Doron Kukliansy, Vered Cohen, Thomas Scialom, Idan Szpektor, Avinatan Hassidim, and Yossi Matias. 2022 · 2022
Cited alongside, same era.
SummaC: Re-visiting NLI-based models for inconsistency detection in summarization
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022 · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. 2022 · 2022
Cited alongside, same era.
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Med-halt: Medical domain hallucination test for large language models
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2023 · 2023
Later among the works it cites.
Summarizing, simplifying, and synthesizing medical evidence using GPT-3 (with varying success)
Chantal Shaib, Millicent Li, Sebastian Joseph, Iain Marshall, Junyi Jessy Li, and Byron Wallace. 2023 · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Later among the works it cites.
The effects of online access to general practice medical records perceived by patients: Longitudinal survey study
Rosa R L C Thielmann, Ciska Hoving, Jochen W L Cals, and Rik Crutzen. 2023 · 2023
Later among the works it cites.
Fine-tuning language models for factuality
Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher Manning, and Chelsea Finn. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
Is ChatGPT a good NLG evaluator? A preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023 · 2023
Later among the works it cites.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. 2024 · 2024
Closest in time.
InfoLossQA: Characterizing and recovering information loss in text simplification
Jan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert, Kyle Lo, Wei Xu, Byron C. Wallace, and Junyi Jessy Li. 2024 · 2024
Closest in time.