Fetching the paper…
Reading the bibliography…
Can ChatGPT provide evidence to support its answers? Does the evidence it suggests actually exist and does it really support its answer? We investigate these questions using a collection of domain-specific knowledge-based questions, specifically prompting ChatGPT to provide both an answer and supporting evidence in the form of references to external sources.
Concerns over use of glyphosate-based herbicides and risks associated with exposures: a consensus statement
John Peterson Myers, Michael N Antoniou, Bruce Blumberg, Lynn Carroll, Theo Colborn, Lorne G Everett, Michael Hansen, Philip J Landrigan, Bruce P Lanphear, Robin Mesnage, et al · 2016
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021 · 2021
Earlier work this paper cites.
Rethinking Search: Making Domain Experts out of Dilettantes
Donald Metzler, Yi Tay, Dara Bahri, and Marc Najork. 2021 · 2021
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Earlier work this paper cites.
Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models
Bernd Bohnet, Vinh Q Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, Kai Hui, et al · 2022
Earlier work this paper cites.
RARR: Researching and Revising What Language Models Say, Using Language Models
Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Y Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, et al · 2022
Earlier work this paper cites.
Teaching language models to support answers with verified quotes
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, et al · 2022
Earlier work this paper cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Earlier work this paper cites.
Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency . 214–229
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al · 2022
Cited alongside, same era.
Enabling Large Language Models to Generate Text with Citations
Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023 · 2023
Cited alongside, same era.
ChatGPT is not all you need. A State of the Art Review of large Generative AI models
Roberto Gozalo-Brizuela and Eduardo C Garrido-Merchan. 2023 · 2023
Cited alongside, same era.
Learning to fake it: limited responses and fabricated references provided by ChatGPT for medical questions
Jocelyn Gravel, Madeleine D’Amours-Gravel, and Esli Osmanlliu. 2023 · 2023
Cited alongside, same era.
How close is chatgpt to human experts? comparison corpus, evaluation, and detection
AgAsk: an agent to help answer farmer’s questions from scientific documents
Bevan Koopman, Ahmed Mourad, Hang Li, Anton van der Vegt, Shengyao Zhuang, Simon Gibson, Yash Dang, David Lawrence, and Guido Zuccon. 2023 · 2023
Closest in time.
Reham Omar, Omij Mangukiya, Panos Kalnis, and Essam Mansour. 2023 · 2023
Closest in time.
ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. In Healthcare , Vol. 11. MDPI, 887
Malik Sallam. 2023 · 2023
Closest in time.
Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agent
Weiwei Sun, Lingyong Yan, Xinyu Ma, Pengjie Ren, Dawei Yin, and Zhaochun Ren. 2023 · 2023
Closest in time.
Evaluation of ChatGPT as a question answering system for answering complex questions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023 · 2023
Cited alongside, same era.
Quality of citation data using the natural language processing tool ChatGPT in rheumatology: creation of false references
Axel J Hueber and Arnd Kleyer. 2023 · 2023
Cited alongside, same era.
Yunjie Ji, Yan Gong, Yiping Peng, Chao Ni, Peiyan Sun, Dongyu Pan, Baochang Ma, and Xiangang Li. 2023a · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023b · 2023
Cited alongside, same era.
Measuring attribution in natural language generation models
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. [n.d.]
Cited in the paper.
Yiming Tan, Dehai Min, Yu Li, Wenbo Li, Nan Hu, Yongrui Chen, and Guilin Qi. 2023 · 2023
Closest in time.
Can ChatGPT write a good boolean query for systematic review literature search?
Shuai Wang, Harrisen Scells, Bevan Koopman, and Guido Zuccon. 2023 · 2023
Closest in time.
Dr ChatGPT, tell me what I want to hear: How prompt knowledge impacts health answer correctness
Guido Zuccon and Bevan Koopman. 2023 · 2023
Closest in time.