Fetching the paper…
Reading the bibliography…
Attributing answers to source documents is an approach used to enhance the verifiability of a model's output in retrieval augmented generation (RAG).
Distilling dense representations for ranking using tightly-coupled teachers
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2020 · 2010
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X. Charles, D. Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson. 2013 · 2013
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, and 1 others. 2016 · 2016
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, and 1 others. 2019 · 2019
Earlier work this paper cites.
BERTScore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Earlier work this paper cites.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, and 1 others. 2020 · 2020
Earlier work this paper cites.
Reducing sentiment bias in language models via counterfactual evaluation
Po-Sen Huang, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, and Pushmeet Kohli. 2020 · 2020
Earlier work this paper cites.
Colbert: Efficient and effective passage search via contextualized late interaction over bert
Omar Khattab and Matei Zaharia. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, and 1 others. 2020 · 2020
Earlier work this paper cites.
Coil: Revisit exact lexical match in information retrieval with contextualized inverted list
Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021 · 2021
Earlier work this paper cites.
Jimmy Lin and Xueguang Ma. 2021 · 2021
Earlier work this paper cites.
KILT: A benchmark for knowledge intensive language tasks
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, and 1 others. 2021 · 2021
Earlier work this paper cites.
Robustness to spurious correlations in text classification via automatically generated counterfactuals
Zhao Wang and Aron Culotta. 2021 · 2021
Earlier work this paper cites.
Attributed question answering: Evaluation and modeling for attributed large language models
Bernd Bohnet, Vinh Q Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Massimiliano Ciaramita, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, and 1 others. 2022 · 2022
Earlier work this paper cites.
Teaching language models to support answers with verified quotes
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, and 1 others. 2022 · 2022
Earlier work this paper cites.
Self-rag: Learning to retrieve, generate, and critique through self-reflection
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023 · 2023
Earlier work this paper cites.
Factuality challenges in the era of large language models
Isabelle Augenstein, Timothy Baldwin, Meeyoung Cha, Tanmoy Chakraborty, Giovanni Luca Ciampaglia, David Corney, Renee DiResta, Emilio Ferrara, Scott Hale, Alon Halevy, and 1 others. 2023 · 2023
Earlier work this paper cites.
Hierarchical evaluation framework: Best practices for human evaluation
Iva Bojic, Jessica Chen, Si Yuan Chang, Qi Chwen Ong, Shafiq Joty, and Josip Car. 2023 · 2023
Earlier work this paper cites.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-Yi Lee. 2023 · 2023
Earlier work this paper cites.
ROBBIE: Robust bias evaluation of large generative language models
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, and Eric Smith. 2023 · 2023
Cited alongside, same era.
Enabling large language models to generate text with citations
Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023 · 2023
Cited alongside, same era.
Bias beyond English: Counterfactual tests for bias in sentiment analysis in four languages
Seraphina Goldfarb-Tarrant, Adam Lopez, Roi Blanco, and Diego Marcheggiani. 2023 · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, and 1 others. 2023 · 2023
Benchmarking large language models in complex question answering attribution using knowledge graphs
Nan Hu, Jiaoyan Chen, Yike Wu, Guilin Qi, Sheng Bi, Tongtong Wu, and Jeff Z. Pan. 2024 · 2024
Closest in time.
Citation: A key to building responsible and accountable large language models
Jie Huang and Kevin Chang. 2024 · 2024
Closest in time.
Adaptive-RAG: Learning to adapt retrieval-augmented large language models through question complexity
Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong C Park. 2024 · 2024
Closest in time.
Source-aware training enables knowledge attribution in language models
Muhammad Khalifa, David Wadden, Emma Strubell, Honglak Lee, Lu Wang, Iz Beltagy, and Hao Peng · 2024
Closest in time.
PlanRAG: A plan-then-retrieval augmented generation for generative large language models as decision makers
Myeonghwa Lee, Seonho An, and Min-Soo Kim. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
HAGRID: A human-LLM collaborative dataset for generative information-seeking with attribution
Ehsan Kamalloo, Aref Jafari, Xinyu Zhang, Nandan Thakur, and Jimmy Lin. 2023 · 2023
Cited alongside, same era.
Evaluating and modeling attribution for cross-lingual question answering
Benjamin Muller, John Wieting, Jonathan H Clark, Tom Kwiatkowski, Sebastian Ruder, Livio Soares, Roee Aharoni, Jonathan Herzig, and Xinyi Wang. 2023 · 2023
Cited alongside, same era.
GPT-4 technical report
OpenAI. 2023 · 2023
Cited alongside, same era.
The troubling emergence of hallucination in large language models-an extensive definition, quantification, and prescriptive remediations
Vipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, SM Towhidul Islam Tonmoy, Aman Chadha, Amit Sheth, and Amitava Das. 2023 · 2023
Cited alongside, same era.
Improving the domain adaptation of retrieval augmented generation (RAG) models for open domain question answering
Shamane Siriwardhana, Rivindu Weerasekera, Elliott Wen, Tharindu Kaluarachchi, Rajib Rana, and Suranga Nanayakkara. 2023 · 2023
Cited alongside, same era.
Learning to filter context for retrieval-augmented generation
Zhiruo Wang, Jun Araki, Zhengbao Jiang, Md Rizwan Parvez, and Graham Neubig. 2023 · 2023
Cited alongside, same era.
Counter-gap: Counterfactual bias evaluation through gendered ambiguous pronouns
Zhongbin Xie, Vid Kocijan, Thomas Lukasiewicz, and Oana-Maria Camburu. 2023 · 2023
Cited alongside, same era.
TRAQ: Trustworthy retrieval augmented question answering via conformal prediction
Shuo Li, Sangdon Park, Insup Lee, and Osbert Bastani. 2024a · 2024
Closest in time.
Towards verifiable generation: A benchmark for knowledge-aware language model attribution
Xinze Li, Yixin Cao, Liangming Pan, Yubo Ma, and Aixin Sun. 2024b · 2024
Closest in time.
AttributionBench: How hard is automatic attribution evaluation?
Yifei Li, Xiang Yue, Zeyi Liao, and Huan Sun. 2024c · 2024
Closest in time.
ExpertQA: Expert-curated questions and attributed answers
Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth. 2024 · 2024
Closest in time.
Exploring reasoning biases in large language models through syllogism: Insights from the NeuBAROCO dataset
Kentaro Ozeki, Risako Ando, Takanobu Morishita, Hirohiko Abe, Koji Mineshima, and Mitsuhiro Okada. 2024 · 2024
Closest in time.
Towards improved multi-source attribution for long-form answer generation
Nilay Patel, Shivashankar Subramanian, Siddhant Garg, Pratyay Banerjee, and Amita Misra. 2024 · 2024
Closest in time.
Bergen: A benchmarking library for retrieval-augmented generation
David Rau, Hervé Déjean, Nadezhda Chirkova, Thibault Formal, Shuai Wang, Vassilina Nikoulina, and Stéphane Clinchant. 2024 · 2024
Closest in time.
Groundedness in retrieval-augmented long-form generation: An empirical study
Alessandro Stolfo. 2024 · 2024
Closest in time.
Hexiang Tan, Fei Sun, Wanli Yang, Yuanzhuo Wang, Qi Cao, and Xueqi Cheng. 2024 · 2024
Closest in time.
Decodingtrust: A comprehensive assessment of trustworthiness in GPT models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, and 1 others. 2024 · 2024
Closest in time.
Benchmarking retrieval-augmented generation for medicine
Guangzhi Xiong, Qiao Jin, Zhiyong Lu, and Aidong Zhang. 2024 · 2024
Closest in time.
Effective large language model adaptation for improved grounding and citation generation
Xi Ye, Ruoxi Sun, Sercan Arik, and Tomas Pfister. 2024 · 2024
Closest in time.
Evidence-driven retrieval augmented response generation for online misinformation
Zhenrui Yue, Huimin Zeng, Yimeng Lu, Lanyu Shang, Yang Zhang, and Dong Wang. 2024 · 2024
Closest in time.
Measuring and addressing indexical bias in information retrieval
Caleb Ziems, William Held, Jane Dwivedi-Yu, and Diyi Yang. 2024 · 2024
Closest in time.