Fetching the paper…
Reading the bibliography…
Since large language models (LLMs) achieve significant success in recent years, the hallucination issue remains a challenge, numerous benchmarks are proposed to detect the hallucination.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019 · 1905
Earlier work this paper cites.
Representing discourse coherence: A corpus-based analysis
Florian Wolf and Edward Gibson. 2004 · 2004
Earlier work this paper cites.
Free-marginal multirater kappa (multirater k [free]): An alternative to fleiss’ fixed-marginal multirater kappa
Justus J Randolph. 2005 · 2005
Earlier work this paper cites.
The dialog state tracking challenge
Jason D. Williams, Antoine Raux, Deepak Ramachandran, and Alan W. Black. 2013 · 2013
Earlier work this paper cites.
Inter-annotator agreement
Ron Artstein. 2017 · 2017
Earlier work this paper cites.
A survey on dialogue systems: Recent advances and new frontiers
Hongshen Chen, Xiaorui Liu, Dawei Yin, and Jiliang Tang. 2017 · 2017
Earlier work this paper cites.
A network-based end-to-end trainable task-oriented dialogue system
Tsung-Hsien Wen, David Vandyke, Nikola Mrksic, Milica Gasic, Lina Maria Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve J. Young. 2017 · 2017
Earlier work this paper cites.
Multiwoz - A large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Pawel Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gasic. 2018 · 2018
Earlier work this paper cites.
A knowledge-grounded neural conversation model
Marjan Ghazvininejad, Chris Brockett, Ming-Wei Chang, Bill Dolan, Jianfeng Gao, Wen-tau Yih, and Michel Galley. 2018 · 2018
Earlier work this paper cites.
The web as a knowledge-base for answering complex questions
Alon Talmor and Jonathan Berant. 2018 · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller. 2019 · 2019
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Survey on evaluation methods for dialogue systems
Jan Deriu, Álvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak. 2021 · 2021
Earlier work this paper cites.
Do we know what we don’t know? studying unanswerable questions beyond squad 2.0
Elior Sulem, Jamaal Hay, and Dan Roth. 2021 · 2021
Earlier work this paper cites.
Adding chit-chat to enhance task-oriented dialogues
Kai Sun, Seungwhan Moon, Paul A. Crook, Stephen Roller, Becka Silvert, Bing Liu, Zhiguang Wang, Honglei Liu, Eunjoon Cho, and Claire Cardie. 2021 · 2021
Earlier work this paper cites.
On the origin of hallucinations in conversational models: Is it the datasets or the models?
Nouha Dziri, Sivan Milton, Mo Yu, Osmar R. Zaïane, and Siva Reddy. 2022b · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Earlier work this paper cites.
A token-level reference-free hallucination detection benchmark for free-form text generation
Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan. 2022 · 2022
Earlier work this paper cites.
Factgraph: Evaluating factuality in summarization with semantic graph representations
Leonardo F. R. Ribeiro, Mengwen Liu, Iryna Gurevych, Markus Dreyer, and Mohit Bansal. 2022 · 2022
Earlier work this paper cites.
A stitch in time saves nine: Enabling early anomaly detection with correlation analysis
Yihao Ang, Qiang Huang, Anthony K. H. Tung, and Zhiyong Huang. 2023 · 2023
Earlier work this paper cites.
Gemini: A family of highly capable multimodal models
Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, and et al. 2023 · 2023
Cited alongside, same era.
The internal state of an LLM knows when its lying
Amos Azaria and Tom M. Mitchell. 2023 · 2023
Cited alongside, same era.
Beyond factuality: A comprehensive evaluation of large language models as knowledge generators
Liang Chen, Yang Deng, Yatao Bian, Zeyu Qin, Bingzhe Wu, Tat-Seng Chua, and Kam-Fai Wong. 2023a · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, and Joseph E. 2023 · 2023
Cited alongside, same era.
Recommender systems in the era of large language models (llms)
Self-contradictory hallucinations of large language models: Evaluation, detection and mitigation
Niels Mündler, Jingxuan He, Slobodan Jenko, and Martin Vechev. 2023 · 2023
Later among the works it cites.
Large language models in medicine: the potentials and pitfalls
Jesutofunmi A. Omiye, Haiwen Gui, Shawheen J. Rezaei, James Zou, and Roxana Daneshjou. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, and Shauli Ravfogel. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenqi Fan, Zihuai Zhao, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Jiliang Tang, and Qing Li. 2023 · 2023
Cited alongside, same era.
Factkb: Generalizable factuality evaluation using language models enhanced with factual knowledge
Shangbin Feng, Vidhisha Balachandran, Yuyang Bai, and Yulia Tsvetkov. 2023b · 2023
Cited alongside, same era.
Are large language models reliable judges? A study on the factuality evaluation capabilities of llms
Xue-Yong Fu, Md. Tahmid Rahman Laskar, Cheng Chen, and Shashi Bhushan TN. 2023 · 2023
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Qianyu Guo, Meng Wang, and Haofen Wang. 2023 · 2023
Cited alongside, same era.
Language models hallucinate, but may excel at fact verification
Jian Guan, Jesse Dodge, David Wadden, Minlie Huang, and Hao Peng. 2023 · 2023
Cited alongside, same era.
Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2023 · 2023
Cited alongside, same era.
Chatgpt for zero-shot dialogue state tracking: A solution or an opportunity?
Michael Heck, Nurul Lubis, Benjamin Matthias Ruppik, Renato Vukovic, Shutong Feng, Christian Geishauser, Hsien-Chin Lin, Carel van Niekerk, and Milica Gasic. 2023 · 2023
Cited alongside, same era.
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang. 2023 · 2023
Cited alongside, same era.
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Later among the works it cites.
Zero-shot information extraction via chatting with chatgpt
Xiang Wei, Xingyu Cui, Ning Cheng, Xiaobin Wang, Xin Zhang, Shen Huang, Pengjun Xie, Jinan Xu, Yufeng Chen, Meishan Zhang, Yong Jiang, and Wenjuan Han. 2023 · 2023
Later among the works it cites.
A multi-task dataset for assessing discourse coherence in chinese essays: Structure, theme, and logic analysis
Hongyi Wu, Xinshu Shen, Man Lan, Shaoguang Mao, Xiaopeng Bai, and Yuanbin Wu. 2023a · 2023
Later among the works it cites.
A new benchmark and reverse validation method for passage-level hallucination detection
Shiping Yang, Renliang Sun, and Xiaojun Wan. 2023 · 2023
Later among the works it cites.
Enhancing uncertainty-based hallucination detection with stronger focus
Tianhang Zhang, Lin Qiu, Qipeng Guo, Cheng Deng, Yue Zhang, Zheng Zhang, Chenghu Zhou, Xinbing Wang, and Luoyi Fu. 2023b · 2023
Later among the works it cites.
Hallucination detection for grounded instruction generation
Lingjun Zhao, Khanh Nguyen, and Hal Daumé III. 2023a · 2023
Later among the works it cites.
Why does chatgpt fall short in providing truthful answers?
Shen Zheng, Jie Huang, and Kevin Chen-Chuan Chang. 2023 · 2023
Later among the works it cites.
Combining discourse coherence with large language models for more inclusive, equitable, and robust task-oriented dialogue
Katherine Atwell, Mert Inan, Anthony B. Sicilia, and Malihe Alikhani. 2024 · 2024
Closest in time.
Reducing hallucination in structured outputs via retrieval-augmented generation
Patrice Béchard and Orlando Marquez Ayala. 2024 · 2024
Closest in time.
Red teaming for large language models at scale: Tackling hallucinations on mathematics tasks
Aleksander Buszydlik, Karol Dobiczek, Michal Teodor Okon, Konrad Skublicki, Philip Lippmann, and Jie Yang. 2024 · 2024
Closest in time.
A regularization-based transfer learning method for information extraction via instructed graph decoder
Kedi Chen, Jie Zhou, Qin Chen, Shunyu Liu, and Liang He. 2024 · 2024
Closest in time.
Navigating hallucinations for reasoning of unintentional activities
Shresth Grover, Vibhav Vineet, and Yogesh S. Rawat. 2024 · 2024
Closest in time.
Language model cascades: Token-level uncertainty and beyond
Neha Gupta, Harikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Aditya Krishna Menon, and Sanjiv Kumar. 2024 · 2024
Closest in time.
Using large language models to assess tutors’ performance in reacting to students making math errors
Sanjit Kakarla, Danielle Thomas, Jionghao Lin, Shivang Gupta, and Kenneth R. Koedinger. 2024 · 2024
Closest in time.
Hallucination detection and hallucination mitigation: An investigation
Junliang Luo, Tianyu Li, Di Wu, Michael Jenkin, Steve Liu, and Gregory Dudek. 2024 · 2024
Closest in time.
Unifying large language models and knowledge graphs: A roadmap
Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu. 2024 · 2024
Closest in time.
A comprehensive survey of hallucination mitigation techniques in large language models
S. M. Towhidul Islam Tonmoy, S. M. Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. 2024 · 2024
Closest in time.
Boosting conversational question answering with fine-grained retrieval-augmentation and self-check
Linhao Ye, Zhikai Lei, Jianghao Yin, Qin Chen, Jie Zhou, and Liang He. 2024 · 2024
Closest in time.
Memorybank: Enhancing large language models with long-term memory
Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. 2024 · 2024
Closest in time.