Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are starting to complement traditional information seeking mechanisms such as web search.
Overview of the medical question answering task at trec 2017 liveqa
Asma Ben Abacha, Eugene Agichtein, Yuval Pinter, and Dina Demner-Fushman · 2017
Earlier work this paper cites.
A hierarchical attention retrieval model for healthcare question answering
Ming Zhu, Aman Ahuja, Wei Wei, and Chandan K Reddy · 2019
Earlier work this paper cites.
Bridging the gap between consumers’ medication questions and trusted answers
Asma Ben Abacha, Yassine Mrabet, Mark Sharp, Travis R Goodwin, Sonya E Shooshan, and Dina Demner-Fushman · 2019
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Evaluating the feasibility of chatgpt in healthcare: an analysis of multiple clinical and research scenarios
Marco Cascella, Jonathan Montomoli, Valentina Bellini, and Elena Bignami · 2023
Earlier work this paper cites.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Earlier work this paper cites.
Large language models in finance: A survey
Yinheng Li, Shaofei Wang, Han Ding, and Hang Chen · 2023
Earlier work this paper cites.
Chatlaw: Open-source legal large language model with integrated external knowledge bases
Jiaxi Cui, Zongjian Li, Yang Yan, Bohua Chen, and Li Yuan · 2023
Earlier work this paper cites.
Chatgpt, 2023
OpenAI · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Earlier work this paper cites.
A survey of hallucination in large foundation models
Vipula Rawte, Amit Sheth, and Amitava Das · 2023
Earlier work this paper cites.
Xmqas: Constructing complex-modified question-answering dataset for robust question understanding
Yuyan Chen, Yanghua Xiao, Zhixu Li, and Bang Liu · 2023
Earlier work this paper cites.
Med-halt: Medical domain hallucination test for large language models
Logesh Kumar Umapathi, Ankit Pal, and Malaikannan Sankarasubbu · 2023
Earlier work this paper cites.
Med-halt: Medical domain hallucination test for large language models
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu · 2023
Earlier work this paper cites.
Siren’s song in the ai ocean: a survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al · 2023
Earlier work this paper cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
OpenAI · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
A comprehensive survey on instruction following
Renze Lou, Kai Zhang, and Wenpeng Yin · 2023
Cited alongside, same era.
Can chatgpt reproduce human-generated labels? a study of social computing tasks
Yiming Zhu, Peixian Zhang, Ehsan-Ul Haq, Pan Hui, and Gareth Tyson · 2023
Cited alongside, same era.
Aligning large multi-modal model with robust instruction tuning
Detecting and evaluating medical hallucinations in large vision language models
Jiawei Chen, Dingkang Yang, Tong Wu, Yue Jiang, Xiaolu Hou, Mingcheng Li, Shunli Wang, Dongling Xiao, Ke Li, and Lihua Zhang · 2024
Closest in time.
Llama-3.1
Meta AI · 2024
Closest in time.
Introducing the next generation of claude, 2024
Anthropic · 2024
Closest in time.
MUFFIN: Curating multi-faceted instructions for improving instruction following
Renze Lou, Kai Zhang, Jian Xie, Yuxuan Sun, Janice Ahn, Hanzi Xu, Yu su, and Wenpeng Yin · 2024
Closest in time.
Mm-soc: Benchmarking multimodal large language models in social media platforms
Yiqiao Jin, Minje Choi, Gaurav Verma, Jindong Wang, and Srijan Kumar · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang · 2023
Cited alongside, same era.
Vibhor Agarwal, Yu Chen, and Nishanth Sastry · 2023
Cited alongside, same era.
Ai in the gray: Exploring moderation policies in dialogic large language models vs. human answers in controversial topics
Vahid Ghafouri, Vibhor Agarwal, Yong Zhang, Nishanth Sastry, Jose Such, and Guillermo Suarez-Tangil · 2023
Cited alongside, same era.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2023
Cited alongside, same era.
Can large language models reason about medical questions?
Valentin Liévin, Christoffer Egeberg Hother, Andreas Geert Motzfeldt, and Ole Winther · 2023
Cited alongside, same era.
Harnessing the power of large language models for natural language to first-order logic translation
Yuan Yang, Siheng Xiong, Ali Payani, Ehsan Shareghi, and Faramarz Fekri · 2023
Cited alongside, same era.
Towards expert-level medical question answering with large language models
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, et al · 2023
Cited alongside, same era.
Hallucination detection: Robustly discerning reliable answers in large language models
Yuyan Chen, Qiang Fu, Yichen Yuan, Zhihao Wen, Ge Fan, Dayiheng Liu, Dongmei Zhang, Zhixu Li, and Yanghua Xiao · 2023
Cited alongside, same era.
Yijia Xiao, Edward Sun, Tianyu Liu, and Wei Wang · 2024
Closest in time.
Competeai: Understanding the competition behaviors in large language model-based agents
Qinlin Zhao, Jindong Wang, Yixuan Zhang, Yiqiao Jin, Kaijie Zhu, Hao Chen, and Xing Xie · 2024
Closest in time.
Can large language model agents simulate human trust behaviors?
Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye, Kai Shu, Adel Bibi, Ziniu Hu, Philip Torr, Bernard Ghanem, and Guohao Li · 2024
Closest in time.
Large language models can learn temporal reasoning
Siheng Xiong, Ali Payani, Ramana Kompella, and Faramarz Fekri · 2024
Closest in time.
Can llms reason in the wild with programs?
Yuan Yang, Siheng Xiong, Ali Payani, Ehsan Shareghi, and Faramarz Fekri · 2024
Closest in time.
When search engine services meet large language models: Visions and challenges
Haoyi Xiong, Jiang Bian, Yuchen Li, Xuhong Li, Mengnan Du, Shuaiqiang Wang, Dawei Yin, and Sumi Helal · 2024
Closest in time.
Ceb: Compositional evaluation benchmark for fairness in large language models
Song Wang, Peng Wang, Tong Zhou, Yushun Dong, Zhen Tan, and Jundong Li · 2024
Closest in time.
Large language models can be good privacy protection learners
Yijia Xiao, Yiqiao Jin, Yushi Bai, Yue Wu, Xianjun Yang, Xiao Luo, Wenchao Yu, Xujiang Zhao, Yanchi Liu, Haifeng Chen, et al · 2024
Closest in time.
Chengyuan Deng, Yiqun Duan, Xin Jin, Heng Chang, Yijun Tian, Han Liu, Henry Peng Zou, Yiqiao Jin, Yijia Xiao, Yichen Wang, et al · 2024
Closest in time.
Codemirage: Hallucinations in code generated by large language models
Vibhor Agarwal, Yulong Pei, Salwa Alamir, and Xiaomo Liu · 2024
Closest in time.
Dr.academy: A benchmark for evaluating questioning capability in education for large language models
Yuyan Chen, Songzhou Yan, Panjun Liu, and Yanghua Xiao · 2024
Closest in time.
Canyu Chen, Baixiang Huang, Zekun Li, Zhaorun Chen, Shiyang Lai, Xiongxiao Xu, Jia-Chen Gu, Jindong Gu, Huaxiu Yao, Chaowei Xiao, et al · 2024
Closest in time.
Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries
Yiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu, Munmun De Choudhury, and Srijan Kumar · 2024
Closest in time.