Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are transforming a wide range of domains, yet verifying their outputs remains a significant challenge, especially for complex open-ended tasks such as consolidation, summarization, and knowledge extraction.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Parallelism in Random Access Machines. In Proceedings of the Tenth Annual ACM Symposium on Theory of Computing (San Diego, CA, USA) (STOC ’78) . Association for Computing Machinery, New York, NY, USA, 114–118
Steven Fortune and James Wyllie. 1978 · 1978
Earlier work this paper cites.
Practical PRAM Programming
Jörg Keller, Christoph Kessler, and Jesper Larsson Träff. 2000 · 2000
Earlier work this paper cites.
Bleu: A Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (Philadelphia, PA, USA) (ACL ’02) , Pierre Isabelle, Eugene Charniak, and Dekang Lin (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries. In Proceedings of the Text Summarization Branches Out Workshop (Barcelona, Spain). Association for Computational Linguistics, Kerrville, TX, USA, 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Parallel Algorithms (2 ed.)
Guy E. Blelloch and Bruce M. Maggs. 2010 · 2010
Earlier work this paper cites.
Neural Text Generation from Structured Data with Application to the Biography Domain. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (Austin, TX, USA) (EMNLP ’16) , Jian Su, Kevin Duh, and Xavier Carreras (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 1203–1213
Rémi Lebret, David Grangier, and Michael Auli. 2016 · 2016
Earlier work this paper cites.
Improved Techniques for Training GANs. In Proceedings of the 30th International Conference on Neural Information Processing Systems (NIPS ’16) (Barcelona, Spain) (Advances in Neural Information Processing Systems, Vol. 29) , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.). Curran Associates, Red Hook, NY, USA, 2234–2242
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016 · 2016
Earlier work this paper cites.
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, CA, USA) (Advances in Neural Information Processing Systems, Vol. 30) , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.). Curran Associates, Red Hook, NY, USA, 6629–6640
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
BLEU Is Not Suitable for the Evaluation of Text Simplification. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (Brussels, Belgium) (EMNLP ’18) , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 738–744
Elior Sulem, Omri Abend, and Ari Rappoport. 2018 · 2018
Earlier work this paper cites.
Language Models as Knowledge Bases?. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (Hong Kong, China) (EMNLP-IJCNLP ’19) , Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 2463–2473
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (Hong Kong, China) (EMNLP-IJCNLP ’19) , Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 3982–3992
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Retrieval-Augmented Language Model Pre-Training. In Proceedings of the 37th International Conference on Machine Learning (ICML ’20) (Virtual Event) (Proceedings of Machine Learning Research, Vol. 119) , Hal Daumé III and Aarti Singh (Eds.). PMLR, New York, NY, USA, 3929–3938
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2020
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the Thirty-Fourth Annual Conference on Neural Information Processing Systems (NeurIPS ’20) (Virtual Event) (Advances in Neural Information Processing Systems, Vol. 33) , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.). Curran Associates, Red Hook, NY, USA, 9459–9474
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Earlier work this paper cites.
COMET: A Neural Framework for MT Evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (Virtual Event) (EMNLP ’20) , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 2685–2702
Ricardo Rei, Craig Stewart, Ana C. Farinha, and Alon Lavie. 2020 · 2020
Earlier work this paper cites.
BLEURT: Learning Robust Metrics for Text Generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (Virtual Event) (ACL ’20) , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 7881–7892
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Earlier work this paper cites.
Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (Virtual Event) (EMNLP ’20) , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 90–121
Brian Thompson and Matt Post. 2020 · 2020
Earlier work this paper cites.
BERTScore: Evaluating Text Generation with BERT. In Proceedings of the Eighth International Conference on Learning Representations (Virtual Event) (ICLR ’20) . OpenReview, Amherst, MA, USA, 43 pages
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
DeBERTa: Decoding-Enhanced BERT With Disentangled Attention. In Proccedings of the Ninth International Conference on Learning Representations (Virtual Event) (ICLR ’21) . OpenReview, Amherst, MA, USA, 21 pages
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML ’21) (Virtual Event) (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, New York, NY, USA, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Earlier work this paper cites.
BARTScore: Evaluating Generated Text as Text Generation. In Proceedings of the Thirty-Fifth Annual Conference on Neural Information Processing Systems (NeurIPS ’21) (Virtual Event) (Advances in Neural Information Processing Systems, Vol. 34) , M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (Eds.). Curran Associates, Red Hook, NY, USA, 27263–27277
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Earlier work this paper cites.
Understanding, Idealization, and Explainable AI
Will Fleisher. 2022 · 2022
Earlier work this paper cites.
SummaC: Re-Visiting NLI-Based Models for Inconsistency Detection in Summarization
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022 · 2022
Earlier work this paper cites.
A Survey of Evaluation Metrics Used for NLG Systems
Ananya B. Sai, Akash Kumar Mohankumar, and Mitesh M. Khapra. 2022 · 2022
Cited alongside, same era.
Explainable Artificial Intelligence (XAI) Post-Hoc Explainability Methods: Risks and Limitations in Non-Discrimination Law
Daniel Vale, Ali El-Sharif, and Muhammed Ali. 2022 · 2022
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Proceedings of the Thirty-Sixth Annual Conference on Neural Information Processing Systems (NeurIPS ’22) (New Orleans, LA, USA) (Advances in Neural Information Processing Systems, Vol. 35) , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.). Curran Associates, Red Hook, NY, USA, 24824–24837
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V. Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Towards a Unified Multi-Dimensional Evaluator for Text Generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (Abu Dhabi, United Arab Emirates) (EMNLP ’22) , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 2023–2038
Dissociating Language and Thought in Large Language Models
Kyle Mahowald, Anna A. Ivanova, Idan A. Blank, Nancy Kanwisher, Joshua B. Tenenbaum, and Evelina Fedorenko. 2024 · 2024
Closest in time.
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (Bangkok, Thailand) (ACL ’24) , Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 10862–10878
Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, KaShun Shum, Randy Zhong, Juntong Song, and Tong Zhang. 2024 · 2024
Closest in time.
Detecting Hallucinations in Large Language Model Generation: A Token Probability Approach. In Proceedings of the 26th International Conference on Artificial Intelligence and Applications (ICAI ’24) (Las Vegas, NV, USA) (Communications in Computer and Information Science (CCIS), Vol. 2252) , Hamid R. Arabnia, Leonidas Deligiannidis, Soheyla Amirian, Farzan Shenavarmasouleh, Farid Ghareh Mohammadi, and David de la Fuente (Eds.). Springer Nature Switzerland, Cham, Switzerland, 154–173
Ernesto Quevedo, Jorge Yero Salazar, Rachel Koerner, Pablo Rivas, and Tomas Cerny. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2022
Cited alongside, same era.
I-Chun Chern, Steffi Chern, Shiqi Chen, Weizhe Yuan, Kehua Feng, Chunting Zhou, Junxian He, Graham Neubig, and Pengfei Liu. 2023 · 2023
Cited alongside, same era.
Halo: Estimation and Reduction of Hallucinations in Open-Source Weak Large Language Models
Mohamed Elaraby, Mengyin Lu, Jacob Dunn, Xueying Zhang, Yu Wang, Shizhu Liu, Pingchuan Tian, Yuping Wang, and Yuxuan Wang. 2023 · 2023
Cited alongside, same era.
Can Large Language Models Explain Themselves? A Study of LLM-Generated Self-Explanations
Shiyuan Huang, Siddarth Mamidanna, Shreedhar Jangam, Yilun Zhou, and Leilani H. Gilpin. 2023 · 2023
Cited alongside, same era.
Towards General Text Embeddings with Multi-Stage Contrastive Learning
Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023 · 2023
Cited alongside, same era.
G-Eval: NLG Evaluation Using GPT-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (Singapore) (EMNLP ’23) , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 2511–2522
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Cited alongside, same era.
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (Singapore) (EMNLP ’23) , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 9004–9017
Potsawee Manakul, Adian Liusie, and Mark Gales. 2023 · 2023
Cited alongside, same era.
AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-Powered Applications. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track (Singapore) (EMNLP ’23) , Mingxuan Wang and Imed Zitouni (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 380–395
Bhaktipriya Radharapu, Kevin Robinson, Lora Aroyo, and Preethi Lahoti. 2023 · 2023
Cited alongside, same era.
A Survey of Hallucination in Large Foundation Models
Vipula Rawte, Amit Sheth, and Amitava Das. 2023 · 2023
Cited alongside, same era.
Closest in time.
Talking about Large Language Models
Murray Shanahan. 2024 · 2024
Closest in time.
Hallucination Index: An Image Quality Metric for Generative Reconstruction Models. In Proceedings of 27th International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI ’24) (Marrakesh, Morocco) (Lecture Notes in Computer Science (LNCS), Vol. 15010) , Marius George Linguraru, Qi Dou, Aasa Feragen, Stamatia Giannarou, Ben Glocker, Karim Lekadir, and Julia A. Schnabel (Eds.). Springer Nature, Cham, Switzerland, 449–458
Matthew Tivnan, Siyeop Yoon, Zhennong Chen, Xiang Li, Dufan Wu, and Quanzheng Li. 2024 · 2024
Closest in time.
A User-Centric Multi-Intent Benchmark for Evaluating Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (Miami, FL, USA) (EMNLP ’24) , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 3588–3612
Jiayin Wang, Fengran Mo, Weizhi Ma, Peijie Sun, Min Zhang, and Jian-Yun Nie. 2024a · 2024
Closest in time.
Improving Text Embeddings with Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (Bangkok, Thailand) (ACL ’24) , Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Kerrville, TX, USA, 11897–11916
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024c · 2024
Closest in time.
C-Pack: Packed Resources For General Chinese Embeddings. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington, DC, USA) (SIGIR ’24) . Association for Computing Machinery, New York, NY, USA, 641–649
Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. 2024 · 2024
Closest in time.
Explainability for Large Language Models: A Survey
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024 · 2024
Closest in time.
Large Language Models for Information Retrieval: A Survey
Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Haonan Chen, Zheng Liu, Zhicheng Dou, and Ji-Rong Wen. 2024 · 2024
Closest in time.
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Saizhuo Wang, Kun Zhang, Yuanzhuo Wang, Wen Gao, Lionel Ni, and Jian Guo. 2025 · 2025
Closest in time.
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025 · 2025
Closest in time.
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models. In Proceedings of the Thirteenth International Conference on Learning Representations (Singapore) (ICLR ’25) . OpenReview, Amherst, MA, USA, 24 pages
Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2025 · 2025
Closest in time.
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-Judge
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. 2025 · 2025
Closest in time.
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
Youngsun Lim, Hojun Choi, and Hyunjung Shim. 2025 · 2025
Closest in time.
SFR-Embedding-Mistral: Enhance Text Retrieval with Transfer Learning
Rui Meng, Ye Liu, Shafiq Rayhan Joty, Caiming Xiong, Yingbo Zhou, and Semih Yavuz. 2024 · 2025
Closest in time.
Large Language Models: A Survey
Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2025 · 2025
Closest in time.
NovaSearch/stella_en_1.5B_v5
NovaSearch. 2024a · 2025
Closest in time.
NovaSearch/stella_en_400M_v5
NovaSearch. 2024b · 2025
Closest in time.
The Rise and Potential of Large Language Model Based Agents: A Survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, Rui Zheng, Xiaoran Fan, Xiao Wang, Limao Xiong, Yuhao Zhou, Weiran Wang, Changhao Jiang, Yicheng Zou, Xiangyang Liu, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang, Qi Zhang, and Tao Gui. 2025 · 2025
Closest in time.
A Survey of Large Language Models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2025 · 2025
Closest in time.
OpenAI Text-Embedding-Large: New Embedding Models and API Updates
Juntang Zhuang, Paul Baltescu, Joy Jiao, Arvind Neelakantan, Andrew Braunstein, Jeff Harris, Logan Kilpatrick, Leher Pathak, Enoch Cheung, Ted Sanders, Yutian Liu, Anushree Agrawal, Andrew Peng, Ian Kivlichan, Mehmet Yatbaz, Madelaine Boyd, Anna-Luisa Brakman, Florencia Leoni Aleman, Henry Head, Molly Lin, Meghan Shah, Chelsea Carlson, Sam Toizer, Ryan Greene, Alison Harmon, Denny Jin, Karolis Kosas, Marie Inuzuka, Peter Bakkum, Barret Zoph, Luke Metz, Jiayi Weng, Randall Lin, Yash Patil, Mianna Chen, Andrew Kondrich, Brydon Eastman, Liam Fedus, John Schulman, Vlad Fomenko, Andrej Karpathy, Aidan Clark, and Owen Campbell-Moore. 2024 · 2025
Closest in time.