Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) like ChatGPT are foundational in various applications due to their extensive knowledge from pre-training and fine-tuning.
Unified Language Model Pre-training for Natural Language Understanding and Generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 1905
Earlier work this paper cites.
A Declarative Metamorphic Testing Framework for Autonomous Driving
Yao Deng, Xi Zheng, Tianyi Zhang, Huai Liu, Guannan Lou, Miryung Kim, and Tsong Yueh Chen. 2020 · 1982
Earlier work this paper cites.
WordNet: A Lexical Database for English
George A. Miller. 1995 · 1995
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation. In Annual Meeting of the Association for Computational Linguistics
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out . Association for Computational Linguistics, Barcelona, Spain, 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Freebase: a collaboratively created graph database for structuring human knowledge. In SIGMOD Conference
Kurt D. Bollacker, Colin Evans, Praveen K. Paritosh, Tim Sturge, and Jamie Taylor. 2008 · 2008
Earlier work this paper cites.
Testing and validating machine learning classifiers by metamorphic testing
Xiaoyuan Xie, Joshua W. K. Ho, Christian Murphy, Gail E. Kaiser, Baowen Xu, and Tsong Yueh Chen. 2011 · 2011
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
Denny Vrandečić and Markus Krötzsch. 2014 · 2014
Earlier work this paper cites.
Hidden Voice Commands. In USENIX Security Symposium
Nicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang, Michael E. Sherr, Clay Shields, David A. Wagner, and Wenchao Zhou. 2016 · 2016
Earlier work this paper cites.
A google self-driving car caused a crash for the first time. [Online]
Chris Ziegler. 2016 · 2016
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
Paul Francis Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Adversarial Examples for Evaluating Reading Comprehension Systems
Robin Jia and Percy Liang. 2017 · 2017
Earlier work this paper cites.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
DeepXplore: Automated Whitebox Testing of Deep Learning Systems
Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Sekhar Jana. 2017 · 2017
Earlier work this paper cites.
The science of fake news
David Lazer, Matthew A. Baum, Yochai Benkler, Adam J. Berinsky, Kelly M. Greenhill, Filippo Menczer, Miriam J. Metzger, Brendan Nyhan, Gordon Pennycook, David M. Rothschild, Michael Schudson, Steven A. Sloman, Cass Robert Sunstein, Emily A. Thorson, Duncan J. Watts, and Jonathan Zittrain. 2018 · 2018
Earlier work this paper cites.
Tesla fatal crash: ’autopilot’ mode sped up car before driver killed, report finds [Online]
Sam Levin. 2018 · 2018
Earlier work this paper cites.
DeepGauge: Multi-Granularity Testing Criteria for Deep Learning Systems
L. Ma, Felix Juefei-Xu, Fuyuan Zhang, Jiyuan Sun, Minhui Xue, Bo Li, Chunyang Chen, Ting Su, Li Li, Yang Liu, Jianjun Zhao, and Yadong Wang. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
DeepRoad: GAN-Based Metamorphic Testing and Input Validation Framework for Autonomous Driving Systems
Mengshi Zhang, Yuqun Zhang, Lingming Zhang, Cong Liu, and Sarfraz Khurshid. 2018 · 2018
Earlier work this paper cites.
DeepBillboard: Systematic Physical-World Testing of Autonomous Driving Systems
Husheng Zhou, Wei Li, Yuankun Zhu, Yuqun Zhang, Bei Yu, Lingming Zhang, and Cong Liu. 2018 · 2018
Earlier work this paper cites.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. In North American Chapter of the Association for Computational Linguistics
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Improving the Robustness of Question Answering Systems to Question Paraphrasing. In Annual Meeting of the Association for Computational Linguistics
Wee Chung Gan and Hwee Tou Ng. 2019 · 2019
Earlier work this paper cites.
Comparing Offline and Online Testing of Deep Neural Networks: An Autonomous Car Case Study
Fitash Ul Haq, Donghwan Shin, Shiva Nejati, and Lionel Claude Briand. 2019 · 2019
Earlier work this paper cites.
Coverage-Guided Testing for Recurrent Neural Networks
Wei Huang, Youcheng Sun, Xing-E. Zhao, James Sharp, Wenjie Ruan, Jie Meng, and Xiaowei Huang. 2019 · 2019
Earlier work this paper cites.
Natural Questions: a Benchmark for Question Answering Research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
R. Thomas McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Earlier work this paper cites.
Language Models as Knowledge Bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2019 · 2019
Cited alongside, same era.
Automatic Testing and Improvement of Machine Translation
Zeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis, and Lu Zhang. 2019 · 2019
Cited alongside, same era.
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, Minneapolis, Minnesota, 4149–4158
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019a · 2019
Cited alongside, same era.
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019b · 2019
Cited alongside, same era.
Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language Models
Yinlin Deng, Chun Xia, Haoran Peng, Chenyuan Yang, and Lingming Zhang. 2022 · 2022
Later among the works it cites.
Efficient Online Testing for DNN-Enabled Systems using Surrogate-Assisted and Many-Objective Optimization
Fitash Ul Haq, Donghwan Shin, and Lionel Claude Briand. 2022 · 2022
Later among the works it cites.
Large Language Models are Few-shot Testers: Exploring LLM-based General Bug Reproduction
Sungmin Kang, Juyeon Yoon, and Shin Yoo. 2022 · 2022
Later among the works it cites.
QATest: A Uniform Fuzzing Framework for Question Answering Systems
Zixi Liu, Yang Feng, Yining Yin, J. Sun, Zhenyu Chen, and Baowen Xu. 2022 · 2022
Later among the works it cites.
Memory-Based Model Editing at Scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning, and Chelsea Finn. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Machine Learning Testing: Survey, Landscapes and Horizons
J Zhang, Mark Harman, Lei Ma, and Yang Liu. 2019 · 2019
Cited alongside, same era.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
Machine Translation Testing via Pathological Invariance
Shashij Gupta. 2020 · 2020
Cited alongside, same era.
A Survey on Knowledge Graphs: Representation, Acquisition, and Applications
Shaoxiong Ji, Shirui Pan, E. Cambria, Pekka Marttinen, and Philip S. Yu. 2020 · 2020
Cited alongside, same era.
Similarity Based Information Retrieval Using Levenshtein Distance Algorithm
Daw Khin Po. 2020 · 2020
Cited alongside, same era.
Testing machine learning based systems: a systematic mapping
Vincenzo Riccio, Gunel Jahangirova, Andrea Stocco, Nargiz Humbatova, Michael Weiss, and Paolo Tonella. 2020 · 2020
Cited alongside, same era.
Automatic Testing and Improvement of Machine Translation
Zeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis, and Lu Zhang. 2020 · 2020
Cited alongside, same era.
White-box Fairness Testing through Adversarial Sampling
Peixin Zhang, Jingyi Wang, Jun Sun, Guoliang Dong, Xinyu Wang, Xingen Wang, Jin Song Dong, and Ting Dai. 2020 · 2020
Cited alongside, same era.
Natural Test Generation for Precise Testing of Question Answering Software
Qingchao Shen, Junjie Chen, J Zhang, Haoyu Wang, Shuang Liu, and Menghan Tian. 2022 · 2022
Later among the works it cites.
AEON: a method for automatic evaluation of NLP test cases
Jen tse Huang, Jianping Zhang, Wenxuan Wang, Pinjia He, Yuxin Su, and Michael R. Lyu. 2022 · 2022
Later among the works it cites.
Machine Learning Testing: Survey, Landscapes and Horizons
J Zhang, Mark Harman, Lei Ma, and Yang Liu. 2022 · 2022
Later among the works it cites.
Can we trust the evaluation on ChatGPT?
Rachith Aiyappa, Jisun An, Haewoon Kwak, and Yong-Yeol Ahn. 2023 · 2023
Later among the works it cites.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023 · 2023
Later among the works it cites.
Towards Autonomous Testing Agents via Conversational Large Language Models
Robert Feldt, Sungmin Kang, Juyeon Yoon, and Shin Yoo. 2023 · 2023
Later among the works it cites.
Constructing Effective In-Context Demonstration for Code Intelligence Tasks: An Empirical Study
Shuzheng Gao, Xinjie Wen, Cuiyun Gao, Wenxuan Wang, and Michael R. Lyu. 2023 · 2023
Later among the works it cites.
ChatGPT Is The Fastest Growing App In The History Of Web Applications
Cindy Gordon. 2023 · 2023
Later among the works it cites.
Zero-shot Faithful Factual Error Correction. In Annual Meeting of the Association for Computational Linguistics
Kung-Hsiang Huang, Hou Pong Chan, and Heng Ji. 2023 · 2023
Later among the works it cites.
Is ChatGPT A Good Translator? A Preliminary Study
Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Xing Wang, and Zhaopeng Tu. 2023 · 2023
Later among the works it cites.
Amr Keleg and Walid Magdy. 2023 · 2023
Later among the works it cites.
A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets. In Annual Meeting of the Association for Computational Linguistics
Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq R. Joty, and J. Huang. 2023 · 2023
Later among the works it cites.
CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language Models
Caroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, and Siddhartha Sen. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Evaluating the Robustness of Machine Reading Comprehension Models to Low Resource Entity Renaming
Clemencia Siro and Tunde Oluwaseyi Ajayi. 2023a · 2023
Later among the works it cites.
BiasAsker: Measuring the Bias in Conversational AI System
Yuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu, Haonan Bai, and Michael R. Lyu. 2023 · 2023
Later among the works it cites.
Validating Multimedia Content Moderation Software via Semantic Fusion
Wenxuan Wang, Jingyuan Huang, Chang Chen, Jiazhen Gu, Jianping Zhang, Weibin Wu, Pinjia He, and Michael R. Lyu. 2023a · 2023
Later among the works it cites.
An Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation Software
Wenxuan Wang, Jingyuan Huang, Jen tse Huang, Chang Chen, Jiazhen Gu, Pinjia He, and Michael R. Lyu. 2023b · 2023
Later among the works it cites.
MTTM: Metamorphic Testing for Textual Content Moderation Software
Wenxuan Wang, Jen tse Huang, Weibin Wu, Jianping Zhang, Yizhan Huang, Shuqing Li, Pinjia He, and Michael R. Lyu. 2023c · 2023
Later among the works it cites.
ChatGPT or Grammarly? Evaluating ChatGPT on Grammatical Error Correction Benchmark
Hao Wu, Wenxuan Wang, Yuxuan Wan, Wenxiang Jiao, and Michael R. Lyu. 2023 · 2023
Later among the works it cites.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al · 2023
Later among the works it cites.