Fetching the paper…
Reading the bibliography…
Powered by advanced Artificial Intelligence (AI) techniques, conversational AI systems, such as ChatGPT and digital assistants like Siri, have been widely deployed in daily life.
Unified Language Model Pre-training for Natural Language Understanding and Generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 1905
Earlier work this paper cites.
Automatic Acquisition of Hyponyms from Large Text Corpora. In International Conference on Computational Linguistics
Marti A. Hearst. 1992 · 1992
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation. In Annual Meeting of the Association for Computational Linguistics
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out . Association for Computational Linguistics, Barcelona, Spain, 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2020 · 2004
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Hidden Voice Commands. In USENIX Security Symposium
Nicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang, Michael E. Sherr, Clay Shields, David A. Wagner, and Wenchao Zhou. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Conference on Empirical Methods in Natural Language Processing
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
A google self-driving car caused a crash for the first time. [Online]
Chris Ziegler. 2016 · 2016
Earlier work this paper cites.
Learning End-to-End Goal-Oriented Dialog
Antoine Bordes and Jason Weston. 2017 · 2017
Earlier work this paper cites.
Towards Deep Learning Models Resistant to Adversarial Attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017 · 2017
Earlier work this paper cites.
DeepXplore: Automated Whitebox Testing of Deep Learning Systems
Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Sekhar Jana. 2017 · 2017
Earlier work this paper cites.
Semantics-Enhanced Task-Oriented Dialogue Translation: A Case Study on Hotel Booking. In International Joint Conference on Natural Language Processing
Longyue Wang, Jinhua Du, Liangyou Li, Zhaopeng Tu, Andy Way, and Qun Liu. 2017 · 2017
Earlier work this paper cites.
Building Task-Oriented Dialogue Systems for Online Shopping. In AAAI Conference on Artificial Intelligence
Zhao Yan, Nan Duan, Peng Chen, M. Zhou, Jianshe Zhou, and Zhoujun Li. 2017 · 2017
Earlier work this paper cites.
Looking Beyond the Surface: A Challenge Set for Reading Comprehension over Multiple Sentences. In North American Chapter of the Association for Computational Linguistics
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Earlier work this paper cites.
Tesla fatal crash: ’autopilot’ mode sped up car before driver killed, report finds [Online]
Sam Levin. 2018 · 2018
Earlier work this paper cites.
Automated Directed Fairness Testing
Sakshi Udeshi, Pryanshu Arora, and Sudipta Chattopadhyay. 2018 · 2018
Earlier work this paper cites.
Identifying and Reducing Gender Bias in Word-Level Language Models. In North American Chapter of the Association for Computational Linguistics
Shikha Bordia and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. In North American Chapter of the Association for Computational Linguistics
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Queens Are Powerful Too: Mitigating Gender Bias in Dialogue Generation. In Conference on Empirical Methods in Natural Language Processing
Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, and Jason Weston. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Learning to Speak and Act in a Fantasy Text Adventure Game
Jack Urbanek, Angela Fan, Siddharth Karamcheti, Saachi Jain, Samuel Humeau, Emily Dinan, Tim Rocktäschel, Douwe Kiela, Arthur D. Szlam, and Jason Weston. 2019 · 2019
Earlier work this paper cites.
DIALOGPT : Large-Scale Generative Pre-training for Conversational Response Generation. In Annual Meeting of the Association for Computational Linguistics
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and William B. Dolan. 2019 · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
Towards a Human-like Open-Domain Chatbot
Daniel De Freitas, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. 2020 · 2020
Cited alongside, same era.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
Machine Translation Testing via Pathological Invariance
Shashij Gupta. 2020 · 2020
Cited alongside, same era.
EVA: An Open-Domain Chinese Dialogue System with Large-Scale Generative Pre-Training
Hao Zhou, Pei Ke, Zheng Zhang, Yuxian Gu, Yinhe Zheng, Chujie Zheng, Yida Wang, Chen Henry Wu, Hao Sun, Xiaocong Yang, Bosi Wen, Xiaoyan Zhu, Minlie Huang, and Jie Tang. 2021 · 2021
Later among the works it cites.
Taylor Swift ’tried to sue’ Microsoft over racist chatbot Tay
Newsbeat BBC. 2019 · 2022
Later among the works it cites.
29 Top Chatbot Statistics For 2022: Usage, Demographics, Trends
Nicola Bleu. 2022 · 2022
Later among the works it cites.
Apple Statistics
David Curry. 2022 · 2022
Later among the works it cites.
How to make a chatbot that isn’t racist or sexist
Will Heaven. 2020 · 2022
Later among the works it cites.
Tencent’s Multilingual Machine Translation System for WMT22 Large-Scale African Languages. In Conference on Machine Translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vincenzo Riccio, Gunel Jahangirova, Andrea Stocco, Nargiz Humbatova, Michael Weiss, and Paolo Tonella. 2020 · 2020
Cited alongside, same era.
Social Bias Frames: Reasoning about Social and Power Implications of Language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Developing a Twitter bot that can join a discussion using state-of-the-art architectures
Yusuf Mücahit Çetinkaya, Ismail Hakki Toroslu, and Hasan Davulcu. 2020 · 2020
Cited alongside, same era.
Coverage-Guided Testing for Recurrent Neural Networks
2021 · 2021
Cited alongside, same era.
Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive Contexts. In Conference on Empirical Methods in Natural Language Processing
Ashutosh Baheti, Maarten Sap, Alan Ritter, and Mark O. Riedl. 2021 · 2021
Cited alongside, same era.
Bias in machine learning software: why? how? what to do?
Joymallya Chakraborty, Suvodeep Majumder, and Tim Menzies. 2021 · 2021
Cited alongside, same era.
Testing Your Question Answering Software via Asking Recursively
Songqiang Chen, Shuo Jin, and Xiaoyuan Xie. 2021 · 2021
Cited alongside, same era.
BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
J. Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021 · 2021
Cited alongside, same era.
Wenxiang Jiao, Zhaopeng Tu, Jiarui Li, Wenxuan Wang, Jen tse Huang, and Shuming Shi. 2022 · 2022
Later among the works it cites.
QATest: A Uniform Fuzzing Framework for Question Answering Systems
Zixi Liu, Yang Feng, Yining Yin, J. Sun, Zhenyu Chen, and Baowen Xu. 2022 · 2022
Later among the works it cites.
RoMe: A Robust Metric for Evaluating Natural Language Generation. In Annual Meeting of the Association for Computational Linguistics
Md. Rashad Al Hasan Rony, Liubov Kovriguina, Debanjan Chaudhuri, Ricardo Usbeck, and Jens Lehmann. 2022 · 2022
Later among the works it cites.
Natural Test Generation for Precise Testing of Question Answering Software
Qingchao Shen, Junjie Chen, J Zhang, Haoyu Wang, Shuang Liu, and Menghan Tian. 2022 · 2022
Later among the works it cites.
Why So Toxic?: Measuring and Triggering Toxic Behavior in Open-Domain Chatbots
Waiman Si, Michael Backes, Jeremy Blackburn, Emiliano De Cristofaro, Gianluca Stringhini, Savvas Zannettou, and Yand Zhang. 2022 · 2022
Later among the works it cites.
"I’m sorry to hear that": finding bias in language models with a holistic descriptor dataset
Eric Michael Smith, Melissa Hall Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022 · 2022
Later among the works it cites.
On the Safety of Conversational Models: Taxonomy, Dataset, and Benchmark
Hao Sun, Guangxuan Xu, Deng Jiawen, Jiale Cheng, Chujie Zheng, Hao Zhou, Nanyun Peng, Xiaoyan Zhu, and Minlie Huang. 2022 · 2022
Later among the works it cites.
LaMDA: Language Models for Dialog Applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam M. Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, Yaguang Li, Hongrae Lee, Huaixiu Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Yanqi Zhou, Chung-Ching Chang, I. A. Krivokon, Willard James Rusch, Marc Pickett, Kathleen S. Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Hartz Søraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Díaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravindran Rajakumar, Alena Butryna, Matthew Lamm, V. O. Kuzmina, Joseph Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Huai hsin Chi, and Quoc Le. 2022 · 2022
Later among the works it cites.
AEON: a method for automatic evaluation of NLP test cases
Jen tse Huang, Jianping Zhang, Wenxuan Wang, Pinjia He, Yuxin Su, and Michael R. Lyu. 2022 · 2022
Later among the works it cites.
Voice Search Statistics: Smart Speakers, Voice Assistants, and Users in 2022
Josh Wardini. 2022 · 2022
Later among the works it cites.
Social bias, discrimination and inequity in healthcare: mechanisms, implications and recommendations
Craig S. Webster, S Taylor, Courtney Anne De Thomas, and Jennifer M Weller. 2022 · 2022
Later among the works it cites.
Machine Learning Testing: Survey, Landscapes and Horizons
J Zhang, Mark Harman, Lei Ma, and Yang Liu. 2022a · 2022
Later among the works it cites.
Improving Adversarial Transferability via Neuron Attribution-based Attacks
Jianping Zhang, Weibin Wu, Jen tse Huang, Yizhan Huang, Wenxuan Wang, Yuxin Su, and Michael R. Lyu. 2022b · 2022
Later among the works it cites.
Is ChatGPT A Good Translator? A Preliminary Study
Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Xing Wang, and Zhaopeng Tu. 2023 · 2023
Closest in time.
MTTM: Metamorphic Testing for Textual Content Moderation Software
Wenxuan Wang, Jen tse Huang, Weibin Wu, Jianping Zhang, Yizhan Huang, Shuqing Li, Pinjia He, and Michael R. Lyu. 2023 · 2023
Closest in time.
ChatGPT or Grammarly? Evaluating ChatGPT on Grammatical Error Correction Benchmark
Hao Wu, Wenxuan Wang, Yuxuan Wan, Wenxiang Jiao, and Michael R. Lyu. 2023 · 2023
Closest in time.
Improving the Transferability of Adversarial Samples by Path-Augmented Method
Jianping Zhang, Jen tse Huang, Wenxuan Wang, Yichen Li, Weibin Wu, Xiaosen Wang, Yuxin Su, and Michael R. Lyu. 2023 · 2023
Closest in time.