Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly integrated into software applications.
Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 1877–1901
Tom Brown, Benjamin Mann, et al · 1901
Earlier work this paper cites.
Case study research: Design and methods . Vol. 5
Robert K Yin. 2009 · 2009
Earlier work this paper cites.
How does web service API evolution affect clients?. In 2013 IEEE 20th International Conference on Web Services . IEEE, 300–307
Jun Li, Yingfei Xiong, Xuanzhe Liu, and Lu Zhang. 2013 · 2013
Earlier work this paper cites.
ModelTracker: Redesigning Performance Analysis Tools for Machine Learning. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (Seoul, Republic of Korea) (CHI ’15) . Association for Computing Machinery, 337–346
Saleema Amershi, Max Chickering, Steven M. Drucker, Bongshin Lee, Patrice Simard, and Jina Suh. 2015 · 2015
Earlier work this paper cites.
Software Engineering (10th ed.)
Ian Sommerville. 2015 · 2015
Earlier work this paper cites.
Deceiving google’s perspective api built for detecting toxic comments
Hossein Hosseini, Sreeram Kannan, Baosen Zhang, and Radha Poovendran. 2017 · 2017
Earlier work this paper cites.
Stress Test Evaluation for Natural Language Inference. In Proceedings of the 27th International Conference on Computational Linguistics . Association for Computational Linguistics, 2340–2353
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Earlier work this paper cites.
Bridging the Gap between ML Solutions and Their Business Requirements Using Feature Interactions. In Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (Tallinn, Estonia) (ESEC/FSE 2019) . Association for Computing Machinery, 1048–1058
Guy Barash, Eitan Farchi, Ilan Jayaraman, Orna Raz, Rachel Tzoref-Brill, and Marcel Zalmanovici. 2019 · 2019
Earlier work this paper cites.
Jigsaw Unintended Bias in Toxicity Classification
cjadams, Daniel Borkan, inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and nithum. 2019 · 2019
Earlier work this paper cites.
Losing Confidence in Quality: Unspoken Evolution of Computer Vision Services. In 2019 IEEE International Conference on Software Maintenance and Evolution (ICSME) . 333–342
Alex Cummaudo, Rajesh Vasa, John Grundy, Mohamed Abdelrazek, and Andrew Cain. 2019 · 2019
Earlier work this paper cites.
Machine learning in the AWS cloud: Add intelligence to applications with Amazon Sagemaker and Amazon Rekognition
Abhishek Mishra. 2019 · 2019
Earlier work this paper cites.
Errudite: Scalable, Reproducible, and Testable Error Analysis. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Korhonen, David Traum, and Lluís Màrquez (Eds.). Association for Computational Linguistics, 747–763
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2019 · 2019
Earlier work this paper cites.
Mitigating Uncertainty in Document Classification. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, 3126–3136
Xuchao Zhang, Fanglan Chen, Chang-Tien Lu, and Naren Ramakrishnan. 2019 · 2019
Earlier work this paper cites.
Debugging Tests for Model Explanations. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS’20) . Curran Associates Inc., Article 60, 13 pages
Julius Adebayo, Michael Muelly, Ilaria Liccardi, and Been Kim. 2020 · 2020
Cited alongside, same era.
Beware the Evolving ‘intelligent’ Web Service! An Integration Architecture Tactic to Guard AI-First Components. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (Virtual Event, USA) (ESEC/FSE 2020) . Association for Computing Machinery, 269–280
Alex Cummaudo, Scott Barnett, Rajesh Vasa, John Grundy, and Mohamed Abdelrazek. 2020 · 2020
Cited alongside, same era.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, 4902–4912
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
How is ChatGPT’s behavior changing over time?
Lingjiao Chen, Matei Zaharia, and James Zou. 2023 · 2023
Closest in time.
InCoder: A Generative Model for Code Infilling and Synthesis. In The Eleventh International Conference on Learning Representations
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, and Mike Lewis. 2023 · 2023
Closest in time.
Challenges and applications of large language models
Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. 2023 · 2023
Closest in time.
Pre-Train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023a · 2023
Closest in time.
Instruction Position Matters in Sequence Generation with Large Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automatic Testing and Improvement of Machine Translation. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering (Seoul, South Korea) (ICSE ’20) . Association for Computing Machinery, 974–985
Zeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis, and Lu Zhang. 2020 · 2020
Cited alongside, same era.
Calibrate Before Use: Improving Few-shot Performance of Language Models. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) , Marina Meila and Tong Zhang (Eds.). PMLR, 12697–12706
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Domino: Discovering systematic errors with cross-modal embeddings
Sabri Eyuboglu, Maya Varma, Khaled Saab, Jean-Benoit Delbrouck, Christopher Lee-Messer, Jared Dunnmon, James Zou, and Christopher Ré. 2022 · 2022
Cited alongside, same era.
Machine Learning in Production: From Models to Products
Christian Kästner. 2022 · 2022
Cited alongside, same era.
Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, 8086–8098
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022 · 2022
Cited alongside, same era.
“Did You Miss My Comment or What?” Understanding Toxicity in Open Source Discussions. In 2022 IEEE/ACM 44th International Conference on Software Engineering (ICSE) . 710–722
Courtney Miller, Sophie Cohen, Daniel Klug, Bogdan Vasilescu, and Christian Kästner. 2022 · 2022
Cited alongside, same era.
Adaptive Testing and Debugging of NLP Models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, 3253–3267
Marco Tulio Ribeiro and Scott Lundberg. 2022 · 2022
Cited alongside, same era.
Improving Machine Translation Systems via Isotopic Replacement. In Proceedings of the 44th International Conference on Software Engineering (Pittsburgh, Pennsylvania) (ICSE ’22) . Association for Computing Machinery, 1181–1192
Zeyu Sun, Jie M. Zhang, Yingfei Xiong, Mark Harman, Mike Papadakis, and Lu Zhang. 2022 · 2022
Cited alongside, same era.
Toxicity detection with generative prompt-based inference
Yau-Shian Wang and Yingshan Chang. 2022 · 2022
Cited alongside, same era.
Yanjun Liu, Xianfeng Zeng, Fandong Meng, and Jie Zhou. 2023b · 2023
Closest in time.
Aditi Mishra, Utkarsh Soni, Anjana Arunkumar, Jinbin Huang, Bum Chul Kwon, and Chris Bryan. 2023 · 2023
Closest in time.
LLM is Like a Box of Chocolates: the Non-determinism of ChatGPT in Code Generation
Shuyin Ouyang, Jie M Zhang, Mark Harman, and Meng Wang. 2023 · 2023
Closest in time.
Experiencing Decreased Performance with ChatGPT-4
radiator57. 2023 · 2023
Closest in time.
Testing language models (and prompts) like we test software
Marco Tulio Ribeiro. 2023 · 2023
Closest in time.
ChatLog: Recording and Analyzing ChatGPT Across Time
Shangqing Tu, Chunyang Li, Jifan Yu, Xiaozhi Wang, Lei Hou, and Juanzi Li. 2023 · 2023
Closest in time.
Beyond Testers’ Biases: Guiding Model Testing with Knowledge Bases using LLMs
Chenyang Yang, Rishabh Rustogi, Rachel Brower-Sinning, Grace A Lewis, Christian Kästner, and Tongshuang Wu. 2023b · 2023
Closest in time.
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, Article 437, 21 pages
J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang. 2023 · 2023
Closest in time.
Towards a Unified Multi-Dimensional Evaluator for Text Generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, 2023–2038
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2038
Closest in time.