Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are being increasingly deployed as part of pipelines that repeatedly process or generate data of some sort.
Ease.Ml/Ci and Ease.Ml/Meter in Action: Towards Data Management for Statistical Generalization
Cedric Renggli, Frances Ann Hubis, Bojan Karlaš, Kevin Schawinski, Wentao Wu, and Ce Zhang. 2019 · 1965
Earlier work this paper cites.
CBC user guide
John Forrest and Robin Lougee-Heimer. 2005 · 2005
Earlier work this paper cites.
Active learning and sampling
Rui Castro and Robert Nowak. 2008 · 2008
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
The Data Linter: Lightweight Automated Sanity Checking for ML Data Sets
Nick Hynes, D. Sculley, and Michael Terry. 2017 · 2017
Earlier work this paper cites.
Data lifecycle challenges in production machine learning: a survey
Neoklis Polyzotis, Sudip Roy, Steven Euijong Whang, and Martin Zinkevich. 2018 · 2018
Earlier work this paper cites.
Automating Large-Scale Data Quality Verification
Sebastian Schelter, Dustin Lange, Philipp Schmidt, Meltem Celikel, Felix Biessmann, and Andreas Grafberger. 2018 · 2018
Earlier work this paper cites.
Data Platform for Machine Learning. In Proceedings of the 2019 International Conference on Management of Data (Amsterdam, Netherlands) (SIGMOD ’19) . Association for Computing Machinery, New York, NY, USA, 1803–1816
Pulkit Agrawal, Rajat Arya, Aanchal Bindal, Sandeep Bhatia, Anupriya Gagneja, Joseph Godlewski, Yucheng Low, Timothy Muss, Mudit Manu Paliwal, Sethu Raman, Vishrut Shah, Bochao Shen, Laura Sugden, Kaiyu Zhao, and Ming-Chuan Wu. 2019 · 2019
Earlier work this paper cites.
Data Validation for Machine Learning. In Proceedings of SysML
Eric Breck, Marty Zinkevich, Neoklis Polyzotis, Steven Whang, and Sudip Roy. 2019b · 2019
Earlier work this paper cites.
Model assertions for monitoring and improving ML models
Daniel Kang, Deepti Raghavan, Peter Bailis, and Matei Zaharia. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Vamsa: Automated Provenance Tracking in Data Science Scripts. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Virtual Event, CA, USA) (KDD ’20) . Association for Computing Machinery, New York, NY, USA, 1542–1551
Mohammad Hossein Namaki, Avrilia Floratou, Fotis Psallidas, Subru Krishnan, Ashvin Agrawal, Yinghui Wu, Yiwen Zhu, and Markus Weimer. 2020 · 2020
Earlier work this paper cites.
Snorkel: Rapid training data creation with weak supervision
Alexander Ratner, Stephen H Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher Ré. 2020 · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A Smith, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021 · 2021
Earlier work this paper cites.
Ask me anything: A simple strategy for prompting language models
Simran Arora, Avanika Narayan, Mayee F Chen, Laurel Orr, Neel Guha, Kush Bhatia, Ines Chami, Frederic Sala, and Christopher Ré. 2022 · 2022
Earlier work this paper cites.
Data distribution debugging in machine learning pipelines
Stefan Grafberger, Paul Groth, Julia Stoyanovich, and Sebastian Schelter. 2022 · 2022
Earlier work this paper cites.
Finding label and model errors in perception data with learned observation assertions. In Proceedings of the 2022 International Conference on Management of Data . 496–505
Daniel Kang, Nikos Arechiga, Sudeep Pillai, Peter D Bailis, and Matei Zaharia. 2022 · 2022
Earlier work this paper cites.
Can foundation models wrangle your data?
Avanika Narayan, Ines Chami, Laurel Orr, Simran Arora, and Christopher Ré. 2022 · 2022
Earlier work this paper cites.
Challenges in deploying machine learning: a survey of case studies
Andrei Paleyes, Raoul-Gabriel Urma, and Neil D Lawrence. 2022 · 2022
Earlier work this paper cites.
Operationalizing machine learning: An interview study
Shreya Shankar, Rolando Garcia, Joseph M Hellerstein, and Aditya G Parameswaran. 2022a · 2022
Cited alongside, same era.
Rethinking streaming machine learning evaluation
Shreya Shankar, Bernease Herman, and Aditya G Parameswaran. 2022b · 2022
Cited alongside, same era.
Prompting gpt-3 to be reliable
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–22
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim. 2023 · 2023
Later among the works it cites.
Langchain AI
Langchain 2023 · 2023
Later among the works it cites.
CODAMOSA: Escaping coverage plateaus in test generation with pre-trained large language models. In International conference on software engineering (ICSE)
Caroline Lemieux, Jeevana Priya Inala, Shuvendu K Lahiri, and Siddhartha Sen. 2023 · 2023
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Later among the works it cites.
Llama Index
Llama Index 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022b · 2022
Cited alongside, same era.
Large language models are human-level prompt engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022 · 2022
Cited alongside, same era.
Anastasios N Angelopoulos, Stephen Bates, Clara Fannjiang, Michael I Jordan, and Tijana Zrnic. 2023 · 2023
Cited alongside, same era.
ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
Ian Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg, and Elena Glassman. 2023 · 2023
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Cited alongside, same era.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2023
Cited alongside, same era.
How is ChatGPT’s behavior changing over time?
Lingjiao Chen, Matei Zaharia, and James Zou. 2023 · 2023
Cited alongside, same era.
Prompt Sapper: A LLM-Empowered Production Tool for Building AI Chains
Yu Cheng, Jieshan Chen, Qing Huang, Zhenchang Xing, Xiwei Xu, and Qinghua Lu. 2023 · 2023
Cited alongside, same era.
Shuyin Ouyang, Jie M Zhang, Mark Harman, and Meng Wang. 2023 · 2023
Later among the works it cites.
Revisiting Prompt Engineering via Declarative Crowdsourcing
Aditya G Parameswaran, Shreya Shankar, Parth Asawa, Naman Jain, and Yujie Wang. 2023 · 2023
Later among the works it cites.
Building Your Own Product Copilot: Challenges, Opportunities, and Needs
Chris Parnin, Gustavo Soares, Rahul Pandita, Sumit Gulwani, Jessica Rich, and Austin Z. Henley. 2023 · 2023
Later among the works it cites.
Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails
Traian Rebedea, Razvan Dinu, Makesh Sreedhar, Christopher Parisien, and Jonathan Cohen. 2023 · 2023
Later among the works it cites.
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2023 · 2023
Later among the works it cites.
An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation
Max Schäfer, Sarah Nadi, Aryaz Eghbali, and Frank Tip. 2023 · 2023
Later among the works it cites.
Proactively Screening Machine Learning Pipelines with ARGUSEYES. In Companion of the 2023 International Conference on Management of Data (Seattle, WA, USA) (SIGMOD ’23) . Association for Computing Machinery, New York, NY, USA, 91–94
Sebastian Schelter, Stefan Grafberger, Shubha Guha, Bojan Karlas, and Ce Zhang. 2023 · 2023
Later among the works it cites.
Exploring the Effectiveness of Large Language Models in Generating Unit Tests
Mohammed Latif Siddiq, Joanna Santos, Ridwanul Hasan Tanvir, Noshin Ulfat, Fahmid Al Rifat, and Vinicius Carvalho Lopes. 2023 · 2023
Later among the works it cites.
DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines
Arnav Singhvi, Manish Shetty, Shangyin Tan, Christopher Potts, Koushik Sen, Matei Zaharia, and Omar Khattab. 2023 · 2023
Later among the works it cites.
Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation
Benjamin Steenhoek, Michele Tufano, Neel Sundaresan, and Alexey Svyatkovskiy. 2023 · 2023
Later among the works it cites.
Software testing with large language model: Survey, landscape, and vision
Junjie Wang, Yuchao Huang, Chunyang Chen, Zhe Liu, Song Wang, and Qing Wang. 2023a · 2023
Later among the works it cites.
Aligning large language models with human: A survey
Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu. 2023c · 2023
Later among the works it cites.
Evaluating NLG Evaluation Metrics: A Measurement Theory Perspective
Ziang Xiao, Susu Zhang, Vivian Lai, and Q Vera Liao. 2023 · 2023
Later among the works it cites.
Why Johnny can’t prompt: how non-AI experts try (and fail) to design LLM prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–21
JD Zamfirescu-Pereira, Richmond Y Wong, Bjoern Hartmann, and Qian Yang. 2023 · 2023
Later among the works it cites.
Wider and deeper llm networks are fairer llm evaluators
Xinghua Zhang, Bowen Yu, Haiyang Yu, Yangyu Lv, Tingwen Liu, Fei Huang, Hongbo Xu, and Yongbin Li. 2023 · 2023
Later among the works it cites.