Fetching the paper…
Reading the bibliography…
Evaluating outputs of large language models (LLMs) is challenging, requiring making -- and making sense of -- many responses.
Prompting is programming: A query language for large language models
Luca Beurer-Kellner, Marc Fischer, and Martin Vechev. 2023 · 1969
Earlier work this paper cites.
Structure mapping in analogy and similarity
Dedre Gentner and Arthur B Markman. 1997 · 1997
Earlier work this paper cites.
Evaluating user interface systems research. In Proceedings of the 20th Annual ACM Symposium on User Interface Software and Technology (Newport, Rhode Island, USA) (UIST ’07) . Association for Computing Machinery, New York, NY, USA, 251–258
Dan R. Olsen. 2007 · 2007
Earlier work this paper cites.
Usability evaluation considered harmful (some of the time). In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Florence, Italy) (CHI ’08) . Association for Computing Machinery, New York, NY, USA, 111–120
Saul Greenberg and Bill Buxton. 2008 · 2008
Earlier work this paper cites.
Necessary conditions of learning
Ference Marton. 2014 · 2014
Earlier work this paper cites.
Intermodulation: Improvisation and Collaborative Art Practice for HCI. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal, QC, Canada) (CHI ’18) . Association for Computing Machinery, New York, NY, USA, 1–13
Laewoo (Leo) Kang, Steven J. Jackson, and Phoebe Sengers. 2018 · 2018
Earlier work this paper cites.
Evaluation Strategies for HCI Toolkit Research. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal, QC, Canada) (CHI ’18) . Association for Computing Machinery, New York, NY, USA, 1–17
David Ledo, Steven Houben, Jo Vermeulen, Nicolai Marquardt, Lora Oehlberg, and Saul Greenberg. 2018 · 2018
Earlier work this paper cites.
Standardizing Reporting of Participant Compensation in HCI: A Systematic Literature Review and Recommendations for the Field. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21) . Association for Computing Machinery, New York, NY, USA, Article 141, 16 pages
Jessica Pater, Amanda Coupe, Rachel Pfafman, Chanda Phelan, Tammy Toscos, and Maia Jacobs. 2021 · 2021
Earlier work this paper cites.
Global reconstruction of language models with linguistic rules–Explainable AI for online consumer reviews
Markus Binder, Bernd Heinrich, Marcus Hopf, and Alexander Schiller. 2022 · 2022
Earlier work this paper cites.
Designing machine learning systems
Chip Huyen. 2022 · 2022
Earlier work this paper cites.
PromptMaker: Prompt-based Prototyping with Large Language Models. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI EA ’22) . Association for Computing Machinery, New York, NY, USA, Article 35, 8 pages
Ellen Jiang, Kristen Olson, Edwin Toh, Alejandra Molina, Aaron Donsbach, Michael Terry, and Carrie J Cai. 2022 · 2022
Earlier work this paper cites.
Designing for Responsible Trust in AI Systems: A Communication Perspective. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT ’22) . Association for Computing Machinery, New York, NY, USA, 1257–1268
Q. Vera Liao and S. Shyam Sundar. 2022 · 2022
Earlier work this paper cites.
Beyond Amazon: Social Justice and Ethical Considerations for Research Compensation
Wing Ng, Ava Anjom, and Joanna M Drinane. 2022 · 2022
Earlier work this paper cites.
Ignore Previous Prompt: Attack Techniques For Language Models. In NeurIPS ML Safety Workshop
Fábio Perez and Ian Ribeiro. 2022 · 2022
Earlier work this paper cites.
Interactive and visual prompt engineering for ad-hoc task adaptation with large language models
Hendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover, Johanna Beyer, Hanspeter Pfister, and Alexander M Rush. 2022 · 2022
Earlier work this paper cites.
Black-box tuning for language-model-as-a-service. In International Conference on Machine Learning . PMLR, 20841–20855
Tianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang, and Xipeng Qiu. 2022 · 2022
Cited alongside, same era.
PromptChainer: Chaining Large Language Model Prompts through Visual Programming. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI EA ’22) . Association for Computing Machinery, New York, NY, USA, Article 359, 10 pages
Tongshuang Wu, Ellen Jiang, Aaron Donsbach, Jeff Gray, Alejandra Molina, Michael Terry, and Carrie J Cai. 2022a · 2022
Cited alongside, same era.
AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22) . Association for Computing Machinery, New York, NY, USA, Article 385, 22 pages
Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022b · 2022
Cited alongside, same era.
Prompt Flow
Microsoft. 2023 · 2023
Closest in time.
Aditi Mishra, Utkarsh Soni, Anjana Arunkumar, Jinbin Huang, Bum Chul Kwon, and Chris Bryan. 2023 · 2023
Closest in time.
Zeno GPT Machine Translation Report
Graham Neubig and Zhiwei He. 2023 · 2023
Closest in time.
OpenAI Playground
OpenAI. 2023 · 2023
Closest in time.
openai/evals
OpenAI. 2023 · 2023
Closest in time.
Supporting Human-AI Collaboration in Auditing LLMs with LLMs
Charvi Rastogi, Marco Tulio Ribeiro, Nicholas King, and Saleema Amershi. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022 · 2022
Cited alongside, same era.
Aleph-Alpha
Aleph-Alpha. 2023 · 2023
Cited alongside, same era.
Grounded copilot: How programmers interact with code-generating models
Shraddha Barke, Michael B James, and Nadia Polikarpova. 2023 · 2023
Cited alongside, same era.
Stephen Brade, Bryan Wang, Mauricio Sousa, Sageev Oore, and Tovi Grossman. 2023 · 2023
Cited alongside, same era.
Jailbreaker: Automated Jailbreak Across Multiple Large Language Model Chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2023 · 2023
Cited alongside, same era.
LangChain
Harrison Chase et al. 2023 · 2023
Cited alongside, same era.
FlowiseAI Build LLMs Apps Easily
FlowiseAI, Inc. 2023 · 2023
Cited alongside, same era.
Graphologue: Exploring Large Language Model Responses with Interactive Diagrams. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23) . Association for Computing Machinery, New York, NY, USA, Article 3, 20 pages
Peiling Jiang, Jude Rayan, Steven P. Dow, and Haijun Xia. 2023 · 2023
Cited alongside, same era.
Cells, Generators, and Lenses: Design Framework for Object-Oriented Interaction with Large Language Models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23) . Association for Computing Machinery, New York, NY, USA, Article 4, 18 pages
Tae Soo Kim, Yoonjoo Lee, Minsuk Chang, and Juho Kim. 2023 · 2023
Cited alongside, same era.
Prompt Space Optimizing Few-shot Reasoning Success with Large Language Models
Fobo Shi, Peijun Qing, Dong Yang, Nan Wang, Youbo Lei, Haonan Lu, and Xiaodong Lin. 2023 · 2023
Closest in time.
Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language Models
Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023 · 2023
Closest in time.
trulens: Evaluate and Track LLM Applications
TruLens. 2023 · 2023
Closest in time.
Vellum The dev platform for production LLM apps
Vellum. 2023 · 2023
Closest in time.
Vercel: Deveop.Preview.Ship
Vercel. 2023 · 2023
Closest in time.
promptfoo: Test your prompts
Ian Webster. 2023 · 2023
Closest in time.
Weights and Biases Docs: Prompts for LLMs
Weights and Biases. 2023 · 2023
Closest in time.
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 437, 21 pages
J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang. 2023 · 2023
Closest in time.