Fetching the paper…
Reading the bibliography…
The text generated by large language models is commonly controlled by prompting, where a prompt prepended to a user's query guides the model's output.
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks, December 2014
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Reading Wikipedia to Answer Open-Domain Questions, April 2017
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes · 2017
Earlier work this paper cites.
SGDR: Stochastic Gradient Descent with Warm Restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Energy and Policy Considerations for Deep Learning in NLP, June 2019
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, September 2020
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization, July 2020
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J. Liu · 2020
Earlier work this paper cites.
Fine-Tuning Language Models from Human Preferences, January 2020
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2020
Earlier work this paper cites.
DEBERTA: DECODING-ENHANCED BERT WITH DISENTANGLED ATTENTION
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding, January 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
How many data points is a prompt worth?
Teven Le Scao and Alexander Rush · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Xiang Lisa Li and Percy Liang · 2021
Earlier work this paper cites.
Trading Off Diversity and Quality in Natural Language Generation
Hugh Zhang, Daniel Duckworth, Daphne Ippolito, and Arvind Neelakantan · 2021
Cited alongside, same era.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, April 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Cited alongside, same era.
Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor, December 2022
Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, March 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Cited alongside, same era.
Prompt Injection attack against LLM-integrated Applications, June 2023
Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu · 2023
Closest in time.
Black Box Adversarial Prompting for Foundation Models, May 2023
Natalie Maus, Patrick Chao, Eric Wong, and Jacob Gardner · 2023
Closest in time.
Introducing the new Bing
Microsoft · 2023
Closest in time.
Language Model Inversion, November 2023
John X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov, and Alexander M. Rush · 2023
Closest in time.
Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities, August 2023
Maximilian Mozes, Xuanli He, Bennett Kleinberg, and Lewis D. Griffin · 2023
Closest in time.
GPT-4 Technical Report, March 2023
OpenAI · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Discovering Language Model Behaviors with Model-Written Evaluations, December 2022
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2022
Cited alongside, same era.
Ignore Previous Prompt: Attack Techniques For Language Models
Fábio Perez and Ian Ribeiro · 2022
Cited alongside, same era.
Prompt injection attacks against GPT-3
Simon Willison · 2022
Cited alongside, same era.
PaLM 2 Technical Report, May 2023
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, et al · 2023
Cited alongside, same era.
INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models, June 2023
Yew Ken Chia, Pengfei Hong, Lidong Bing, and Soujanya Poria · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Cited alongside, same era.
Prompting GitHub Copilot Chat to become your personal AI assistant for accessibility, October 2023
Ed Summers Dugas, Jesse · 2023
Cited alongside, same era.
Not what you’ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, May 2023
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Cited alongside, same era.
A List of Leaked System Prompts
Matt Rickard · 2023
Closest in time.
Stanford alpaca: An instruction-following LLaMA model, 2023
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
These are Microsoft’s Bing AI secret rules and why it says it’s named Sydney
Tom Warren · 2023
Closest in time.
Jailbroken: How Does LLM Safety Training Fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Closest in time.
Large Language Models Are Human-Level Prompt Engineers, March 2023
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba · 2023
Closest in time.
Universal and Transferable Adversarial Attacks on Aligned Language Models, July 2023
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Closest in time.
Introducing the next generation of Claude
Anthropic · 2024
Closest in time.