Fetching the paper…
Reading the bibliography…
Deploying large language models (LMs) can pose hazards from harmful outputs such as toxic or false text.
Preference formation
James N Druckman and Arthur Lupia · 2000
Earlier work this paper cites.
Toxic comment classification challenge, 2017
C.J. Adams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark Mcdonald, and Will Cukierski · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou · 2017
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang · 2018
Earlier work this paper cites.
Fever: a large-scale dataset for fact extraction and verification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Multifc: A real-world multi-domain dataset for evidence-based fact checking of claims
Isabelle Augenstein, Christina Lioma, Dongsheng Wang, Lucas Chaves Lima, Casper Hansen, Christian Hansen, and Jakob Grue Simonsen · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Generating natural language adversarial examples through probability weighted word saliency
Shuhuai Ren, Yihe Deng, Kun He, and Wanxiang Che · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Earlier work this paper cites.
Quantifying hypothesis space misspecification in learning from human–robot demonstrations and physical corrections
Andreea Bobu, Andrea Bajcsy, Jaime F Fisac, Sampada Deglurkar, and Anca D Dragan · 2020
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald · 2020
Earlier work this paper cites.
Kilt: a benchmark for knowledge intensive language tasks
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, et al · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh · 2020
Earlier work this paper cites.
Universal adversarial attacks with natural triggers for text classification
Liwei Song, Xinwei Yu, Hsuan-Tung Peng, and Karthik Narasimhan · 2020
Earlier work this paper cites.
Muppet: Massive multi-task representations with pre-finetuning
Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta · 2021
Earlier work this paper cites.
Improving question answering model robustness with synthetic adversarial data generation
Max Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel, Pontus Stenetorp, and Douwe Kiela · 2021
Earlier work this paper cites.
Parrot: Paraphrase generation for nlu., 2021
Prithiviraj Damodaran · 2021
Earlier work this paper cites.
Hard choices in artificial intelligence
Roel Dobbe, Thomas Krendl Gilbert, and Yonatan Mintz · 2021
Earlier work this paper cites.
Truthful ai: Developing and governing ai that does not lie
Owain Evans, Owen Cotton-Barratt, Lukas Finnveden, Adam Bales, Avital Balwit, Peter Wills, Luca Righetti, and William Saunders · 2021
Cited alongside, same era.
Choice set misspecification in reward inference
Rachel Freedman, Rohin Shah, and Anca Dragan · 2021
Cited alongside, same era.
Gradient-based adversarial attacks against text transformers
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela · 2021
Cited alongside, same era.
Hurdles to progress in long-form question answering
Kalpesh Krishna, Aurko Roy, and Mohit Iyyer · 2021
Cited alongside, same era.
Toward human readable prompt tuning: Kubrick’s the shining is a good movie, and a good prompt too?
Weijia Shi, Xiaochuang Han, Hila Gonen, Ari Holtzman, Yulia Tsvetkov, and Luke Zettlemoyer · 2022
Later among the works it cites.
Surge ai, 2023
Surge AI · 2023
Closest in time.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al · 2023
Closest in time.
Chatgpt outperforms crowd-workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli · 2023
Closest in time.
Ground (less) truth: A causal framework for proxy labels in human-algorithm decision-making
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stephanie Lin, Jacob Hilton, and Owain Evans · 2021
Cited alongside, same era.
Ethical aspects of multi-stakeholder recommendation systems
Silvia Milano, Mariarosaria Taddeo, and Luciano Floridi · 2021
Cited alongside, same era.
Creak: A dataset for commonsense reasoning over entity knowledge
Yasumasa Onoe, Michael JQ Zhang, Eunsol Choi, and Greg Durrett · 2021
Cited alongside, same era.
Probing toxic content in large pre-trained language models
Nedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song, and Dit-Yan Yeung · 2021
Cited alongside, same era.
Retrieval augmentation reduces hallucination in conversation
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Cited alongside, same era.
Transformer reinforcement learning x
CarperAI · 2022
Cited alongside, same era.
Luke Guerdan, Amanda Coston, Zhiwei Steven Wu, and Kenneth Holstein · 2023
Closest in time.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Closest in time.
Automatically auditing large language models via discrete optimization
Erik Jones, Anca Dragan, Aditi Raghunathan, and Jacob Steinhardt · 2023
Closest in time.
Pretraining language models with human preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez · 2023
Closest in time.
Gpt-4 vs. gpt-3.5: A concise showdown
Anis Koubaa · 2023
Closest in time.
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar · 2023
Closest in time.
Still no lie detector for language models: Probing empirical and conceptual roadblocks
BA Levinstein and Daniel A Herrmann · 2023
Closest in time.
Multi-step jailbreaking privacy attacks on chatgpt
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, and Yangqiu Song · 2023
Closest in time.
Jailbreaking chatgpt via prompt engineering: An empirical study, 2023
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu · 2023
Closest in time.
Flirt: Feedback loop in-context red teaming
Ninareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu, Shalini Ghosh, Richard Zemel, Kai-Wei Chang, Aram Galstyan, and Rahul Gupta · 2023
Closest in time.
Chat gpt ”dan” (and other ”jailbreaks”)
A.J. Oneal · 2023
Closest in time.
Introducing chatgpt, 2023
OpenAI · 2023
Closest in time.
Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks, 2023
Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya, and Monojit Choudhury · 2023
Closest in time.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Closest in time.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Closest in time.