Machine learning and security
Clarence Chio and David Freeman. 2018 · 2018
Earlier work this paper cites.
Language models are few-shot learners. In NeurIPS
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In Findings of EMNLP
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Learning to summarize with human feedback. In NeurIPS
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Earlier work this paper cites.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In the ACM conference on Fairness, Accountability, and Transparency
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
How social engineers use persuasion principles during vishing attacks
Keith S Jones, Miriam E Armstrong, McKenna K Tornblad, and Akbar Siami Namin. 2021 · 2021
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models. In ACL | IJCNLP
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Earlier work this paper cites.
“Was it “stated” or was it “claimed”?: How linguistic bias affects generative language models. In EMNLP
Roma Patel and Ellie Pavlick. 2021 · 2021
Earlier work this paper cites.
Understanding emails and drafting responses–An approach using GPT-3
Jonas Thiergart, Stefan Huber, and Thomas Übellacker. 2021 · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al · 2021
Earlier work this paper cites.
Language models as agent models. In Findings of EMNLP
Jacob Andreas. 2022 · 2022
Earlier work this paper cites.
Position:“Real Attackers Don’t Compute Gradients”: Bridging the Gap Between Adversarial ML Research and Practice. In SaTML
Giovanni Apruzzese, Hyrum Anderson, Savino Dambra, David Freeman, Fabio Pierazzi, and Kevin Roundy. 2022 · 2022
Earlier work this paper cites.
Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures. In S&P
Eugene Bagdasaryan and Vitaly Shmatikov. 2022 · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Bad characters: Imperceptible nlp attacks. In S&P
Nicholas Boucher, Ilia Shumailov, Ross Anderson, and Nicolas Papernot. 2022 · 2022
Earlier work this paper cites.
Evaluating Human-Language Model Interaction
Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, et al · 2022
Earlier work this paper cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods. In ACL
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback. In NeurIPS
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, et al · 2022
Earlier work this paper cites.
Discovering Language Model Behaviors with Model-Written Evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al · 2022
Earlier work this paper cites.