Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have gradually become the gateway for people to acquire new knowledge.
Gestalt psychology
Wolfgang Köhler. 1943 · 1943
Earlier work this paper cites.
A theory of cognitive dissonance
Leon Festinger. 1957 · 1957
Earlier work this paper cites.
Compliance without pressure: the foot-in-the-door technique
Jonathan L. Freedman and Scott C. Fraser. 1966 · 1966
Earlier work this paper cites.
Self-perception: An alternative interpretation of cognitive dissonance phenomena
Daryl J. Bem. 1967 · 1967
Earlier work this paper cites.
Argumentation and persuasion in the cognitive coherence theory
Philippe Pasquier, Iyad Rahwan, Frank Dignum, and Liz Sonenberg. 2006 · 2006
Earlier work this paper cites.
Machine behaviour
Iyad Rahwan, Manuel Cebrian, Nick Obradovich, Josh Bongard, Jean-François Bonnefon, Cynthia Breazeal, Jacob W. Crandall, Nicholas A. Christakis, Iain D. Couzin, Matthew O. Jackson, Nicholas R. Jennings, Ece Kamar, Isabel M. Kloumann, Hugo Larochelle, David Lazer, Richard McElreath, Alan Mislove, David C. Parkes, Alex ‘Sandy’ Pentland, Margaret E. Roberts, Azim Shariff, Joshua B. Tenenbaum, and Michael Wellman. 2019 · 2019
Earlier work this paper cites.
Using large language models to simulate multiple humans and replicate human subject studies
Gati Aher, RosaI. Arriaga, and Adam Tauman Kalai. 2022 · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, John Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, E Perez, Jamie Kerr, Jared Mueller, Jeff Ladish, J Landau, Kamal Ndousse, Kamilė Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noem’i Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, T. J. Henighan, Tristan Hume, Sam Bowman, Zac Hatfield-Dodds, Benjamin Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom B. Brown, and Jared Kaplan. 2022 · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, John Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Benjamin Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zachary Dodds, T. J. Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom B. Brown, Nicholas Joseph, Sam McCandlish, Christopher Olah, Jared Kaplan, and Jack Clark. 2022 · 2022
Earlier work this paper cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Vi’egas, Hanspeter Pfister, and Martin Wattenberg. 2022 · 2022
Earlier work this paper cites.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Earlier work this paper cites.
Playing repeated games with large language models
Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, Matthias Bethge, and Eric Schulz. 2023 · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. 2023 · 2023
Cited alongside, same era.
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang. 2023 · 2023
Cited alongside, same era.
Explore and Browse ChatGPT Prompts on FlowGPT
FlowGPT. 2023 · 2023
Cited alongside, same era.
ChatGPT in Grandma Mode will Spill All Your Secrets
Anirudh V. K. 2023 · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Later among the works it cites.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. 2023 · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, and Xinyu Xing. 2023 · 2023
Later among the works it cites.
Jade: A linguistic-based safety evaluation platform for llm
Mi Zhang, Xudong Pan, and Min Yang. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wes Gurnee and Max Tegmark. 2023 · 2023
Cited alongside, same era.
Thilo Hagendorff. 2023 · 2023
Cited alongside, same era.
Jailbreak Chat
Jailbreak. 2023 · 2023
Cited alongside, same era.
Open sesame! universal black box jailbreaking of large language models
Raz Lapid, Ron Langberg, and Moshe Sipper. 2023 · 2023
Cited alongside, same era.
Samuel Marks and Max Tegmark. 2023 · 2023
Cited alongside, same era.
Llama-2-7b-chat-hf ⋅ \cdot Hugging Face
Meta. 2023 · 2023
Cited alongside, same era.
Xinyue Shen, Zeyuan Johnson Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023 · 2023
Cited alongside, same era.
Large language models understand and can be enhanced by emotional stimuli
Cheng Li, Jindong Wang, Kaijie Zhu, Yixuan Zhang, Wenxin Hou, Jianxun Lian, and Xingxu Xie. 2023a
Cited in the paper.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
Rethinking machine ethics - can llms perform moral reasoning through the lens of moral theories?
Jingyan Zhou, Minda Hu, Junan Li, Xiaoying Zhang, Xixin Wu, Irwin King, and Helen M. Meng. 2023 · 2023
Later among the works it cites.
Claude 2
2024 · 2024
Closest in time.
THUDM/chatglm2-6b ⋅ \cdot Hugging Face
2024a · 2024
Closest in time.
THUDM/chatglm3-6b ⋅ \cdot Hugging Face
2024b · 2024
Closest in time.
Masterkey: Automated jailbreak across multiple large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2024 · 2024
Closest in time.