Fetching the paper…
Reading the bibliography…
While large language models (LLMs) have demonstrated increasing power, they have also given rise to a wide range of harmful behaviors.
The use of worked examples as a substitute for problem solving in learning algebra
John Sweller and Graham A Cooper. 1985 · 1985
Earlier work this paper cites.
Cognitive load during problem solving: Effects on learning
John Sweller. 1988 · 1988
Earlier work this paper cites.
Variability of worked examples and transfer of geometrical problem-solving skills: A cognitive-load approach
Fred GWC Paas and Jeroen JG Van Merriënboer. 1994 · 1994
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020 · 2006
Earlier work this paper cites.
Word order
Matthew S Dryer. 2007 · 2007
Earlier work this paper cites.
The radicalization risks of gpt-3 and advanced neural language models
Kris McGuffie and Alex Newhouse. 2020 · 2009
Earlier work this paper cites.
Writing about testing worries boosts exam performance in the classroom
Gerardo Ramirez and Sian L Beilock. 2011 · 2011
Earlier work this paper cites.
Cognitive load theory
John Sweller. 2011 · 2011
Earlier work this paper cites.
Eyeclosure helps memory by reducing cognitive load and enhancing visualisation
Annelies Vredeveldt, Graham J Hitch, and Alan D Baddeley. 2011 · 2011
Earlier work this paper cites.
Visual environment, attention allocation, and learning in young children: When too much of a good thing may be bad
Anna V Fisher, Karrie E Godwin, and Howard Seltman. 2014 · 2014
Earlier work this paper cites.
On difficulties of cross-lingual transfer with order differences: A case study on dependency parsing
Wasi Ahmad, Zhisong Zhang, Xuezhe Ma, Eduard Hovy, Kai-Wei Chang, and Nanyun Peng. 2019 · 2019
Earlier work this paper cites.
Cognitive architecture and instructional design: 20 years later
John Sweller, Jeroen JG van Merriënboer, and Fred Paas. 2019 · 2019
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Earlier work this paper cites.
Cognitive-load theory: Methods to manage working memory load in the learning of complex tasks
Fred Paas and Jeroen JG van Merriënboer. 2020 · 2020
Earlier work this paper cites.
From theory to practice: the application of cognitive load theory to the practice of medicine
Adam Szulewski, Daniel Howes, Jeroen JG van Merriënboer, and John Sweller. 2020 · 2020
Earlier work this paper cites.
Large language models associate muslims with violence
Abubakar Abid, Maheen Farooqi, and James Zou. 2021 · 2021
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022 · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022 · 2022
Cited alongside, same era.
Understanding dataset difficulty with v-usable information
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. 2022 · 2022
Cited alongside, same era.
An overview of bard: an early experiment with generative ai
James Manyika. 2023 · 2023
Closest in time.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng. 2023 · 2023
Closest in time.
Test-time backdoor mitigation for black-box large language models with defensive demonstrations
Wenjie Mo, Jiashu Xu, Qin Liu, Jiongxiao Wang, Jun Yan, Chaowei Xiao, and Muhao Chen. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
New models and developer products announced at devday
OpenAI. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, et al. 2022a · 2022
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
The hacking of chatgpt is just getting started
Matt Burgess. 2023 · 2023
Cited alongside, same era.
Defending against alignment-breaking attacks via robustly aligned llm
Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023 · 2023
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023 · 2023
Cited alongside, same era.
Enhancing chat language models by scaling high-quality instructional conversations
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023 · 2023
Cited alongside, same era.
Piotr Pęzik, Agnieszka Mikołajczyk, Adam Wawrzyński, Filip Żarnecki, Bartłomiej Nitoń, and Maciej Ogrodniczuk. 2023 · 2023
Closest in time.
Huachuan Qiu, Shuai Zhang, Anqi Li, Hongliang He, and Zhenzhong Lan. 2023 · 2023
Closest in time.
Smoothllm: Defending large language models against jailbreaking attacks
Alexander Robey, Eric Wong, Hamed Hassani, and George J Pappas. 2023 · 2023
Closest in time.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, commercially usable llms
MosaicML NLP Team. 2023 · 2023
Closest in time.
Chatgpt doesn’t have permissions to run programs
themirrazz. 2023 · 2023
Closest in time.
Dan is my new friend
walkerspider. 2023 · 2023
Closest in time.
Wizardlm: Empowering large language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023 · 2023
Closest in time.
Low-resource languages jailbreak gpt-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. 2023 · 2023
Closest in time.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2023 · 2023
Closest in time.
Benchmarking large language models for news summarization
Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B Hashimoto. 2023 · 2023
Closest in time.
Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing. 2023 · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Closest in time.