Fetching the paper…
Reading the bibliography…
In this paper, we present JADE, a targeted linguistic fuzzing platform which strengthens the linguistic complexity of seed questions to simultaneously and consistently break a wide range of widely-used LLMs categorized in three groups: eight open-sourced Chinese, six commercial Chinese and four commercial English LLMs.
A model and an hypothesis for language structure
Victor H. Yngve · 1960
Earlier work this paper cites.
Some syntactic determinants of sentential complexity
Jerry A. Fodor and Merrill F. Garrett · 1967
Earlier work this paper cites.
Some syntactic determinants of sentential complexity, ii : Verb structure
Jerry A. Fodor, Merrill F. Garrett, and Thomas G. Bever · 1968
Earlier work this paper cites.
Language and problems of knowledge
Noam Chomsky · 1987
Earlier work this paper cites.
De Gruyter Mouton, Berlin, Boston, 1996
Deep Structure, Surface Structure and Semantic Interpretation · 1996
Earlier work this paper cites.
Assessing Vocabulary
John Read · 2000
Earlier work this paper cites.
Syntactic structures
Noam Chomsky · 2002
Earlier work this paper cites.
On operationalizing syntactic complexity
Benedikt Szmrecsanyi · 2004
Earlier work this paper cites.
Understanding and Measuring Morphological Complexity
Matthew Baerman, Dunstan Brown, and Greville G. Corbett · 2015
Earlier work this paper cites.
https://github.com/nikitakit/self-attentive-parser
nikitakit/self-attentive-parser (github repository), 2018 · 2018
Earlier work this paper cites.
Constituency parsing with a self-attentive encoder
Nikita Kitaev and Dan Klein · 2018
Earlier work this paper cites.
Multilingual constituency parsing with self-attention and pre-training
Nikita Kitaev, Steven Cao, and Dan Klein · 2019
Earlier work this paper cites.
Tree transformer: Integrating tree structures into self-attention
Yau-Shian Wang, Hung yi Lee, and Yun-Nung (Vivian) Chen · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, et al · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramèr, Eric Wallace, et al · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie C. Lin, Jacob Hilton, and Owain Evans · 2021
Earlier work this paper cites.
A word on machine ethics: A response to jiang et al. (2021)
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams · 2021
Earlier work this paper cites.
https://www.forbes.com/sites/davidbirch/2022/12/08/chatgpt-is-a-window-into-the-real-future-of-financial-services/?sh=1df5295f59e2
ChatGPT Is A Window Into The Real Future Of Financial Services, 2022 · 2022
Earlier work this paper cites.
https://openai.com/blog/chatgpt
Introducing ChatGPT, 2022 · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, et al · 2022
Cited alongside, same era.
Ctap for chinese:a linguistic complexity feature automatic calculation platform
Yue Cui, Junhui Zhu, Liner Yang, Xuezhi Fang, Xiaobin Chen, Yujie Wang, and Erhong Yang · 2022
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, et al · 2022
Cited alongside, same era.
Pile of law: Learning responsible data filtering from the law and a 256gb open-source legal dataset
Peter Henderson, Mark S. Krass, Lucia Zheng, Neel Guha, Christopher D. Manning, Dan Jurafsky, and Daniel E. Ho · 2022
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan · 2023
Closest in time.
Evaluating superhuman models with consistency checks
Lukas Fluri, Daniel Paleka, and Florian Tramèr · 2023
Closest in time.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi · 2023
Closest in time.
Meta semantic template for evaluation of large language models
Yachuan Liu, Liang Chen, Jindong Wang, Qiaozhu Mei, and Xing Xie · 2023
Closest in time.
Jailbreaking chatgpt via prompt engineering: An empirical study
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Wenliang Dai, Andrea Madotto, and Pascale Fung · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, et al · 2022
Cited alongside, same era.
Hidden trigger backdoor attack on nlp models via linguistic style manipulation
Xudong Pan, Mi Zhang, Beina Sheng, Jiaming Zhu, and Min Yang · 2022
Cited alongside, same era.
Red teaming language models with language models
Ethan Perez, Saffron Huang, H. Francis Song, et al · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamile Lukosiute, et al · 2022
Cited alongside, same era.
Transformer grammars: Augmenting transformer language models with syntactic inductive biases at scale
Laurent Sartran, Samuel Barrett, Adhiguna Kuncoro, Milovs Stanojevi’c, Phil Blunsom, and Chris Dyer · 2022
Cited alongside, same era.
The moral integrity corpus: A benchmark for ethical dialogue systems
Caleb Ziems, Jane A. Yu, Yi-Chia Wang, Alon Y. Halevy, and Diyi Yang · 2022
Cited alongside, same era.
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu · 2023
Closest in time.
A sentence is worth a thousand pictures: Can large language models understand human language?, 2023
Gary Marcus, Evelina Leivada, and Elliot Murphy · 2023
Closest in time.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, et al · 2023
Closest in time.
Xinyu Shen, Zeyuan Johnson Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Closest in time.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou · 2023
Closest in time.
Safety assessment of chinese large language models
Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, and Minlie Huang · 2023
Closest in time.
Moss: Training conversational language models from synthetic data
Tianxiang Sun, Xiaotian Zhang, Zhengfu He, et al · 2023
Closest in time.
Chatgpt is fun, but not an author
H. Holden Thorp · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin R. Stone, et al · 2023
Closest in time.
Do-not-answer: A dataset for evaluating safeguards in llms
Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin · 2023
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Closest in time.
Cvalues: Measuring the values of chinese large language models from safety to responsibility
Guohai Xu, Jiayi Liu, Mingshi Yan, Haotian Xu, Jinghui Si, Zhuoran Zhou, Peng Yi, Xing Gao, Jitao Sang, Rong Zhang, Ji Zhang, Chao Peng, Feiyan Huang, and Jingren Zhou · 2023
Closest in time.
Large language models as optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen · 2023
Closest in time.
Gptfuzzer : Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, and Xinyu Xing · 2023
Closest in time.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Weirong Ye, Neil Zhenqiang Gong, Yue Zhang, and Xingxu Xie · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Closest in time.