Fetching the paper…
Reading the bibliography…
While automatic prompt generation methods have recently received significant attention, their robustness remains poorly understood.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 1909
Earlier work this paper cites.
Building a question answering test collection
Ellen M Voorhees and Dawn M Tice. 2000 · 2000
Earlier work this paper cites.
Mining and summarizing customer reviews
Minqing Hu and Bing Liu. 2004 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Fernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton, Benoît Sagot, and Lucia Specia. 2020 · 2005
Earlier work this paper cites.
Exploitingclassrelationshipsforsentimentcate gorizationwithrespectratingsales
L PaNgB. 2005 · 2005
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Optimizing statistical machine translation for text simplification
Wei Xu, Courtney Napoles, Ellie Pavlick, Quanze Chen, and Chris Callison-Burch. 2016 · 2016
Earlier work this paper cites.
Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
A sentence-level text adversarial attack algorithm against iiot based smart grid
Jialiang Dong, Zhitao Guan, Longfei Wu, Xiaojiang Du, and Mohsen Guizani. 2021 · 2021
Earlier work this paper cites.
Cross-task generalization via natural language crowdsourcing instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021 · 2021
Earlier work this paper cites.
Evaluating the robustness of neural language models to input perturbations
Milad Moradi and Matthias Samwald. 2021 · 2021
Earlier work this paper cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021 · 2021
Earlier work this paper cites.
Rlprompt: Optimizing discrete text prompts with reinforcement learning
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric P Xing, and Zhiting Hu. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Automatic chain of thought prompting in large language models
Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022 · 2022
Cited alongside, same era.
Large language models are human-level prompt engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022 · 2022
Cited alongside, same era.
Generative ai text classification using ensemble llm approaches
Harika Abburi, Michael Suesserman, Nirmala Pudota, Balaji Veeramani, Edward Bowen, and Sanmitra Bhattacharya. 2023 · 2023
Cited alongside, same era.
Instructzero: Efficient instruction optimization for black-box large language models
An llm can fool itself: A prompt-based adversarial attack
Xilie Xu, Keyi Kong, Ning Liu, Lizhen Cui, Di Wang, Jingfeng Zhang, and Mohan Kankanhalli. 2023 · 2023
Later among the works it cites.
Llm lies: Hallucinations are not bugs, but features as adversarial examples
Jia-Yu Yao, Kun-Peng Ning, Zhen-Hui Liu, Mu-Nan Ning, Yu-Yang Liu, and Li Yuan. 2023 · 2023
Later among the works it cites.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Yue Zhang, Neil Zhenqiang Gong, et al. 2023 · 2023
Later among the works it cites.
Openagi: When llm meets domain experts
Yingqiang Ge, Wenyue Hua, Kai Mei, Juntao Tan, Shuyuan Xu, Zelong Li, Yongfeng Zhang, et al. 2024 · 2024
Closest in time.
The impact of reasoning step length on large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lichang Chen, Jiuhai Chen, Tom Goldstein, Heng Huang, and Tianyi Zhou. 2023 · 2023
Cited alongside, same era.
Revisit input perturbation problems for llms: A unified robustness evaluation framework for noisy slot filling task
Guanting Dong, Jinxu Zhao, Tingfeng Hui, Daichi Guo, Wenlong Wang, Boqi Feng, Yueyan Qiu, Zhuoma Gongque, Keqing He, Zechen Wang, et al. 2023 · 2023
Cited alongside, same era.
Towards sentence level inference attack against pre-trained language models
Kang Gu, Ehsanul Kabir, Neha Ramsurrun, Soroush Vosoughi, and Shagufta Mehnaz. 2023 · 2023
Cited alongside, same era.
Connecting large language models with evolutionary algorithms yields powerful prompt optimizers
Qingyan Guo, Rui Wang, Junliang Guo, Bei Li, Kaitao Song, Xu Tan, Guoqing Liu, Jiang Bian, and Yujiu Yang. 2023 · 2023
Cited alongside, same era.
Certifying llm safety against adversarial prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li, Soheil Feizi, and Himabindu Lakkaraju. 2023 · 2023
Cited alongside, same era.
Summedits: measuring llm ability at factual reasoning through the lens of summarization
Philippe Laban, Wojciech Kryściński, Divyansh Agarwal, Alexander Richard Fabbri, Caiming Xiong, Shafiq Joty, and Chien-Sheng Wu. 2023 · 2023
Cited alongside, same era.
Tuna: Instruction tuning using feedback from large language models
Haoran Li, Yiran Liu, Xingxing Zhang, Wei Lu, and Furu Wei. 2023 · 2023
Cited alongside, same era.
Eureka: Human-level reward design via coding large language models
Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023 · 2023
Cited alongside, same era.
Mingyu Jin, Qinkai Yu, Dong Shu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, and Mengnan Du. 2024 · 2024
Closest in time.
Large language model sentinel: Advancing adversarial robustness by llm agent
Guang Lin and Qibin Zhao. 2024 · 2024
Closest in time.
Large language models as evolutionary optimizers
Shengcai Liu, Caishun Chen, Xinghua Qu, Ke Tang, and Yew-Soon Ong. 2024 · 2024
Closest in time.
NoiseBench: Benchmarking the impact of real label noise on named entity recognition
Elena Merdjanovska, Ansar Aynetdinov, and Alan Akbik. 2024 · 2024
Closest in time.
Autonomous llm-enhanced adversarial attack for text-to-motion
Honglei Miao, Fan Ma, Ruijie Quan, Kun Zhan, and Yi Yang. 2024 · 2024
Closest in time.
Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot llm assessment
Vyas Raina, Adian Liusie, and Mark Gales. 2024 · 2024
Closest in time.
Targeted latent adversarial training improves robustness to persistent harmful behaviors in llms
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, et al. 2024 · 2024
Closest in time.
Efficient adversarial training in llms with continuous attacks
Sophie Xhonneux, Alessandro Sordoni, Stephan Günnemann, Gauthier Gidel, and Leo Schwinn. 2024 · 2024
Closest in time.
An LLM can fool itself: A prompt-based adversarial attack
Xilie Xu, Keyi Kong, Ning Liu, Lizhen Cui, Di Wang, Jingfeng Zhang, and Mohan Kankanhalli. 2024 · 2024
Closest in time.
Evaluating the validity of word-level adversarial attacks with large language models
Huichi Zhou, Zhaoyang Wang, Hongtao Wang, Dongping Chen, Wenhan Mu, and Fangyuan Zhang. 2024b · 2024
Closest in time.