Fetching the paper…
Reading the bibliography…
Standard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for desirable generations.
Text processing like humans do: Visually attacking and shielding nlp systems
Steffen Eger, Gözde Gül Şahin, Andreas Rücklé, Ji-Ung Lee, Claudia Schulz, Mohsen Mesgar, Krishnkant Swarnkar, Edwin Simpson, and Iryna Gurevych · 1903
Earlier work this paper cites.
Paws: Paraphrase adversaries from word scrambling
Yuan Zhang, Jason Baldridge, and Luheng He · 1904
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
Jiazhu Dai, Chuanshuai Chen, and Yufeng Li · 1905
Earlier work this paper cites.
Does data augmentation lead to positive margin?
Shashank Rajput, Zhili Feng, Zachary Charles, Po-Ling Loh, and Dimitris Papailiopoulos · 1905
Earlier work this paper cites.
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits · 1907
Earlier work this paper cites.
Achieving verified robustness to symbol substitutions via interval bound propagation
Po-Sen Huang, Robert Stanforth, Johannes Welbl, Chris Dyer, Dani Yogatama, Sven Gowal, Krishnamurthy Dvijotham, and Pushmeet Kohli · 1909
Earlier work this paper cites.
Certified robustness to adversarial word substitutions
Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang · 1909
Earlier work this paper cites.
Natural language adversarial defense through synonym encoding
Xiaosen Wang, Jin Hao, Yichen Yang, and Kun He · 1909
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
Wordnet: a lexical database for english
George Miller · 1995
Earlier work this paper cites.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al · 2004
Earlier work this paper cites.
Frequency-guided word substitutions for detecting textual adversarial examples
Maximilian Mozes, Pontus Stenetorp, Bennett Kleinberg, and Lewis Griffin · 2004
Earlier work this paper cites.
Badnl: Backdoor attacks against NLP models
Xiaoyi Chen, Ahmed Salem, Michael Backes, Shiqing Ma, and Yang Zhang · 2006
Earlier work this paper cites.
Contextual diversity for active learning
Sharat Agarwal, Himanshu Arora, Saket Anand, and Chetan Arora · 2008
Earlier work this paper cites.
Re-evaluating evaluation in text summarization
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig · 2010
Earlier work this paper cites.
Super-samples from kernel herding
Yutian Chen, Max Welling, and Alex Smola · 2010
Earlier work this paper cites.
A sweet rabbit hole by darcy: Using honeypots to detect universal trigger’s adversarial attacks
Thai Le, Noseong Park, and Dongwon Lee · 2011
Earlier work this paper cites.
Ppdb: The paraphrase database
Juri Ganitkevitch, Benjamin Van Durme, and Chris Callison-Burch · 2013
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Generating natural language adversarial examples
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani B Srivastava, and Kai-Wei Chang · 2018
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk · 2018
Cited alongside, same era.
Black-box generation of adversarial text sequences to evade deep learning classifiers
Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi · 2018
Cited alongside, same era.
Categorizing variants of goodhart’s law
David Manheim and Scott Garrabrant · 2018
Cited alongside, same era.
Semantically equivalent adversarial rules for debugging nlp models
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Cited alongside, same era.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese · 2018
Enriching a model’s notion of belief using a persistent memory
Nora Kassner, Oyvind Tafjord, Hinrich Schutze, and Peter Clark · 2021
Later among the works it cites.
WebGPT: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Later among the works it cites.
Onion: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun · 2021
Later among the works it cites.
Crafting adversarial examples for neural machine translation
Xinze Zhang, Junzhe Zhang, Zhenhua Chen, and Kun He · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Synthetic QA corpora generation with roundtrip consistency
Chris Alberti, Daniel Andor, Emily Pitler, Jacob Devlin, and Michael Collins · 2019
Cited alongside, same era.
Be consistent! improving procedural text comprehension using label consistency
Xinya Du, Bhavana Dalvi Mishra, Niket Tandon, Antoine Bosselut, Wen-tau Yih, Peter Clark, and Claire Cardie · 2019
Cited alongside, same era.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
NLP augmentation, 2019
Edward Ma · 2019
Cited alongside, same era.
Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham · 2019
Cited alongside, same era.
Generating natural language adversarial examples through probability weighted word saliency
Shuhuai Ren, Yihe Deng, Kun He, and Wanxiang Che · 2019
Cited alongside, same era.
Are red roses red? evaluating consistency of question-answering models
Marco Túlio Ribeiro, Carlos Guestrin, and Sameer Singh · 2019
Cited alongside, same era.
Leo Gao, John Schulman, and Jacob Hilton · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trbacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al · 2022
Later among the works it cites.
Quark: controllable text generation with reinforced unlearning
Ximing Lu, Sean Welleck, Jack Hessel, Liwei Jiang, Lianhui Qin, Peter West, Prithviraj Ammanabrolu, and Yejin Choi · 2022
Later among the works it cites.
Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi · 2022
Later among the works it cites.
Enhancing self-consistency and performance of pre-trained language models through natural language inference
Eric Mitchell, Joseph J. Noh, Siyan Li, William S. Armstrong, Ananth Agarwal, Patrick Liu, Chelsea Finn, and Christopher D. Manning · 2022
Later among the works it cites.
Training Language Models to Follow Instructions with Human Feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Stackllama: An rl fine-tuned llama model for stack exchange question and answering, 2023
Edward Beeching, Younes Belkada, Kashif Rasul, Lewis Tunstall, Leandro von Werra, Nazneen Rajani, and Nathan Lambert · 2023
Closest in time.
Raft: reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang · 2023
Closest in time.
Aligning language models with preferences through f-divergence minimization
Dongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen, Nahyeon Ryu, and Marc Dymetman · 2023
Closest in time.
Few-shot preference learning for human-in-the-loop RL
Donald Joseph Hejna III and Dorsa Sadigh · 2023
Closest in time.
Pretraining language models with human preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez · 2023
Closest in time.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi · 2023
Closest in time.
Expertqa: Expert-curated questions and attributed answers
Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Textshield: Beyond successfully detecting adversarial sentences in text classification
Lingfeng Shen, Ze Zhang, Haiyun Jiang, and Ying Chen · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
Banghua Zhu, Jiantao Jiao, and Michael Jordan · 2023
Closest in time.