Fetching the paper…
Reading the bibliography…
Text detoxification, a variant of style transfer tasks, finds useful applications in online social media.
Linear hinge loss and average margin
Claudio Gentile and Manfred K. K Warmuth. 1998 · 1998
Earlier work this paper cites.
Toxic comment classification challenge
cjadams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, nithum, and Will Cukierski. 2017 · 2017
Earlier work this paper cites.
Style transfer in text: Exploration and evaluation
Zhenxin Fu, Xiaoye Tan, Nanyun Peng, Dongyan Zhao, and Rui Yan. 2018 · 2018
Earlier work this paper cites.
Delete, retrieve, generate: a simple approach to sentiment and style transfer
Juncen Li, Robin Jia, He He, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Fighting offensive language on social media with unsupervised text style transfer
Cicero Nogueira dos Santos, Igor Melnyk, and Inkit Padhi. 2018 · 2018
Earlier work this paper cites.
Jigsaw unintended bias in toxicity classification
cjadams, Daniel Borkan, inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and nithum. 2019 · 2019
Earlier work this paper cites.
A dual reinforcement learning framework for unsupervised text style transfer
Fuli Luo, Peng Li, Jie Zhou, Pengcheng Yang, Baobao Chang, Xu Sun, and Zhifang Sui. 2019 · 2019
Earlier work this paper cites.
Neural Network Acceptability Judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
Beyond BLEU: Training neural machine translation with semantic similarity
John Wieting, Taylor Berg-Kirkpatrick, Kevin Gimpel, and Graham Neubig. 2019 · 2019
Earlier work this paper cites.
Mask and infill: Applying masked language model for sentiment transfer
Xing Wu, Tao Zhang, Liangjun Zang, Jizhong Han, and Songlin Hu. 2019 · 2019
Earlier work this paper cites.
A probabilistic formulation of unsupervised text style transfer
Junxian He, Xinyi Wang, Graham Neubig, and Taylor Berg-Kirkpatrick. 2020 · 2020
Cited alongside, same era.
Jigsaw multilingual toxic comment classification
Ian Kivlichan, Jeffrey Sorensen, Lucy Vasserman Julia Elliott, Martin Görner, and Phil Culliton. 2020 · 2020
Cited alongside, same era.
Stable style transformer: Delete and generate approach with encoder-decoder for text style transfer
Joosung Lee. 2020 · 2020
Cited alongside, same era.
Towards a friendly online community: An unsupervised style transfer framework for profanity redaction
Minh Tran, Yipeng Zhang, and Mohammad Soleymani. 2020 · 2020
Cited alongside, same era.
Text detoxification using large pre-trained neural models
David Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva, Olga Kozlova, Nikita Semenov, and Alexander Panchenko. 2021 · 2021
Cited alongside, same era.
Overview of pan 2024: Multi-author writing style analysis, multilingual text detoxification, oppositional thinking analysis, and generative ai authorship verification: Extended abstract
Janek Bevendorff, Xavier Bonet Casals, Berta Chulvi, Daryna Dementieva, Ashaf Elnagar, Dayne Freitag, Maik Fröbe, Damir Korenčić, Maximilian Mayerl, Animesh Mukherjee, Alexander Panchenko, Martin Potthast, Francisco Rangel, Paolo Rosso, Alisa Smirnova, Efstathios Stamatatos, Benno Stein, Mariona Taulé, Dmitry Ustalov, Matti Wiegmann, and Eva Zangerle. 2024 · 2024
Closest in time.
Parl: A unified framework for policy alignment in reinforcement learning from human feedback
Souradip Chakraborty, Amrit Singh Bedi, Alec Koppel, Dinesh Manocha, Huazheng Wang, Mengdi Wang, and Furong Huang. 2024 · 2024
Closest in time.
Self-play fine-tuning converts weak language models to strong language models
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu. 2024 · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Léo Laugier, John Pavlopoulos, Jeffrey Sorensen, and Lucas Dixon. 2021 · 2021
Cited alongside, same era.
ParaDetox: Detoxification with parallel data
Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, and Alexander Panchenko. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos. 2023 · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, John A. Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuan-Fang Li, Scott M. Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023 · 2023
Cited alongside, same era.
The Confluence of Networks, Games, and Learning a Game-Theoretic Framework for Multiagent Decision Making Over Networks
Tao Li, Guanze Peng, Quanyan Zhu, and Tamer Baar. 2022a
Cited in the paper.
The role of information structures in game-theoretic multi-agent learning
Tao Li, Yuhan Zhao, and Quanyan Zhu. 2022b
Cited in the paper.
Closest in time.
Symbiotic game and foundation models for cyber deception operations in strategic cyber warfare
Tao Li and Quanyan Zhu. 2024 · 2024
Closest in time.
Nash learning from human feedback
Rémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Zhaohan Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, and Bilal Piot. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2024 · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal. 2024 · 2024
Closest in time.
A systematic review of toxicity in large language models: Definitions, datasets, detectors, detoxification methods and challenges
Guillermo Villate-Castillo, Javier Del Ser Lorente, Borja Sanz Urquijo, et al. 2024 · 2024
Closest in time.