Fetching the paper…
Reading the bibliography…
Language models have shown promise in various tasks but can be affected by undesired data during training, fine-tuning, or alignment.
Learning with noisy labels
Nagarajan Natarajan, Inderjit S Dhillon, Pradeep K Ravikumar, and Ambuj Tewari · 2013
Earlier work this paper cites.
Training deep neural networks on noisy labels with bootstrapping
Scott Reed, Honglak Lee, Dragomir Anguelov, Christian Szegedy, Dumitru Erhan, and Andrew Rabinovich · 2014
Earlier work this paper cites.
Classification with noisy labels by importance reweighting
Tongliang Liu and Dacheng Tao · 2015
Earlier work this paper cites.
Are you a racist or am I seeing things? annotator influence on hate speech detection on Twitter
Zeerak Waseem · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber · 2017
Earlier work this paper cites.
Learning from noisy labels with distillation
Yuncheng Li, Jianchao Yang, Yale Song, Liangliang Cao, Jiebo Luo, and Li-Jia Li · 2017
Earlier work this paper cites.
Learning with confident examples: Rank pruning for robust classification with noisy labels
Curtis G. Northcutt, Tailin Wu, and Isaac L. Chuang · 2017
Earlier work this paper cites.
Making deep neural networks robust to label noise: A loss correction approach
Giorgio Patrini, Alessandro Rozza, Aditya Krishna Menon, Richard Nock, and Lizhen Qu · 2017
Earlier work this paper cites.
Toward robustness against label noise in training deep discriminative neural networks
Arash Vahdat · 2017
Earlier work this paper cites.
Learning from noisy large-scale datasets with minimal supervision
Andreas Veit, Neil Alldrin, Gal Chechik, Ivan Krasin, Abhinav Gupta, and Serge Belongie · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Co-teaching: Robust training of deep neural networks with extremely noisy labels
Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, and Masashi Sugiyama · 2018
Earlier work this paper cites.
Ai safety via debate, 2018
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Challenges for toxic comment classification: An in-depth error analysis
Betty van Aken, Julian Risch, Ralf Krestel, and Alexander Löser · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant · 2019
Earlier work this paper cites.
Deep self-learning from noisy labels
Jiangfan Han, Ping Luo, and Xiaogang Wang · 2019
Earlier work this paper cites.
When Does Label Smoothing Help?
Rafael Müller, Simon Kornblith, and Geoffrey Hinton · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Selfie: Refurbishing unclean samples for robust deep learning
Hwanjun Song, Minseok Kim, and Jae-Gil Lee · 2019
Earlier work this paper cites.
How does disagreement help generalization against label corruption?
Xingrui Yu, Bo Han, Jiangchao Yao, Gang Niu, Ivor W Tsang, and Masashi Sugiyama · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
The Alignment Problem: Machine Learning and Human Values
Brian Christian · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Social biases in NLP models as barriers for persons with disabilities
Ben Hutchinson, Vinodkumar Prabhakaran, Emily Denton, Kellie Webster, Yu Zhong, and Stephen Denuyl · 2020
Earlier work this paper cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Cited alongside, same era.
Does label smoothing mitigate label noise?
Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, and Sanjiv Kumar · 2020
Cited alongside, same era.
The radicalization risks of gpt-3 and advanced neural language models, 2020
Kris McGuffie and Alex Newhouse · 2020
Cited alongside, same era.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman · 2020
Cited alongside, same era.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2020
Cited alongside, same era.
Illustrating reinforcement learning from human feedback (rlhf)
Nathan Lambert, Louis Castricato, Leandro von Werra, and Alex Havrilla · 2022
Later among the works it cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes, 2022
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, and Nat McAleese · 2022
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback, 2022
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Combating noisy labels by agreement: A joint training method with co-regularization
Hongxin Wei, Lei Feng, Xiangyu Chen, and Bo An · 2020
Cited alongside, same era.
Impact of politically biased data on hate speech classification
Maximilian Wich, Jan Bauer, and Georg Groh · 2020
Cited alongside, same era.
Part-dependent label noise: Towards instance-dependent label noise
Xiaobo Xia, Tongliang Liu, Bo Han, Nannan Wang, Mingming Gong, Haifeng Liu, Gang Niu, Dacheng Tao, and Masashi Sugiyama · 2020
Cited alongside, same era.
Large language models associate muslims with violence
Abubakar Abid, Maheen Farooqi, and James Zou · 2021
Cited alongside, same era.
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach · 2021
Cited alongside, same era.
Learning with instance-dependent label noise: A sample sieve approach
Hao Cheng, Zhaowei Zhu, Xingyu Li, Yifei Gong, Xing Sun, and Yang Liu · 2021
Cited alongside, same era.
Can cross entropy loss be robust to label noise?
Lei Feng, Senlin Shu, Zhuoyi Lin, Fengmao Lv, Li Li, and Bo An · 2021
Cited alongside, same era.
Later among the works it cites.
Assessing multilingual fairness in pre-trained multimodal representations
Jialu Wang, Yang Liu, and Xin Wang · 2022
Later among the works it cites.
Taxonomy of risks posed by language models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2022
Later among the works it cites.
Understanding and mitigating the label noise in pre-training on downstream tasks
Hao Chen, Jindong Wang, Ankit Shah, Ran Tao, Hongxin Wei, Xing Xie, Masashi Sugiyama, and Bhiksha Raj · 2023
Closest in time.
Mitigating memorization of noisy labels via regularization between representations
Hao Cheng, Zhaowei Zhu, Xing Sun, and Yang Liu · 2023
Closest in time.
Pku-beaver: Constrained value-aligned llm via safe rlhf
Juntao Dai, Xuehai Pan, Jiaming Ji, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Closest in time.
Candidate label set pruning: A data-centric perspective for deep partial-label learning
Shuo He, Chaojie Wang, Guowu Yang, and Lei Feng · 2023
Closest in time.
Our approach to alignment research, 2023
Jeffrey Wu Jan Leike, John Schulman · 2023
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Chi Zhang, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Closest in time.
The importance of human-labeled data in the era of llms
Yang Liu · 2023
Closest in time.
Identifiability of label noise transition matrix
Yang Liu, Hao Cheng, and Kun Zhang · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
T2IAT: Measuring valence and stereotypical biases in text-to-image generation
Jialu Wang, Xinyue Liu, Zonglin Di, Yang Liu, and Xin Wang · 2023
Closest in time.
Freeal: Towards human-free active learning in the era of large language models
Ruixuan Xiao, Yiwen Dong, Junbo Zhao, Runze Wu, Minmin Lin, Gang Chen, and Haobo Wang · 2023
Closest in time.
Parameter-efficient cross-lingual transfer of vision and language models via translation-based alignment
Zhen Zhang, Jialu Wang, and Xin Wang · 2023
Closest in time.
Weak proxies are sufficient and preferable for fairness with missing sensitive attributes
Zhaowei Zhu, Yuanshun Yao, Jiankai Sun, Hang Li, and Yang Liu · 2023
Closest in time.
Robusttsf: Towards theory and design of robust time series forecasting with anomalies
Hao Cheng, Qingsong Wen, Yang Liu, and Liang Sun · 2024
Closest in time.
Human-instruction-free llm self-alignment with limited samples
Hongyi Guo, Yuanshun Yao, Wei Shen, Jiaheng Wei, Xiaoying Zhang, Zhaoran Wang, and Yang Liu · 2024
Closest in time.
Fair classifiers without fair training: An influence-guided data sampling approach
Jinlong Pang, Jialu Wang, Zhaowei Zhu, Yuanshun Yao, Chen Qian, and Yang Liu · 2024
Closest in time.
Early stopping against label noise without validation data
Suqin Yuan, Lei Feng, and Tongliang Liu · 2024
Closest in time.