Fetching the paper…
Reading the bibliography…
Language models trained on large-scale unfiltered datasets curated from the open web acquire systemic biases, prejudices, and harmful views from their training data.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2013
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context, 2016
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
The risk of racial bias in hate speech detection
M. Sap, D. Card, S. Gabriel, Y. Choi, and N. A. Smith · 2019
Earlier work this paper cites.
Challenges and frontiers in abusive content detection
B. Vidgen, A. Harris, D. Nguyen, R. Tromble, S. Hale, and H. Margetts · 2019
Cited alongside, same era.
Language models are few-shot learners, 2020
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models, 2020
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2020
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?, 2021
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Cited alongside, same era.
Documenting the english colossal clean crawled corpus, 2021
J. Dodge, M. Sap, A. Marasovic, W. Agnew, G. Ilharco, D. Groeneveld, and M. Gardner · 2021
Closest in time.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp, 2021
T. Schick, S. Udupa, and H. Schütze · 2021
Closest in time.
Process for adapting language models to society (palms) with values-targeted datasets
I. Solaiman and C. Dennison · 2021
Closest in time.
Universal adversarial triggers for attacking and analyzing nlp, 2021
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh · 2021
Closest in time.
Censorship of online encyclopedias: Implications for nlp models
E. Yang and M. E. Roberts · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…