A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep Reinforcement Learning from Human Preferences,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
S. Hisamoto, M. Post, and K. Duh, “Membership Inference Attacks on Sequence-to-Sequence Models: Is My Data in Your Machine Translation System?” Transactions of the Association for Computational Linguistics , vol. 8, pp. 49–63, 2020
2020
Earlier work this paper cites.
X. Pan, M. Zhang, S. Ji, and M. Yang, “Privacy Risks of General-Purpose Language Models,” in 2020 IEEE Symposium on Security and Privacy (SP) . IEEE, 2020, pp. 1314–1331
2020
Earlier work this paper cites.
J. Dhamala, T. Sun, V. Kumar, S. Krishna, Y. Pruksachatkun, K.-W. Chang, and R. Gupta, “BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation,” in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , 2021, pp. 862–872
2021
Earlier work this paper cites.
P. P. Liang, C. Wu, L.-P. Morency, and R. Salakhutdinov, “Towards Understanding and Mitigating Social Biases in Language Models,” in International Conference on Machine Learning . PMLR, 2021, pp. 6565–6576
2021
Earlier work this paper cites.
Y. Chen, C. Mahoney, I. Grasso, E. Wali, A. Matthews, T. Middleton, M. Njie, and J. Matthews, “Gender Bias and Under-Representation in Natural Language Processing Across Human Languages,” in Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society , 2021, pp. 24–34
2021
Earlier work this paper cites.
A. Tamkin, M. Brundage, J. Clark, and D. Ganguli, “Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models,” arXiv preprint arXiv:2102.02503 , 2021
Original
2021
Earlier work this paper cites.
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al. , “On the Opportunities and Risks of Foundation Models,” arXiv preprint arXiv:2108.07258 , 2021
Original
2021
Earlier work this paper cites.
W. I. Cho, J. Kim, J. Yang, and N. S. Kim, “Towards Cross-Lingual Generalization of Translation Gender Bias,” in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , 2021, pp. 449–457
2021
Earlier work this paper cites.
H. R. Kirk, Y. Jun, F. Volpin, H. Iqbal, E. Benussi, F. Dreyer, A. Shtedritski, and Y. Asano, “Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language Models,” Advances in Neural Information Processing Systems , vol. 34, pp. 2611–2624, 2021
2021
Earlier work this paper cites.
A. Xu, E. Pathak, E. Wallace, S. Gururangan, M. Sap, and D. Klein, “Detoxifying Language Models Risks Marginalizing Minority Voices,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 2390–2397
2021
Earlier work this paper cites.
H. Ngo, C. Raterink, J. G. Araújo, I. Zhang, C. Chen, A. Morisot, and N. Frosst, “Mitigating Harm in Language Models with Conditional-Likelihood Filtration,” arXiv preprint arXiv:2108.07790 , 2021
Original
2021
Earlier work this paper cites.
N. Carlini, F. Tramèr, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, Ú. Erlingsson, A. Oprea, and C. Raffel, “Extracting Training Data from Large Language Models,” in 30th USENIX Security Symposium (USENIX Security 21) . USENIX Association, Aug 2021, pp. 2633–2650
2021
Earlier work this paper cites.