Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are trained on vast, uncurated datasets that contain various forms of biases and language reinforcing harmful stereotypes that may be subsequently inherited by the models themselves.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
Bender, E. M.; Gebru, T.; McMillan-Major, A.; and Shmitchell, S. 2021 · 2021
Earlier work this paper cites.
Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets
Blodgett, S. L.; Lopez, G.; Olteanu, A.; Sim, R.; and Wallach, H. 2021 · 2021
Earlier work this paper cites.
Quantifying Social Biases in NLP: A Generalization and Empirical Comparison of Extrinsic Fairness Metrics
Czarnowska, P.; Vyas, Y.; and Shah, K. 2021 · 2021
Earlier work this paper cites.
Towards understanding and mitigating social biases in language models
Liang, P. P.; Wu, C.; Morency, L.-P.; and Salakhutdinov, R. 2021 · 2021
Earlier work this paper cites.
A survey on bias and fairness in machine learning
Mehrabi, N.; Morstatter, F.; Saxena, N.; Lerman, K.; and Galstyan, A. 2021 · 2021
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Nadeem, M.; Bethke, A.; and Reddy, S. 2021 · 2021
Earlier work this paper cites.
Recipes for Building an Open-Domain Chatbot
Roller, S.; Dinan, E.; Goyal, N.; Ju, D.; Williamson, M.; Liu, Y.; Xu, J.; Ott, M.; Smith, E. M.; Boureau, Y.-L.; and Weston, J. 2021 · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Weidinger, L.; Mellor, J.; Rauh, M.; Griffin, C.; Uesato, J.; Huang, P.-S.; Cheng, M.; Glaese, M.; Balle, B.; Kasirzadeh, A.; et al. 2021 · 2021
Earlier work this paper cites.
SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing
Yu, T.; Zhang, R.; Polozov, A.; Meek, C.; and Awadallah, A. H. 2021 · 2021
Earlier work this paper cites.
Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models
Cao, Y. T.; Sotnikova, A.; Daumé III, H.; Rudinger, R.; and Zou, L. 2022 · 2022
Earlier work this paper cites.
Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models
Delobelle, P.; Tokpo, E.; Calders, T.; and Berendt, B. 2022 · 2022
Earlier work this paper cites.
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Ganguli, D.; Lovitt, L.; Kernion, J.; Askell, A.; Bai, Y.; Kadavath, S.; Mann, B.; Perez, E.; Schiefer, N.; Ndousse, K.; Jones, A.; Bowman, S.; Chen, A.; Conerly, T.; DasSarma, N.; Drain, D.; Elhage, N.; El-Showk, S.; Fort, S.; Hatfield-Dodds, Z.; Henighan, T.; Hernandez, D.; Hume, T.; Jacobson, J.; Johnston, S.; Kravec, S.; Olsson, C.; Ringer, S.; Tran-Johnson, E.; Amodei, D.; Brown, T.; Joseph, N.; McCandlish, S.; Olah, C.; Kaplan, J.; and Clark, J. 2022 · 2022
Earlier work this paper cites.
Maieutic prompting: Logically consistent reasoning with recursive explanations
Jung, J.; Qin, L.; Welleck, S.; Brahman, F.; Bhagavatula, C.; Bras, R. L.; and Choi, Y. 2022 · 2022
Cited alongside, same era.
Holistic Evaluation of Language Models
Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; Newman, B.; Yuan, B.; Yan, B.; Zhang, C.; Cosgrove, C.; Manning, C. D.; Ré, C.; Acosta-Navas, D.; Hudson, D. A.; Zelikman, E.; Durmus, E.; Ladhak, F.; Rong, F.; Ren, H.; Yao, H.; Wang, J.; Santhanam, K.; Orr, L.; Zheng, L.; Yuksekgonul, M.; Suzgun, M.; Kim, N.; Guha, N.; Chatterji, N.; Khattab, O.; Henderson, P.; Huang, Q.; Chi, R.; Xie, S. M.; Santurkar, S.; Ganguli, S.; Hashimoto, T.; Icard, T.; Zhang, T.; Chaudhary, V.; Wang, W.; Li, X.; Mai, Y.; Zhang, Y.; and Koreeda, Y. 2022 · 2022
Cited alongside, same era.
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
Min, S.; Lyu, X.; Holtzman, A.; Artetxe, M.; Lewis, M.; Hajishirzi, H.; and Zettlemoyer, L. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E. H.; Le, Q. V.; and Zhou, D. 2022 · 2022
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; de las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; Lavaud, L. R.; Lachaux, M.-A.; Stock, P.; Scao, T. L.; Lavril, T.; Wang, T.; Lacroix, T.; and Sayed, W. E. 2023 · 2023
Closest in time.
ChatGPT for good? On opportunities and challenges of large language models for education
Kasneci, E.; Seßler, K.; Küchemann, S.; Bannert, M.; Dementieva, D.; Fischer, F.; Gasser, U.; Groh, G.; Günnemann, S.; Hüllermeier, E.; et al. 2023 · 2023
Closest in time.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Kwon, W.; Li, Z.; Zhuang, S.; Sheng, Y.; Zheng, L.; Yu, C. H.; Gonzalez, J. E.; Zhang, H.; and Stoica, I. 2023 · 2023
Closest in time.
RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Lee, H.; Phatale, S.; Mansoor, H.; Mesnard, T.; Ferret, J.; Lu, K.; Bishop, C.; Hall, E.; Carbune, V.; Rastogi, A.; and Prakash, S. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
OPT: Open Pre-trained Transformer Language Models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; Mihaylov, T.; Ott, M.; Shleifer, S.; Shuster, K.; Simig, D.; Koura, P. S.; Sridhar, A.; Wang, T.; and Zettlemoyer, L. 2022 · 2022
Cited alongside, same era.
Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models
Cheng, M.; Durmus, E.; and Jurafsky, D. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Cited alongside, same era.
Can Instruction Fine-Tuned Language Models Identify Social Bias through Prompting?
Dige, O.; Tian, J.-J.; Emerson, D. B.; and Khattak, F. K. 2023 · 2023
Cited alongside, same era.
“So what if ChatGPT wrote it?” Multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy
Dwivedi, Y.; Kshetri, N.; Hughes, L.; Slade, E.; Jeyaraj, A.; Kar, A.; Baabdullah, A.; Koohang, A.; Raghavan, V.; Ahuja, M.; Albanna, H.; Albashrawi, M.; Al-Busaidi, A.; Balakrishnan, J.; Barlette, Y.; Basu, S.; Bose, I.; Brooks, L.; Buhalis, D.; Carter, L.; Chowdhury, S.; Crick, T.; Cunningham, S.; Davies, G.; Davison, R.; Dé, R.; Dennehy, D.; Duan, Y.; Dubey, R.; Dwivedi, R.; Edwards, J.; Flavián, C.; Gauld, R.; Grover, V.; Hu, M.; Janssen, M.; Jones, P.; Junglas, I.; Khorana, S.; Kraus, S.; Larsen, K.; Latreille, P.; Laumer, S.; Malik, F.; Mardani, A.; Mariani, M.; Mithas, S.; Mogaji, E.; Nord, J.; O’Connor, S.; Okumus, F.; Pagani, M.; Pandey, N.; Papagiannidis, S.; Pappas, I.; Pathak, N.; Pries-Heje, J.; Raman, R.; Rana, N.; Rehm, S.; Ribeiro-Navarrete, S.; Richter, A.; Rowe, F.; Sarker, S.; Stahl, B.; Tiwari, M.; van der Aalst, W.; Venkatesh, V.; Viglia, G.; Wade, M.; Walton, P.; Wirtz, J.; and Wright, R. 2023 · 2023
Cited alongside, same era.
The capacity for moral self-correction in large language models
Ganguli, D.; Askell, A.; Schiefer, N.; Liao, T. I.; Lukošiūtė, K.; Chen, A.; Goldie, A.; Mirhoseini, A.; Olsson, C.; and Hernandez, D. 2023 · 2023
Cited alongside, same era.
Towards Reasoning in Large Language Models: A Survey
Huang, J.; and Chang, K. C.-C. 2023 · 2023
Cited alongside, same era.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; Joseph, N.; Kadavath, S.; Kernion, J.; Conerly, T.; El-Showk, S.; Elhage, N.; Hatfield-Dodds, Z.; Hernandez, D.; Hume, T.; Johnston, S.; Kravec, S.; Lovitt, L.; Nanda, N.; Olsson, C.; Amodei, D.; Brown, T.; Clark, J.; McCandlish, S.; Olah, C.; Mann, B.; and Kaplan, J. 2022a
Cited in the paper.
Mökander, J.; Schuett, J.; Kirk, H. R.; and Floridi, L. 2023 · 2023
Closest in time.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Srivastava, A.; Rastogi, A.; Rao, A.; Shoeb, A. A. M.; Abid, A.; Fisch, A.; Brown, A. R.; Santoro, A.; Gupta, A.; Garriga-Alonso, A.; et al. 2023 · 2023
Closest in time.
Soft-prompt Tuning for Large Language Models to Evaluate Bias
Tian, J.-J.; Emerson, D.; Miyandoab, S. Z.; Pandya, D.; Seyyed-Kalantari, L.; and Khattak, F. K. 2023 · 2023
Closest in time.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q. V.; Chi, E. H.; Narang, S.; Chowdhery, A.; and Zhou, D. 2023 · 2023
Closest in time.
Jiang, A. Q.; Sablayrolles, A.; Roux, A.; Mensch, A.; Savary, B.; Bamford, C.; Chaplot, D. S.; de las Casas, D.; Hanna, E. B.; Bressand, F.; Lengyel, G.; Bour, G.; Lample, G.; Lavaud, L. R.; Saulnier, L.; Lachaux, M.-A.; Stock, P.; Subramanian, S.; Yang, S.; Antoniak, S.; Scao, T. L.; Gervet, T.; Lavril, T.; Wang, T.; Lacroix, T.; and Sayed, W. E. 2024 · 2024
Closest in time.
Large language models are zero-shot reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2024 · 2024
Closest in time.
BBQ: A hand-built bias benchmark for question answering
Parrish, A.; Chen, A.; Nangia, N.; Padmakumar, V.; Phang, J.; Thompson, J.; Htut, P. M.; and Bowman, S. 2022 · 2086
Closest in time.