Fetching the paper…
Reading the bibliography…
Large language models produce human-like text that drive a growing number of applications.
Toward fairness in AI for people with disabilities: A research roadmap
Anhong Guo, Ece Kamar, Jennifer Wortman Vaughan, Hanna M. Wallach, and Meredith Ringel Morris · 1907
Earlier work this paper cites.
Marked and unmarked: A choice between unequals in semiotic structure
Linda R. Waugh · 1982
Earlier work this paper cites.
Black Feminist Thought: Knowledge, consciousness, and the politics of empowerment
Patricia Hill Collins · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Donald Martin Jr., Vinodkumar Prabhakaran, Jill Kuhlberg, Andrew Smart, and William S. Isaac · 2006
Earlier work this paper cites.
Decolonial AI: decolonial theory as sociotechnical foresight in artificial intelligence
Shakir Mohamed, Marie-Therese Png, and William Isaac · 2007
Earlier work this paper cites.
Sentiwordnet 3.0: an enhanced lexical resource for sentiment analysis and opinion mining
Stefano Baccianella, Andrea Esuli, and Fabrizio Sebastiani · 2010
Earlier work this paper cites.
On achieving and evaluating language-independence in nlp
Emily M Bender · 2011
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Su Lin Blodgett, Lisa Green, and Brendan O’Connor · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
Offensive political dog whistles: you know them when you hear them. or do you?, 2016
Ian Olasov · 2016
Earlier work this paper cites.
The problem with bias: Allocative versus representational harms in machine learning
Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach · 2017
Earlier work this paper cites.
Israel arrests palestinian because facebook translated ’good morning’ to ’attack them’
Yotam Berger · 2017
Earlier work this paper cites.
A dataset and classifier for recognizing social media English
Su Lin Blodgett, Johnny Wei, and Brendan O’Connor · 2017
Earlier work this paper cites.
The trouble with bias. keynote at neurips, 2017
Kate Crawford · 2017
Earlier work this paper cites.
Better discussions with imperfect machine learning models, September 2017
Jigsaw · 2017
Earlier work this paper cites.
A survey on hate speech detection using natural language processing
Anna Schmidt and Michael Wiegand · 2017
Earlier work this paper cites.
Ex machina: Personal attacks seen at scale
Ellery Wulczyn, Nithum Thain, and Lucas Dixon · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman · 2018
Earlier work this paper cites.
Twitter universal dependency parsing for african-american and mainstream american english
Su Lin Blodgett, Johnny Wei, and Brendan O’Connor · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Earlier work this paper cites.
The measure and mismeasure of fairness: A critical review of fair machine learning
Sam Corbett-Davies and Sharad Goel · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman · 2018
Earlier work this paper cites.
Gender recognition or gender reductionism? the social implications of embedded gender recognition systems
Foad Hamidi, Morgan Klaus Scheuerman, and Stacy M. Branham · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme · 2018
Earlier work this paper cites.
Crowdsourcing subjective tasks: the case study of understanding toxicity in online discussions
Lora Aroyo, Lucas Dixon, Nithum Thain, Olivia Redfield, and Rachel Rosen · 2019
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman · 2019
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Fairness-aware ranking in search & recommendation systems with application to linkedin talent search
Sahin Cem Geyik, Stuart Ambler, and Krishnaram Kenthapadi · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Earlier work this paper cites.
Dissecting racial bias in an algorithm that guides health decisions for 70 million people
Ziad Obermeyer and Sendhil Mullainathan · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Fairness and abstraction in sociotechnical systems
Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi · 2019
Cited alongside, same era.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng · 2019
Cited alongside, same era.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
The importance of modeling social factors of language: Theory and practice
Dirk Hovy and Diyi Yang · 2021
Later among the works it cites.
Towards accountability for machine learning datasets: Practices from software engineering and infrastructure
Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell · 2021
Later among the works it cites.
On transferability of bias mitigation effects in language model fine-tuning
Xisen Jin, Francesco Barbieri, Brendan Kennedy, Aida Mostafazadeh Davani, Leonardo Neves, and Xiang Ren · 2021
Later among the works it cites.
Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, and Geoffrey Irving · 2021
Later among the works it cites.
You keep using that word: Ways of thinking about gender in computing research
Os Keyes, Chandler May, and Annabelle Carrell · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Conversational assistants and gender stereotypes: Public perceptions and desiderata for voice personas
Amanda Cercas Curry, Judy Robertson, and Verena Rieser · 2020
Cited alongside, same era.
A snapshot of the frontiers of fairness in machine learning
Alexandra Chouldechova and Aaron Roth · 2020
Cited alongside, same era.
Opengpt-2: open language models and implications of generated text
Vanya Cohen and Aaron Gokaslan · 2020
Cited alongside, same era.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu · 2020
Cited alongside, same era.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Cited alongside, same era.
Investigating African-American Vernacular English in transformer-based text generation
Sophie Groenwold, Lily Ou, Aesha Parekh, Samhita Honnavalli, Sharon Levy, Diba Mirza, and William Yang Wang · 2020
Cited alongside, same era.
A distributional approach to controlled text generation
Muhammad Khalifa, Hady Elsahar, and Marc Dymetman · 2021
Later among the works it cites.
GeDi: Generative discriminator guided sequence generation, 2021
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Rajani · 2021
Later among the works it cites.
Pitfalls of static language modelling
Angeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Sebastian Ruder, Dani Yogatama, Kris Cao, Tomás Kociský, Susannah Young, and Phil Blunsom · 2021
Later among the works it cites.
Capturing covertly toxic speech via crowdsourcing
Alyssa Lees, Daniel Borkan, Ian Kivlichan, Jorge Nario, and Tesh Goyal · 2021
Later among the works it cites.
Towards understanding and mitigating social biases in language models
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2021
Later among the works it cites.
Jurassic-1: Technical details and evaluation
Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham · 2021
Later among the works it cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2021
Later among the works it cites.
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan · 2021
Later among the works it cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2021
Later among the works it cites.
Bbq: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, James Bradbury, Matthew Johnson, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving · 2021
Later among the works it cites.
Re-imagining algorithmic fairness in india and beyond
Nithya Sambasivan, Erin Arnesen, Ben Hutchinson, Tulsee Doshi, and Vinodkumar Prabhakaran · 2021
Later among the works it cites.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in NLP
Timo Schick, Sahana Udupa, and Hinrich Schütze · 2021
Later among the works it cites.
Societal biases in language generation: Progress and challenges
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng · 2021
Later among the works it cites.
It began as an ai-fueled dungeon game. it got much darker
Tom Simonite · 2021
Later among the works it cites.
Process for adapting language models to society (palms) with values-targeted datasets
Irene Solaiman and Christy Dennison · 2021
Later among the works it cites.
Reliability testing for natural language processing systems
Samson Tan, Shafiq Joty, Kathy Baxter, Araz Taeihagh, Gregory A Bennett, and Min-Yen Kan · 2021
Later among the works it cites.
Fairness for unobserved characteristics: Insights from technological impacts on queer communities
Nenad Tomasev, Kevin R. McKee, Jackie Kay, and Shakir Mohamed · 2021
Later among the works it cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2021
Later among the works it cites.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Later among the works it cites.
Toxicity detection can be sensitive to the conversational context
Alexandros Xenos, John Pavlopoulos, Ion Androutsopoulos, Lucas Dixon, Jeffrey Sorensen, and Leo Laugier · 2021
Later among the works it cites.
Detoxifying language models risks marginalizing minority voices
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein · 2021
Later among the works it cites.
Bot-adversarial dialogue for safe conversational agents
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
Handling and presenting harmful text, 2022
Leon Derczynski, Hannah Rose Kirk, Abeba Birhane, and Bertie Vidgen · 2022
Closest in time.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar · 2022
Closest in time.
Challenges and strategies in cross-cultural nlp
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, et al · 2022
Closest in time.
Anoop K., Manjary P. Gangan, Deepak P., and Lajish V. L · 2022
Closest in time.