Fetching the paper…
Reading the bibliography…
Language serves as a powerful tool for the manifestation of societal belief systems.
Best-worst scaling: A model for the largest difference judgments. working paper
J. J. Louviere. 1991 · 1991
Earlier work this paper cites.
Questions and Answers in Attitude Surveys: Experiments on Question Form, Wording, and Context
Stanley Presser and Howard Schuman. 1996 · 1996
Earlier work this paper cites.
Fightin’words: Lexical feature selection and evaluation for identifying the content of political conflict
Burt L Monroe, Michael P Colaresi, and Kevin M Quinn. 2008 · 2008
Earlier work this paper cites.
Maxdiff analysis: Simple counting,individual-level logit, and hb. sawtooth software, inc
B. Orme. 2009 · 2009
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011 · 2011
Earlier work this paper cites.
Semeval-2012 task 2: Measuring degrees of relational similarity
David A. Jurgens, Peter D. Turney, Saif M. Mohammad, and Keith J. Holyoak. 2012 · 2012
Earlier work this paper cites.
Embracing ambiguity: A comparison of annotation methodologies for crowdsourcing word sense labels
David Jurgens. 2013 · 2013
Earlier work this paper cites.
Best-worst scaling: theory and methods
T.N. Flynn and A.A.J. Marley. 2014 · 2014
Earlier work this paper cites.
Sentiment analysis of short informal text
Svetlana Kiritchenko, Xiaodan Zhu, and Saif Mohammad. 2014 · 2014
Earlier work this paper cites.
Best-Worst Scaling: Theory, Methods and Applications
Jordan J. Louviere, Terry N. Flynn, and A. A. J. Marley. 2015 · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
Capturing reliable fine-grained sentiment associations by crowdsourcing and best–worst scaling
Svetlana Kiritchenko and Saif M. Mohammad. 2016 · 2016
Earlier work this paper cites.
Commercial content moderation: Digital laborers’ dirty work
Sarah T Roberts. 2016 · 2016
Earlier work this paper cites.
Deceiving google’s perspective api built for detecting toxic comments
Hossein Hosseini, Sreeram Kannan, Baosen Zhang, and Radha Poovendran. 2017 · 2017
Earlier work this paper cites.
Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation
Svetlana Kiritchenko and Saif Mohammad. 2017 · 2017
Earlier work this paper cites.
Emotion intensities in tweets
Saif Mohammad and Felipe Bravo-Marquez. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
Understanding emotions: A dataset of tweets to study interactions between affect categories
Saif Mohammad and Svetlana Kiritchenko. 2018 · 2018
Cited alongside, same era.
Mind the gap: A balanced corpus of gendered ambiguous pronouns
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge. 2018 · 2018
Cited alongside, same era.
Big BiRD: A large, fine-grained, bigram relatedness dataset for examining semantic composition
Shima Asaadi, Saif Mohammad, and Svetlana Kiritchenko. 2019 · 2019
Cited alongside, same era.
Good secretaries, bad truck drivers? occupational gender stereotypes in sentiment analysis
Jayadev Bhaskaran and Isha Bhallamudi. 2019 · 2019
Cited alongside, same era.
On measuring gender bias in translation of gender-neutral pronouns
Won Ik Cho, Ji Won Kim, Seok Min Kim, and Nam Soo Kim. 2019 · 2019
Cited alongside, same era.
Ruddit: Norms of offensiveness for English Reddit comments
Rishav Hada, Sohi Sudhir, Pushkar Mishra, Helen Yannakoudakis, Saif M. Mohammad, and Ekaterina Shutova. 2021 · 2021
Later among the works it cites.
Collecting a large-scale gender bias dataset for coreference resolution and machine translation
Shahar Levy, Koren Lazar, and Gabriel Stanovsky. 2021 · 2021
Later among the works it cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Later among the works it cites.
The fabrics of machine moderation: Studying the technical, normative, and organizational structure of perspective api
Bernhard Rieder and Yarden Skop. 2021 · 2021
Later among the works it cites.
A survey on gender bias in natural language processing
Karolina Stanczak and Isabelle Augenstein. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
The knowref coreference corpus: Removing gender and number cues for difficult pronominal anaphora resolution
Ali Emami, Paul Trichelair, Adam Trischler, Kaheer Suleman, Hannes Schulz, and Jackie Chi Kit Cheung. 2019 · 2019
Cited alongside, same era.
Mitigating gender bias in natural language processing: Literature review
Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. 2019 · 2019
Cited alongside, same era.
Challenges and frontiers in abusive content detection
Bertie Vidgen, Alex Harris, Dong Nguyen, Rebekah Tromble, Scott Hale, and Helen Margetts. 2019 · 2019
Cited alongside, same era.
Language (technology) is power: A critical survey of “bias” in nlp
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Cited alongside, same era.
Convokit: A toolkit for the analysis of conversations
Jonathan P Chang, Caleb Chiam, Liye Fu, Andrew Wang, Justine Zhang, and Cristian Danescu-Niculescu-Mizil. 2020 · 2020
Cited alongside, same era.
Quantifying intimacy in language
Jiaxin Pei and David Jurgens. 2020 · 2020
Cited alongside, same era.
Analyzing the effects of annotator gender across nlp tasks
Laura Biester, Vanita Sharma, Ashkan Kazemi, Naihao Deng, Steven Wilson, and Rada Mihalcea. 2022 · 2022
Later among the works it cites.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022 · 2022
Later among the works it cites.
Two contrasting data annotation paradigms for subjective nlp tasks
Paul Röttger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022 · 2022
Later among the works it cites.
The lack of theory is painful: Modeling harshness in peer review comments
Rajeev Verma, Rajarshi Roychoudhury, and Tirthankar Ghosal. 2022 · 2022
Later among the works it cites.
Peering through preferences: Unraveling feedback acquisition for aligning large language models
Hritik Bansal, John Dang, and Aditya Grover. 2023 · 2023
Closest in time.
Tinystories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li. 2023 · 2023
Closest in time.
Do llms understand user preferences? evaluating llms on user rating prediction
Wang-Cheng Kang, Jianmo Ni, Nikhil Mehta, Maheswaran Sathiamoorthy, Lichan Hong, Ed Chi, and Derek Zhiyuan Cheng. 2023 · 2023
Closest in time.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023 · 2023
Closest in time.
Corgi-pm: A chinese corpus for gender bias probing and mitigation
Ge Zhang, Yizhi Li, Yaoyao Wu, Linyuan Zhang, Chenghua Lin, Jiayi Geng, Shi Wang, and Jie Fu. 2023 · 2023
Closest in time.
Cobra frames: Contextual reasoning about effects and harms of offensive statements
Xuhui Zhou, Hao Zhu, Akhila Yerukola, Thomas Davidson, Jena D Hwang, Swabha Swayamdipta, and Maarten Sap. 2023 · 2023
Closest in time.