Fetching the paper…
Reading the bibliography…
Language models (LMs) automatically learn word embeddings during pre-training on language corpora.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Learning effective and interpretable semantic models using non-negative sparse embedding
Brian Murphy, Partha Talukdar, and Tom Mitchell. 2012 · 1950
Earlier work this paper cites.
RedditBias: A real-world resource for bias evaluation and debiasing of conversational language models
Soumya Barikeri, Anne Lauscher, Ivan Vulić, and Goran Glavaš. 2021 · 1955
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 1967
Earlier work this paper cites.
A maximization technique occurring in the statistical analysis of probabilistic functions of markov chains
Leonard E Baum, Ted Petrie, George Soules, and Norman Weiss. 1970 · 1970
Earlier work this paper cites.
Producing high-dimensional semantic spaces from lexical co-occurrence
Kevin Lund and Curt Burgess. 1996 · 1996
Earlier work this paper cites.
Reading tea leaves: How humans interpret topic models
Jonathan Chang, Sean Gerrish, Chong Wang, Jordan Boyd-Graber, and David Blei. 2009 · 2009
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Retrofitting word vectors to semantic lexicons
Manaal Faruqui, Jesse Dodge, Sujay K Jauhar, Chris Dyer, Eduard Hovy, and Noah A Smith. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016 · 2016
Earlier work this paper cites.
Word embedding calculus in meaningful ultradense subspaces
Sascha Rothe and Hinrich Schütze. 2016 · 2016
Earlier work this paper cites.
Elucidating conceptual properties from word embeddings
Kyoung-Rok Jang and Sung-Hyon Myaeng. 2017 · 2017
Earlier work this paper cites.
Rotated word vector representations and their interpretability
Sungjoon Park, JinYeong Bak, and Alice Oh. 2017 · 2017
Earlier work this paper cites.
Delete, retrieve, generate: a simple approach to sentiment and style transfer
Juncen Li, Robin Jia, He He, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018 · 2018
Earlier work this paper cites.
Semantic structure and interpretability of word embeddings
Lütfi Kerem Şenel, Ihsan Utlu, Veysel Yücesoy, Aykut Koc, and Tolga Cukur. 2018 · 2018
Earlier work this paper cites.
Analogies explained: Towards understanding word embeddings
Carl Allen and Timothy Hospedales. 2019 · 2019
Earlier work this paper cites.
Rotate king to get queen: Word relationships as orthogonal transformations in embedding space
Kawin Ethayarajh. 2019 · 2019
Earlier work this paper cites.
Openwebtext corpus (2019)
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. 2019 · 2019
Cited alongside, same era.
Word2Sense: Sparse interpretable word embeddings
Abhishek Panigrahi, Harsha Vardhan Simhadri, and Chiranjib Bhattacharyya. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2019 · 2019
Cited alongside, same era.
Queens are powerful too: Mitigating gender bias in dialogue generation
Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, and Jason Weston. 2020 · 2020
Cited alongside, same era.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2020
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2021 · 2021
Later among the works it cites.
Fudge: Controlled text generation with future discriminators
Kevin Yang and Dan Klein. 2021 · 2021
Later among the works it cites.
Debiasing isn’t enough! – on the effectiveness of debiasing MLMs and their social biases in downstream tasks
Masahiro Kaneko, Danushka Bollegala, and Naoaki Okazaki. 2022 · 2022
Later among the works it cites.
Gradient-based constrained sampling from language models
Sachin Kumar, Biswajit Paria, and Yulia Tsvetkov. 2022 · 2022
Later among the works it cites.
An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models
Nicholas Meade, Elinor Poole-Dayan, and Siva Reddy. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020 · 2020
Cited alongside, same era.
Towards debiasing sentence representations
Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020 · 2020
Cited alongside, same era.
Debiasing pre-trained contextualised embeddings
Masahiro Kaneko and Danushka Bollegala. 2021 · 2021
Cited alongside, same era.
Learning interpretable word embeddings via bidirectional alignment of dimensions with semantic concepts
Lütfi Kerem Şenel, Furkan Şahinuç, Veysel Yücesoy, Hinrich Schütze, Tolga Çukur, and Aykut Koç. 2022 · 2022
Later among the works it cites.
Extracting latent steering vectors from pretrained language models
Nishant Subramani, Nivedita Suresh, and Matthew E Peters. 2022 · 2022
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023 · 2023
Closest in time.
NORMSAGE: Multi-lingual multi-cultural norm discovery from conversations on-the-fly
Yi Fung, Tuhin Chakrabarty, Hao Guo, Owen Rambow, Smaranda Muresan, and Heng Ji. 2023 · 2023
Closest in time.
Backpack language models
John Hewitt, John Thickstun, Christopher Manning, and Percy Liang. 2023 · 2023
Closest in time.
Platypus: Quick, cheap, and powerful refinement of llms
Ariel N Lee, Cole J Hunter, and Nataniel Ruiz. 2023 · 2023
Closest in time.
Defining a new NLP playground
Sha Li, Chi Han, Pengfei Yu, Carl Edwards, Manling Li, Xingyao Wang, Yi Fung, Charles Yu, Joel Tetreault, Eduard Hovy, and Heng Ji. 2023a · 2023
Closest in time.
Social-group-agnostic bias mitigation via the stereotype content model
Ali Omrani, Alireza Salkhordeh Ziabari, Charles Yu, Preni Golazizian, Brendan Kennedy, Mohammad Atari, Heng Ji, and Morteza Dehghani. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Adept: A debiasing prompt framework
Ke Yang, Charles Yu, Yi Fung, Manling Li, and Heng Ji. 2023 · 2023
Closest in time.
Unlearning bias in language models by partitioning gradients
Charles Yu, Sullam Jeoung, Anish Kasi, Pengfei Yu, and Heng Ji. 2023 · 2023
Closest in time.
Controlled text generation with natural language instructions
Wangchunshu Zhou, Yuchen Eleanor Jiang, Ethan Wilcox, Ryan Cotterell, and Mrinmaya Sachan. 2023 · 2023
Closest in time.
Massively multi-cultural knowledge acquisition & lm benchmarking
Yi Fung, Ruining Zhao, Jae Doo, Chenkai Sun, and Heng Ji. 2024 · 2024
Closest in time.