Fetching the paper…
Reading the bibliography…
Transformer language models have received widespread public attention, yet their generated text is often surprising even to NLP researchers.
Does BERT agree? evaluating knowledge of structure dependence through agreement relations
Bacon, Geoff and Terry Regier. 2019 · 1908
Earlier work this paper cites.
Syntax: A Generative Introduction
Carnie, Andrew. 2002 · 2002
Earlier work this paper cites.
Poverty of the stimulus revisited
Berwick, Robert, Paul Pietroski, Beracah Yankama, and Noam Chomsky. 2011 · 2011
Earlier work this paper cites.
The non-redundant contributions of Marr’s three levels of analysis for explaining information-processing mechanisms
Bechtel, William and Oron Shagrir. 2015 · 2015
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Clark, Kevin, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Earlier work this paper cites.
Always keep your target in mind: Studying semantics and improving performance of neural lexical substitution
Arefyev, Nikolay, Boris Sheludko, Alexander Podolskiy, and Alexander Panchenko. 2020 · 2020
Earlier work this paper cites.
Unmasking contextual stereotypes: Measuring and mitigating BERT’s gender bias
Bartl, Marion, Malvina Nissim, and Albert Gatt. 2020 · 2020
Earlier work this paper cites.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Bender, Emily M. and Alexander Koller. 2020 · 2020
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Blodgett, Su Lin, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, Tom, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
Pretrained language model embryology: The birth of ALBERT
Chiang, Cheng-Han, Sung-Feng Huang, and Hung-yi Lee. 2020 · 2020
Earlier work this paper cites.
Persistent anti-muslim bias in large language models
Abid, Abubakar, Maheen Farooqi, and James Zou. 2021 · 2021
Earlier work this paper cites.
Adolphs, Leonard, Shehzaad Dhuliawala, and Thomas Hofmann. 2021 · 2021
Earlier work this paper cites.
The language model understood the prompt was ambiguous: Probing syntactic uncertainty through generation
Aina, Laura and Tal Linzen. 2021 · 2021
Earlier work this paper cites.
ALL dolphins are intelligent and SOME are friendly: Probing BERT for nouns’ semantic properties and their prototypicality
Apidianaki, Marianna and Aina Garí Soler. 2021 · 2021
Earlier work this paper cites.
How reliable are model diagnostics?
Aribandi, Vamsi, Yi Tay, and Donald Metzler. 2021 · 2021
Earlier work this paper cites.
PROST: Physical reasoning about objects through space and time
Aroca-Ouellette, St’ephane, Cory Paik, Alessandro Roncone, and Katharina Kann. 2021 · 2021
Earlier work this paper cites.
Assessing political prudence of open-domain chatbots
Bang, Yejin, Nayeon Lee, Etsuko Ishii, Andrea Madotto, and Pascale Fung. 2021 · 2021
Earlier work this paper cites.
Probing pre-trained language models for semantic attributes and their values
Beloucif, Meriem and Chris Biemann. 2021 · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, Emily M., Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Earlier work this paper cites.
Thinking aloud: Dynamic context generation improves zero-shot reasoning performance of GPT-2
Betz, Gregor, Kyle Richardson, and C. Voigt. 2021 · 2021
Earlier work this paper cites.
Is incoherence surprising? targeted evaluation of coherence prediction from language models
Beyer, Anne, Sharid Loáiciga, and David Schlangen. 2021 · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, Rishi, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, S. Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen A. Creel, Jared Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren E. Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas F. Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, O. Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, Ananya Kumar, Faisal Ladhak, Mina Lee, Tony Lee, Jure Leskovec, Isabelle Levent, Xiang Lisa Li, Xuechen Li, Tengyu Ma, Ali Malik, Christopher D. Manning, Suvir P. Mirchandani, Eric Mitchell, Zanele Munyikwa, Suraj Nair, Avanika Narayan, Deepak Narayanan, Benjamin Newman, Allen Nie, Juan Carlos Niebles, Hamed Nilforoshan, J. F. Nyarko, Giray Ogut, Laurel Orr, Isabel Papadimitriou, Joon Sung Park, Chris Piech, Eva Portelance, Christopher Potts, Aditi Raghunathan, Robert Reich, Hongyu Ren, Frieda Rong, Yusuf H. Roohani, Camilo Ruiz, Jack Ryan, Christopher R’e, Dorsa Sadigh, Shiori Sagawa, Keshav Santhanam, Andy Shih, Krishna Parasuram Srinivasan, Alex Tamkin, Rohan Taori, Armin W. Thomas, Florian Tramèr, Rose E. Wang, William Wang, Bohan Wu, Jiajun Wu, Yuhuai Wu, Sang Michael Xie, Michihiro Yasunaga, Jiaxuan You, Matei A. Zaharia, Michael Zhang, Tianyi Zhang, Xikun Zhang, Yuhui Zhang, Lucia Zheng, Kaitlyn Zhou, and Percy Liang. 2021 · 2021
Cited alongside, same era.
Build a medical sentence matching application using BERT and Amazon SageMaker
On the role of bidirectionality in language model pre-training
Artetxe, Mikel, Jingfei Du, Naman Goyal, Luke Zettlemoyer, and Veselin Stoyanov. 2022 · 2022
Later among the works it cites.
Probing classifiers: Promises, shortcomings, and advances
Belinkov, Yonatan. 2022 · 2022
Later among the works it cites.
Analogy generation by prompting large language models: A case study of instructgpt
Bhavya, Bhavya, Jinjun Xiong, and ChengXiang Zhai. 2022 · 2022
Later among the works it cites.
Improving language models by retrieving from trillions of tokens
Borgeaud, Sebastian, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego De Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack Rae, Erich Elsen, and Laurent Sifre. 2022 · 2022
Later among the works it cites.
The dangers of underclaiming: Reasons for caution when reporting how NLP systems fail
Bowman, Samuel. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Broyde, Joshua and Claire Palmer. 2021 · 2021
Cited alongside, same era.
Isotropy in the contextual embedding space: Clusters and manifolds
Cai, Xingyu, Jiaji Huang, Yu-Lan Bian, and Kenneth Ward Church. 2021 · 2021
Cited alongside, same era.
Knowledgeable or educated guess? revisiting language models as knowledge bases
Cao, Boxi, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu. 2021 · 2021
Cited alongside, same era.
Extracting training data from large language models
Carlini, Nicholas, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, Alina Oprea, and Colin Raffel. 2021 · 2021
Cited alongside, same era.
Convolutions and self-attention: Re-interpreting relative positions in pre-trained language models
Chang, Tyler, Yifan Xu, Weijian Xu, and Zhuowen Tu. 2021 · 2021
Cited alongside, same era.
Look at that! BERT can be easily distracted from paying attention to morphosyntax
Chaves, Rui P. and Stephanie N. Richter. 2021 · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Chen, Mark, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde, Jared Kaplan, Harrison Edwards, Yura Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, F. Such, D. Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William H. Guss, Alex Nichol, I. Babuschkin, S. Balaji, Shantanu Jain, A. Carr, J. Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, M. Knight, Miles Brundage, Mira Murati, Katie Mayer, P. Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021 · 2021
Cited alongside, same era.
Relating neural text degeneration to exposure bias
Chiang, Ting-Rui and Yun-Nung Chen. 2021 · 2021
Cited alongside, same era.
Modeling the influence of verb aspect on the activation of typical event locations with BERT
Cho, Won Ik, Emmanuele Chersoni, Yu-Yin Hsu, and Chu-Ren Huang. 2021 · 2021
Cited alongside, same era.
Stepmothers are mean and academics are pretentious: What do pretrained language models learn about you?
Choenni, Rochelle, Ekaterina Shutova, and Robert van Rooij. 2021 · 2021
Cited alongside, same era.
How conservative are language models? adapting to the introduction of gender-neutral pronouns
Brandl, Stephanie, Ruixiang Cui, and Anders Søgaard. 2022 · 2022
Later among the works it cites.
What does it mean for a language model to preserve privacy?
Brown, Hannah, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr. 2022 · 2022
Later among the works it cites.
Can prompt probe pretrained language models? understanding the invisible risks from a causal view
Cao, Boxi, Hongyu Lin, Xianpei Han, Fangchao Liu, and Le Sun. 2022 · 2022
Later among the works it cites.
Identifying and manipulating the personality traits of language models
Caron, Graham and Shashank Srivastava. 2022 · 2022
Later among the works it cites.
The geometry of multilingual language model representations
Chang, Tyler, Zhuowen Tu, and Benjamin Bergen. 2022 · 2022
Later among the works it cites.
Word acquisition in neural language models
Chang, Tyler A. and Benjamin K. Bergen. 2022 · 2022
Later among the works it cites.
Chen, Kaiping, Anqi Shao, Jirayu Burapacheep, and Yixuan Li. 2022 · 2022
Later among the works it cites.
The grammar-learning trajectories of neural language models
Choshen, Leshem, Guy Hacohen, Daphna Weinshall, and Omri Abend. 2022 · 2022
Later among the works it cites.
PaLM: Scaling language modeling with Pathways
Chowdhery, Aakanksha, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam M. Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C. Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Firat, Michele Catasta, Jason Wei, Kathleen S. Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022 · 2022
Later among the works it cites.
Buy tesla, sell ford: Assessing implicit stock market preference in pre-trained language models
Chuang, Chengyu and Yi Yang. 2022 · 2022
Later among the works it cites.
Black-box language model explanation by context length probing
Cífka, Ondřej and Antoine Liutkus. 2022 · 2022
Later among the works it cites.
Out of one, many: Using language models to simulate human samples
Argyle, Lisa P., Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023 · 2023
Closest in time.
Using cognitive psychology to understand GPT-3
Binz, Marcel and Eric Schulz. 2023 · 2023
Closest in time.
Quantifying memorization across neural language models
Carlini, Nicholas, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2023 · 2023
Closest in time.
2023
Closest in time.