BERT has a mouth, and it must speak: BERT as a Markov random field language model
Alex Wang and Kyunghyun Cho · 2019
Later among the works it cites.
Linguistic Analysis of Pretrained Sentence Encoders with Acceptability Judgments
Original
Alex Warstadt and Samuel R. Bowman · 2019
Later among the works it cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
What Don’t RNN Language Models Learn About Filler-Gap Dependencies?
Rui P. Chaves · 2020
Later among the works it cites.
UNITER: UNiversal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Later among the works it cites.
SyntaxGym: An Online Platform for Targeted Evaluation of Language Models
Jon Gauthier, Jennifer Hu, Ethan Wilcox, Peng Qian, and Roger Levy · 2020
Later among the works it cites.
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2020
Later among the works it cites.
A Systematic Assessment of Syntactic Generalization in Neural Language Models
Jennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox, and Roger Levy · 2020
Later among the works it cites.
COGS: A Compositional Generalization Challenge Based on Semantic Interpretation
Najoung Kim and Tal Linzen · 2020
Later among the works it cites.
Multi-agent Communication meets Natural Language: Synergies between Functional and Structural Language Learning
Angeliki Lazaridou, Anna Potapenko, and Olivier Tieleman · 2020
Later among the works it cites.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Christopher D. Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy · 2020
Later among the works it cites.
Does Syntax Need to Grow on Trees? Sources of Hierarchical Inductive Bias in Sequence-to-Sequence Networks
R. Thomas McCoy, Robert Frank, and Tal Linzen · 2020
Later among the works it cites.
Recurrent babbling: evaluating the acquisition of grammar from limited input data
Ludovica Pannitto and Aurélie Herbelot · 2020
Later among the works it cites.
Information-Theoretic Probing for Linguistic Structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell · 2020
Later among the works it cites.
Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Later among the works it cites.
Capacity, bandwidth, and compositionality in emergent language learning
Cinjon Resnick, Abhinav Gupta, Jakob Foerster, Andrew M Dai, and Kyunghyun Cho · 2020
Later among the works it cites.
“LazImpa”: Lazy and Impatient neural agents learn to communicate efficiently
Mathieu Rita, Rahma Chaabouni, and Emmanuel Dupoux · 2020
Later among the works it cites.
A Primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2020
Later among the works it cites.
WinoGrande: An Adversarial Winograd Schema Challenge at Scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2020
Later among the works it cites.
Masked Language Model Scoring
Julian Salazar, Davis Liang, Toan Q. Nguyen, and Katrin Kirchhoff · 2020
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2020
Later among the works it cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Later among the works it cites.
Information-Theoretic Probing with Minimum Description Length
Elena Voita and Ivan Titov · 2020
Later among the works it cites.
Can neural networks acquire a structural bias from raw linguistic data?
Alex Warstadt and Samuel R Bowman · 2020
Later among the works it cites.
Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually)
Alex Warstadt, Yian Zhang, Haau-Sing Li, Haokun Liu, and Samuel R Bowman · 2020
Later among the works it cites.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew G Wilson and Pavel Izmailov · 2020
Later among the works it cites.
On the proper role of linguistically-oriented deep net analysis in linguistic theorizing
Original
Marco Baroni · 2021
Later among the works it cites.
The grammar-learning trajectories of neural language models
Original
Leshem Choshen, Guy Hacohen, Daphna Weinshall, and Omri Abend · 2021
Later among the works it cites.
(What) Can Deep Learning Contribute to Theoretical Linguistics?
Gabe Dupre · 2021
Later among the works it cites.
Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov · 2021
Later among the works it cites.
The Natural Stories corpus: a reading-time corpus of English texts containing rare syntactic constructions
Richard Futrell, Edward Gibson, Harry J. Tily, Idan Blank, Anastasia Vishnevetsky, Steven T. Piantadosi, and Evelina Fedorenko · 2021
Later among the works it cites.
Using lexical context to discover the noun category: Younger children have it easier
Philip A. Huebner and Jon A. Willits · 2021
Later among the works it cites.
Effect of Visual Extensions on Natural Language Understanding in Vision-and-Language Models
Taichi Iki and Akiko Aizawa · 2021
Later among the works it cites.
MDETR-modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion · 2021
Later among the works it cites.
Generative Spoken Language Modeling from Raw Audio
Original
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, and Emmanuel Dupoux · 2021
Later among the works it cites.
Syntactic Structure from Deep Learning
Tal Linzen and Marco Baroni · 2021
Later among the works it cites.
Predicting Inductive Biases of Fine-tuned Models
Charles Lovering, Rohan Jha, Tal Linzen, and Ellie Pavlick · 2021
Later among the works it cites.
Deep subjecthood: Higher-order grammatical features in multilingual BERT
Isabel Papadimitriou, Ethan A. Chi, Richard Futrell, and Kyle Mahowald · 2021
Later among the works it cites.
Transformers Generalize Linearly
Original
Jackson Petty and Robert Frank · 2021
Later among the works it cites.
How much pretraining data do language models need to learn syntax?
Original
Laura Pérez-Mayos, Miguel Ballesteros, and Leo Wanner · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Original
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, and others · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Original
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, and others · 2021
Later among the works it cites.
The nature of the semantic stimulus: the acquisition of every as a case study
Ezer Rasin and Athulya Aravind · 2021
Later among the works it cites.
Flava: A foundational language and vision alignment model
Original
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2021
Later among the works it cites.
SAYCam: A Large, Longitudinal Audiovisual Dataset Recorded From the Infant’s Perspective
Jessica Sullivan, Michelle Mei, Andrew Perfors, Erica Wojcik, and Michael C. Frank · 2021
Later among the works it cites.
A Targeted Assessment of Incremental Processing in Neural Language Models and Humans
Ethan Wilcox, Pranali Vani, and Roger Levy · 2021
Later among the works it cites.
Does Vision-and-Language Pretraining Improve Lexical Grounding?
Tian Yun, Chen Sun, and Ellie Pavlick · 2021
Later among the works it cites.
When Do You Need Billions of Words of Pretraining Data?
Original
Yian Zhang, Alex Warstadt, Xiaocheng Li, and Samuel R. Bowman · 2021
Later among the works it cites.
Emergent communication at scale
Rahma Chaabouni, Florian Strub, Florent Altché, Eugene Tarassov, Corentin Tallec, Elnaz Davoodi, Kory Wallace Mathewson, Olivier Tieleman, Angeliki Lazaridou, and Bilal Piot · 2022
Closest in time.
Word Acquisition in Neural Language Models
Tyler A. Chang and Benjamin K. Bergen · 2022
Closest in time.
PaLM: Scaling language modeling with pathways
Original
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, and others · 2022
Closest in time.
Early phonetic learning from ecological audio: domain-general versus domain-specific mechanisms
Marvin Lavechin, Maureen de Seyssel, Marianne Métais, Florian Metze, Abdelrahman Mohamed, Hervé BREDIN, Emmanuel Dupoux, and Alejandrina Cristia · 2022
Closest in time.
One model for the learning of language
Yuan Yang and Steven T. Piantadosi · 2022
Closest in time.