Language models are unsupervised multitask learners
A. Radford, Jeffrey Wu, R. Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Understanding learning dynamics of language models with SVCCA
Naomi Saphra and Adam Lopez. 2019 · 2019
Later among the works it cites.
Wikimatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2019 · 2019
Later among the works it cites.
Evaluating gender bias in machine translation
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. 2019 · 2019
Later among the works it cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Quantity doesn’t buy quality syntax with neural language models
Marten van Schijndel, Aaron Mueller, and Tal Linzen. 2019 · 2019
Later among the works it cites.
Blimp: The benchmark of linguistic minimal pairs for english
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Quantifying attention flow in transformers
Samira Abnar and Willem Zuidema. 2020 · 2020
Later among the works it cites.
Classifying syntactic errors in learner language
Leshem Choshen, Dmitry Nikolaev, Yevgeni Berzak, and Omri Abend. 2020 · 2020
Later among the works it cites.
Active Learning for BERT: An Empirical Study
Liat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch, Lena Dankin, Leshem Choshen, Marina Danilevsky, Ranit Aharonov, Yoav Katz, and Noam Slonim. 2020 · 2020
Later among the works it cites.
Let’s agree to agree: Neural networks share classification order on real datasets
Guy Hacohen, Leshem Choshen, and D. Weinshall. 2020 · 2020
Later among the works it cites.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020 · 2020
Later among the works it cites.
What do position embeddings learn? an empirical study of pre-trained language model positional encoding
Yu-An Wang and Yun-Nung Chen. 2020 · 2020
Later among the works it cites.
The gem benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin P. Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh D. Dhole, Wanyu Du, Esin Durmus, Ondrej Dusek, Chris C. Emezue, Varun Gangal, Cristina Garbacea, Tatsunori B. Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Rubungo Andre Niyongabo, Salomey Osei, Ankur P. Parikh, Laura Perez-Beltrachini, Niranjan Rao, Vikas Raunak, Juan Diego Rodríguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Closest in time.
Principal components bias in deep neural networks
Original
Guy Hacohen and Daphna Weinshall. 2021 · 2021
Closest in time.
Probing across time: What does roberta know and when?
L. Liu, Yizhong Wang, Jungo Kasai, Hannaneh Hajishirzi, and Noah A. Smith. 2021 · 2021
Closest in time.
Making transformers solve compositional tasks
Original
Santiago Ontan’on, Joshua Ainslie, Vaclav Cvicek, and Zachary Kenneth Fisher. 2021 · 2021
Closest in time.
When deep classifiers agree: Analyzing correlations between learning order and image statistics
Original
Iuliia Pliushch, Martin Mundt, Nicolas Lupp, and Visvanathan Ramesh. 2021 · 2021
Closest in time.
Mediators in determining what processing BERT performs first
Aviv Slobodkin, Leshem Choshen, and Omri Abend. 2021 · 2021
Closest in time.
Active learning on a budget: Opposite strategies suit high and low budgets
Original
Guy Hacohen, Avihu Dekel, and Daphna Weinshall. 2022 · 2022
Closest in time.
Scaling laws under the microscope: Predicting transformer performance from small scale experiments
Original
Maor Ivgi, Yair Carmon, and Jonathan Berant. 2022 · 2022
Closest in time.