Fetching the paper…
Reading the bibliography…
Large Transformer-based models were shown to be reducible to a smaller number of self-attention heads and layers.
Sofia Serrano and Noah A. Smith. 2019 · 1906
Earlier work this paper cites.
TinyBERT: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Pruning a BERT-based Question Answering Model
JS McCarley. 2019 · 1910
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2020 · 1910
Earlier work this paper cites.
Do attention heads in BERT track syntactic dependencies?
Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R Bowman. 2019 · 1911
Earlier work this paper cites.
R. Thomas McCoy, Junghyun Min, and Tal Linzen. 2019a · 1911
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2002
Earlier work this paper cites.
Compressing BERT: Studying the effects of weight pruning on transfer learning
Mitchell A. Gordon, Kevin Duh, and Nicholas Andrews. 2020 · 2002
Earlier work this paper cites.
Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning
Peter Henderson, Jieru Hu, Joshua Romoff, Emma Brunskill, Dan Jurafsky, and Joelle Pineau. 2020 · 2002
Earlier work this paper cites.
A Primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020b · 2002
Earlier work this paper cites.
Information-Theoretic Probing with Minimum Description Length
Elena Voita and Ivan Titov. 2020 · 2003
Earlier work this paper cites.
Attention Module is Not Only a Weight: Analyzing Transformers with Vector Norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020 · 2004
Earlier work this paper cites.
Learning Music Helps You Read: Using Transfer to Study Linguistic Structure in Language Models
Isabel Papadimitriou and Dan Jurafsky. 2020 · 2004
Earlier work this paper cites.
The PASCAL Recognising Textual Entailment Challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
W.B. Dolan and C. Brockett. 2005 · 2005
Earlier work this paper cites.
The Second Pascal Recognising Textual Entailment Challenge
R. Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. 2006 · 2006
Earlier work this paper cites.
The Lottery Ticket Hypothesis for Pre-trained BERT Networks
Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin. 2020 · 2007
Earlier work this paper cites.
The Third PASCAL Recognizing Textual Entailment Challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007 · 2007
Earlier work this paper cites.
The Fifth PASCAL Recognizing Textual Entailment Challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. 2009 · 2009
Cited alongside, same era.
The Winograd Schema Challenge
Hector J Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
Pruning convolutional neural networks for resource efficient inference
Open Sesame: Getting inside BERT’s Linguistic Knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 2019
Later among the works it cites.
Linguistic Knowledge and Transferability of Contextual Representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Are Sixteen Heads Really Better than One?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
One ticket to win them all: Generalizing lottery ticket initializations across datasets and optimizers
Ari Morcos, Haonan Yu, Michela Paganini, and Yuandong Tian. 2019 · 2019
Later among the works it cites.
Probing Neural Network Comprehension of Natural Language Arguments
Timothy Niven and Hung-Yu Kao. 2019 · 2019
Later among the works it cites.
DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
SNIP: Single-shot network pruning based on connection sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip Torr. 2018 · 2018
Cited alongside, same era.
Country-level social cost of carbon
Katharine Ricke, Laurent Drouet, Ken Caldeira, and Massimo Tavoni. 2018 · 2018
Cited alongside, same era.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018 · 2018
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2017 · 2018
Cited alongside, same era.
What Does BERT Look at? An Analysis of BERT’s Attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
BERT Rediscovers the Classical NLP Pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
Analyzing the Structure of Attention in a Transformer Language Model
Jesse Vig and Yonatan Belinkov. 2019 · 2019
Later among the works it cites.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
Neural Network Acceptability Judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Attention is not not Explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
Are All Layers Created Equal?
Chiyuan Zhang, Samy Bengio, and Yoram Singer. 2019 · 2019
Later among the works it cites.
Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask
Hattie Zhou, Janice Lan, Rosanne Liu, and Jason Yosinski. 2019 · 2019
Later among the works it cites.
On Identifiability in Transformers
Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020 · 2020
Closest in time.
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020 · 2020
Closest in time.
ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Closest in time.
Playing the lottery with rewards and multiple languages: Lottery tickets in RL and NLP
Haonan Yu, Sergey Edunov, Yuandong Tian, and Ari S. Morcos. 2020 · 2020
Closest in time.