Fetching the paper…
Reading the bibliography…
Transformer-based models have pushed state of the art in many areas of NLP, but our understanding of what is behind their success is still limited.
Assessing BERT’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Cross-Lingual Language Model Pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
Learning and Evaluating General Linguistic Intelligence
Dani Yogatama, Cyprien de Masson d’Autume, Jerome Connor, Tomas Kocisky, Mike Chrzanowski, Lingpeng Kong, Angeliki Lazaridou, Wang Ling, Lei Yu, Chris Dyer, and Phil Blunsom. 2019 · 1901
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 1902
Earlier work this paper cites.
Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 1903
Earlier work this paper cites.
ERNIE: Enhanced Representation through Knowledge Integration
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu. 2019b · 1904
Earlier work this paper cites.
Visualizing Attention in Transformer-Based Language Representation Models
Jesse Vig. 2019 · 1904
Earlier work this paper cites.
Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, and Cho-Jui Hsieh. 2019 · 1904
Earlier work this paper cites.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019b · 1905
Earlier work this paper cites.
Pre-Training with Whole Word Masking for Chinese BERT
Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Ziqing Yang, Shijin Wang, and Guoping Hu. 2019 · 1906
Earlier work this paper cites.
Inducing syntactic trees from BERT representations
Rudolf Rosa and David Mareček. 2019 · 1906
Earlier work this paper cites.
Sofia Serrano and Noah A. Smith. 2019 · 1906
Earlier work this paper cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 1906
Earlier work this paper cites.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2019 · 1907
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni. 2019 · 1907
Earlier work this paper cites.
ERNIE 2.0: A Continual Pre-Training Framework for Language Understanding
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2019c · 1907
Earlier work this paper cites.
Well-Read Students Learn Better: The Impact of Student Initialization on Knowledge Distillation
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 1908
Earlier work this paper cites.
StructBERT: Incorporating Language Structures into Pre-Training for Deep Language Understanding
Wei Wang, Bin Bi, Ming Yan, Chen Wu, Zuyi Bao, Liwei Peng, and Luo Si. 2019a · 1908
Earlier work this paper cites.
Symmetric Regularization based BERT for Pair-Wise Semantic Reasoning
Xingyi Cheng, Weidi Xu, Kunlong Chen, Wei Wang, Bin Bi, Ming Yan, Chen Wu, Luo Si, Wei Chu, and Taifeng Wang. 2019 · 1909
Earlier work this paper cites.
Reweighted Proximal Pruning for Large-Scale Language Representation
Fu-Ming Guo, Sijia Liu, Finlay S. Mungall, Xue Lin, and Yanzhi Wang. 2019 · 1909
Earlier work this paper cites.
TinyBERT: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Mixout: Effective regularization to finetune large-scale pretrained language models
Cheolhyoung Lee, Kyunghyun Cho, and Wanmo Kang. 2019 · 1909
Earlier work this paper cites.
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2019 · 1909
Earlier work this paper cites.
Small and Practical BERT Models for Sequence Labeling
Henry Tsai, Jason Riesa, Melvin Johnson, Naveen Arivazhagan, Xin Li, and Amelia Archer. 2019 · 1909
Earlier work this paper cites.
Do NLP Models Know Numbers? Probing Numeracy in Embeddings
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, and Matt Gardner. 2019b · 1909
Earlier work this paper cites.
Does BERT Make Any Sense? Interpretable Word Sense Disambiguation with Contextualized Embeddings
Gregor Wiedemann, Steffen Remus, Avi Chawla, and Chris Biemann. 2019 · 1909
Earlier work this paper cites.
Extreme Language Model Compression with Optimal Subwords and Shared Projections
Sanqiang Zhao, Raghav Gupta, Yang Song, and Denny Zhou. 2019 · 1909
Earlier work this paper cites.
FreeLB: Enhanced Adversarial Training for Language Understanding
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2019 · 1909
Earlier work this paper cites.
Knowledge Distillation from Internal Representations
Gustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Yao, Xing Fan, and Edward Guo. 2019 · 1910
Earlier work this paper cites.
Whatcha lookin’at? DeepLIFTing BERT’s Attention in Question Answering
Ekaterina Arkhangelskaia and Sourav Dutta. 2019 · 1910
Earlier work this paper cites.
exBERT: A Visual Analysis Tool to Explore Learned Representations in Transformers Models
Benjamin Hoover, Hendrik Strobelt, and Sebastian Gehrmann. 2019 · 1910
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Universal Text Representation from BERT: An Empirical Study
Xiaofei Ma, Zhiguo Wang, Patrick Ng, Ramesh Nallapati, and Bing Xiang. 2019 · 1910
Earlier work this paper cites.
Structured Pruning of a BERT-based Question Answering Model
J. S. McCarley, Rishav Chakravarti, and Avirup Sil. 2020 · 1910
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019a · 1910
Earlier work this paper cites.
What does BERT learn from multiple-choice reading comprehension datasets?
Chenglei Si, Shuohang Wang, Min-Yen Kan, and Jing Jiang. 2019a · 1910
Earlier work this paper cites.
What does BERT Learn from Multiple-Choice Reading Comprehension Datasets?
Chenglei Si, Shuohang Wang, Min-Yen Kan, and Jing Jiang. 2019b · 1910
Earlier work this paper cites.
SesameBERT: Attention for Anywhere
Ta-Chun Su and Hsiang-Chih Cheng. 2019 · 1910
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-Art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2020 · 1910
Earlier work this paper cites.
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 1910
Earlier work this paper cites.
On the Cross-lingual Transferability of Monolingual Representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2019 · 1911
Earlier work this paper cites.
Understanding Multi-Head Attention in Abstractive Summarization
Joris Baan, Maartje ter Hoeve, Marlies van der Wees, Anne Schuth, and Maarten de Rijke. 2019 · 1911
Earlier work this paper cites.
Inducing Relational Knowledge from BERT
Zied Bouraoui, Jose Camacho-Collados, and Steven Schockaert. 2019 · 1911
Earlier work this paper cites.
Unsupervised Cross-Lingual Representation Learning at Scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1911
Earlier work this paper cites.
Do attention heads in BERT track syntactic dependencies?
Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R Bowman. 2019 · 1911
Earlier work this paper cites.
Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao. 2019a · 1911
Earlier work this paper cites.
How Can We Know What Language Models Know?
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2019b · 1911
Earlier work this paper cites.
What do you mean, BERT? assessing BERT as a distributional semantics model
Timothee Mickus, Denis Paperno, Mathieu Constant, and Kees van Deemeter. 2019 · 1911
Earlier work this paper cites.
BERT is not a knowledge base (yet): Factual knowledge vs. name-based reasoning in unsupervised qa
Nina Poerner, Ulli Waltinger, and Hinrich Schütze. 2019 · 1911
Earlier work this paper cites.
Dhanasekar Sundararaman, Vivek Subramanian, Guoyin Wang, Shijing Si, Dinghan Shen, Dong Wang, and Lawrence Carin. 2019 · 1911
Earlier work this paper cites.
KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation
Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhiyuan Liu, Juanzi Li, and Jian Tang. 2020c · 1911
Earlier work this paper cites.
How Can BERT Help Lexical Semantics Tasks?
Yile Wang, Leyang Cui, and Yue Zhang. 2020d · 1911
Earlier work this paper cites.
Deepening Hidden Representations from Pre-Trained Language Models for Natural Language Understanding
Junjie Yang and Hai Zhao. 2019 · 1911
Earlier work this paper cites.
Improving BERT Fine-tuning with Embedding Normalization
Wenxuan Zhou, Junyi Du, and Xiang Ren. 2019 · 1911
Earlier work this paper cites.
What Does My QA Model Know? Devising Controlled Probes using Expert Knowledge
Kyle Richardson and Ashish Sabharwal. 2019 · 1912
Earlier work this paper cites.
oLMpics – On what Language Model Pre-Training Captures
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2019 · 1912
Earlier work this paper cites.
WaLDORf: Wasteless Language-model Distillation On Reading-comprehension
James Yi Tian, Alexander P Kreuzer, Pai-Hung Chen, and Hans-Martin Will. 2019 · 1912
Earlier work this paper cites.
Cross-Lingual Ability of Multilingual BERT: An Empirical Study
Zihan Wang, Stephen Mayhew, Dan Roth, et al. 2019b · 1912
Cited alongside, same era.
On the comparability of Pre-Trained Language Models
Matthias Aßenmacher and Christian Heumann. 2020 · 2001
Cited alongside, same era.
Power-bert: Accelerating BERT inference for classification tasks
Saurabh Goyal, Anamitra Roy Choudhary, Venkatesan Chakaravarthy, Saurabh ManishRaje, Yogish Sabharwal, and Ashish Verma. 2020 · 2001
Cited alongside, same era.
Wei-Tsung Kao, Tsung-Han Wu, Po-Han Chi, Chun-Cheng Hsieh, and Hung-Yi Lee. 2020 · 2001
Cited alongside, same era.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, Djamé Seddah, Samuel Unicomb, Gerardo Iñiguez, Márton Karsai, Yannick Léo, Márton Karsai, Carlos Sarraute, Éric Fleury, et al. 2019 · 2019
Later among the works it cites.
75 Languages, 1 Model: Parsing Universal Dependencies Universally
Dan Kondratyuk and Milan Straka. 2019 · 2019
Later among the works it cites.
A mutual information maximization perspective of language representation learning
Lingpeng Kong, Cyprien de Masson d’Autume, Lei Yu, Wang Ling, Zihang Dai, and Dani Yogatama. 2019 · 2019
Later among the works it cites.
Revealing the Dark Secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Later among the works it cites.
Open Sesame: Getting inside BERT’s Linguistic Knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 2019
Later among the works it cites.
Linguistic Knowledge and Transferability of Contextual Representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Songhao Piao, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2020 · 2002
Cited alongside, same era.
Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2002
Cited alongside, same era.
Compressing large-scale transformer-based models: A case study on BERT
Prakhar Ganesh, Yao Chen, Xin Lou, Mohammad Ali Khan, Yin Yang, Deming Chen, Marianne Winslett, Hassan Sajjad, and Preslav Nakov. 2020 · 2002
Cited alongside, same era.
Compressing BERT: Studying the effects of weight pruning on transfer learning
Mitchell A Gordon, Kevin Duh, and Nicholas Andrews. 2020 · 2002
Cited alongside, same era.
REALM: Retrieval-Augmented Language Model Pre-Training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2002
Cited alongside, same era.
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E Gonzalez. 2020 · 2002
Cited alongside, same era.
Fixed Encoder Self-Attention Patterns in Transformer-Based Machine Translation
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann. 2020 · 2002
Cited alongside, same era.
How Much Knowledge Can You Pack Into the Parameters of a Language Model?
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2002
Cited alongside, same era.
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019a · 2019
Later among the works it cites.
On Measuring Social Biases in Sentence Encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019 · 2019
Later among the works it cites.
Are Sixteen Heads Really Better than One?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
Probing Neural Network Comprehension of Natural Language Arguments
Timothy Niven and Hung-Yu Kao. 2019 · 2019
Later among the works it cites.
Knowledge Enhanced Contextual Word Representations
Matthew E. Peters, Mark Neumann, Robert Logan, Roy Schwartz, Vidur Joshi, Sameer Singh, and Noah A. Smith. 2019a · 2019
Later among the works it cites.
To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks
Matthew E. Peters, Sebastian Ruder, and Noah A. Smith. 2019b · 2019
Later among the works it cites.
Language Models as Knowledge Bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Later among the works it cites.
Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-Data Tasks
Jason Phang, Thibault Févry, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019b · 2019
Later among the works it cites.
BERT and PALs: Projected Attention Layers for Efficient Adaptation in Multi-Task Learning
Asa Cooper Stickland and Iain Murray. 2019 · 2019
Later among the works it cites.
Energy and Policy Considerations for Deep Learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 2019
Later among the works it cites.
Patient Knowledge Distillation for BERT Model Compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019a · 2019
Later among the works it cites.
Quantity doesn’t buy quality syntax with neural language models
Marten van Schijndel, Aaron Mueller, and Tal Linzen. 2019 · 2019
Later among the works it cites.
Analyzing the Structure of Attention in a Transformer Language Model
Jesse Vig and Yonatan Belinkov. 2019 · 2019
Later among the works it cites.
Elena Voita, Rico Sennrich, and Ivan Titov. 2019a · 2019
Later among the works it cites.
Universal Adversarial Triggers for Attacking and Analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019a · 2019
Later among the works it cites.
Investigating BERT’s Knowledge of Language: Five Analysis Methods with NPIs
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, et al. 2019 · 2019
Later among the works it cites.
Attention is not not Explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
Conditional BERT Contextual Augmentation
Xing Wu, Shangwen Lv, Liangjun Zang, Jizhong Han, and Songlin Hu. 2019b · 2019
Later among the works it cites.
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
ERNIE: Enhanced Language Representation with Informative Entities
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019 · 2019
Later among the works it cites.
What’s in a Name? Are BERT Named Entity Representations just as Good for any other Name?
Sriram Balasubramanian, Naman Jain, Gaurav Jindal, Abhijeet Awasthi, and Sunita Sarawagi. 2020 · 2020
Closest in time.
Interpreting Pretrained Contextualized Representations via Reductions to Static Embeddings
Rishi Bommasani, Kelly Davis, and Claire Cardie. 2020 · 2020
Closest in time.
On Identifiability in Transformers
Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020 · 2020
Closest in time.
ELECTRA: Pre-Training Text Encoders as Discriminators Rather Than Generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Closest in time.
TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence Selection
Siddhant Garg, Thuy Vu, and Alessandro Moschitti. 2020 · 2020
Closest in time.
Span Selection Pre-training for Question Answering
Michael Glass, Alfio Gliozzo, Rishav Chakravarti, Anthony Ferritto, Lin Pan, G P Shrivatsa Bhargav, Dinesh Garg, and Avi Sil. 2020 · 2020
Closest in time.
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020 · 2020
Closest in time.
SpanBERT: Improving Pre-Training by Representing and Predicting Spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Closest in time.
Are pre-trained language models aware of phrases? simple but strong baselines for grammar induction
Taeuk Kim, Jihun Choi, Daniel Edmiston, and Sang-goo Lee. 2020 · 2020
Closest in time.
Thieves on Sesame Street! Model Extraction of BERT-Based APIs
Kalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot, and Mohit Iyyer. 2020 · 2020
Closest in time.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020b · 2020
Closest in time.
Transformers to Learn Hierarchical Contexts in Multiparty Dialogue for Span-based Question Answering
Changmao Li and Jinho D. Choi. 2020 · 2020
Closest in time.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Christopher D. Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy. 2020 · 2020
Closest in time.
Contextual and Non-Contextual Word Embeddings: An in-depth Linguistic Investigation
Alessio Miaschi and Felice Dell’Orletta. 2020 · 2020
Closest in time.
Turing-NLG: A 17-billion-parameter language model by microsoft
Microsoft. 2020 · 2020
Closest in time.
When BERT Plays the Lottery, All Tickets Are Winning
Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020 · 2020
Closest in time.
Improving Transformer Models by Reordering their Sublayers
Ofir Press, Noah A. Smith, and Omer Levy. 2020 · 2020
Closest in time.
Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?
Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, and Samuel R. Bowman. 2020 · 2020
Closest in time.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Closest in time.
Probing Natural Language Inference Models through Semantic Fragments
Kyle Richardson, Hai Hu, Lawrence S. Moss, and Ashish Sabharwal. 2020 · 2020
Closest in time.
Getting Closer to AI Complete Question Answering: A Set of Prerequisite Real Tasks
Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky. 2020 · 2020
Closest in time.
BERTRAM: Improved Word Embeddings Have Big Impact on Contextualized Model Performance
Timo Schick and Hinrich Schütze. 2020 · 2020
Closest in time.
Assessing the Benchmarking Capacity of Machine Reading Comprehension Datasets
Saku Sugawara, Pontus Stenetorp, Kentaro Inui, and Akiko Aizawa. 2020 · 2020
Closest in time.
MobileBERT: Task-Agnostic Compression of BERT for Resource Limited Devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Closest in time.
Document Classification by Word Embeddings of BERT
Hirotaka Tanaka, Hiroyuki Shinnou, Rui Cao, Jing Bai, and Wen Ma. 2020 · 2020
Closest in time.
A Cross-Task Analysis of Text Span Representations
Shubham Toshniwal, Haoyue Shi, Bowen Shi, Lingyu Gao, Karen Livescu, and Kevin Gimpel. 2020 · 2020
Closest in time.
David Vilares, Michalina Strzyz, Anders Søgaard, and Carlos Gómez-Rodríguez. 2020 · 2020
Closest in time.
Can neural networks acquire a structural bias from raw linguistic data?
Alex Warstadt and Samuel R. Bowman. 2020 · 2020
Closest in time.
Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERT
Zhiyong Wu, Yun Chen, Ben Kao, and Qun Liu. 2020 · 2020
Closest in time.
TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020 · 2020
Closest in time.
Semantics-aware BERT for Language Understanding
Zhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li, Shuailiang Zhang, Xi Zhou, and Xiang Zhou. 2020 · 2020
Closest in time.
How does BERT’s attention change when you fine-tune? An analysis methodology and a case study in negation scope
Yiyun Zhao and Steven Bethard. 2020 · 2020
Closest in time.
Evaluating Commonsense in Pre-Trained Language Models
Xuhui Zhou, Yue Zhang, Leyang Cui, and Dandan Huang. 2020 · 2020
Closest in time.