Fetching the paper…
Reading the bibliography…
Large-scale language models have recently demonstrated impressive empirical performance.
Multi-task Deep Neural Networks for Natural Language Understanding
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao · 1901
Earlier work this paper cites.
ERNIE: Enhanced Representation through Knowledge Integration
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu · 1904
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Patient Knowledge Distillation for BERT Model Compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 1908
Earlier work this paper cites.
GraphMix: Improved Training of GNNs for Semi-Supervised Learning
Vikas Verma, Meng Qu, Alex Lamb, Yoshua Bengio, Juho Kannala, and Jian Tang · 1909
Earlier work this paper cites.
Attentive Student Meets Multi-task Teacher: Improved Knowledge Distillation for Pretrained Models
Linqing Liu, Huan Wang, Jimmy Lin, Richard Socher, and Caiming Xiong · 1911
Earlier work this paper cites.
Transformation Invariance in Pattern Recognition–Tangent Distance and Tangent Propagation
Patrice Y Simard, Yann A LeCun, John S Denker, and Bernard Victorri · 1998
Earlier work this paper cites.
A Model of Inductive Bias Learning
J. Baxter · 2000
Earlier work this paper cites.
Best Practices for Convolutional Neural Networks Applied to Visual Document Analysis
PY Simard, D Steinkraus, and JC Platt · 2003
Earlier work this paper cites.
The PASCAL Recognising Textual Entailment Challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2005
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases
William B Dolan and Chris Brockett · 2005
Earlier work this paper cites.
The Second PASCAL Recognising Textual Entailment Challenge
R Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor · 2006
Earlier work this paper cites.
The Third PASCAL Recognizing Textual Entailment Challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan · 2007
Earlier work this paper cites.
Visualizing Data using t-SNE
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
The Fifth PASCAL Recognizing Textual Entailment Challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo · 2009
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality over a Sentiment Treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Cited alongside, same era.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Foundations of Machine Learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Cited alongside, same era.
A Broad-coverage Challenge Corpus for Sentence Understanding through Inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2018
Cited alongside, same era.
QANet: Combining Local Convolution with Global Self-attention for Reading Comprehension
GLUE: A Multi-task Benchmark and Analysis Platform for Natural Language Understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2019
Later among the works it cites.
EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks
Jason Wei and Kai Zou · 2019
Later among the works it cites.
Conditional BERT Contextual Augmentation
Xing Wu, Shangwen Lv, Liangjun Zang, Jizhong Han, and Songlin Hu · 2019
Later among the works it cites.
Unsupervised Data Augmentation for Consistency Training
Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V Le · 2019
Later among the works it cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V Le · 2018
Cited alongside, same era.
mixup: Beyond Empirical Risk Minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2018
Cited alongside, same era.
BAM! Born-again Multi-task Networks for Natural Language Understanding
Kevin Clark, Minh-Thang Luong, Urvashi Khandelwal, Christopher D Manning, and Quoc V Le · 2019
Cited alongside, same era.
Augmenting Data with Mixup for Sentence Classification: An Empirical Study
Hongyu Guo, Yongyi Mao, and Richong Zhang · 2019
Cited alongside, same era.
TinyBERT: Distilling BERT for Natural Language Understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2019
Cited alongside, same era.
Submodular Optimization-based Diverse Paraphrasing and Its Effectiveness in Data Augmentation
Ashutosh Kumar, Satwik Bhattamishra, Manik Bhandari, and Partha Talukdar · 2019
Cited alongside, same era.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 2019
Cited alongside, same era.
CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo · 2019
Later among the works it cites.
Extreme Language Model Compression with Optimal Subwords and Shared Projections
Sanqiang Zhao, Raghav Gupta, Yang Song, and Denny Zhou · 2019
Later among the works it cites.
UniLMv2: Pseudo-masked Language Models for Unified Language Model Pre-training
Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Songhao Piao, Jianfeng Gao, Ming Zhou, et al · 2020
Closest in time.
MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text Classification
Jiaao Chen, Zichao Yang, and Diyi Yang · 2020
Closest in time.
ELECTRA: Pre-training Text Encoders as Discriminators Rather than Generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning · 2020
Closest in time.
AugMix: A Simple Method to Improve Robustness and Uncertainty under Data Shift
Dan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan · 2020
Closest in time.
SpanBERT: Improving Pre-training by Representing and Predicting Spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy · 2020
Closest in time.
Adversarial Vertex Mixup: Toward Better Adversarially Robust Generalization
Saehyung Lee, Hyungyu Lee, and Sungroh Yoon · 2020
Closest in time.
CoDA: Contrast-enhanced and Diversity-promoting Data Augmentation for Natural Language Understanding
Yanru Qu, Dinghan Shen, Yelong Shen, Sandra Sajeev, Jiawei Han, and Weizhu Chen · 2020
Closest in time.
Dinghan Shen, Mingzhi Zheng, Yelong Shen, Yanru Qu, and Weizhu Chen · 2020
Closest in time.
MobileBERT: A Compact Task-agnostic BERT for Resource-limited Devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou · 2020
Closest in time.
Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation from a Blackbox Model
Dongdong Wang, Yandong Li, Liqiang Wang, and Boqing Gong · 2020
Closest in time.