Fetching the paper…
Reading the bibliography…
Beyond the success story of pre-trained language models (PrLMs) in recent natural language processing, they are susceptible to over-fitting due to unusual large model size.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 1993
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder · 2003
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew D. Zeiler, Sixin Zhang, Yann LeCun, and Rob Fergus · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
A* sampling
Chris J. Maddison, Daniel Tarlow, and Tom Minka · 2014
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Avrim Blum, Nika Haghtalab, and Ariel D. Procaccia · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V. Le · 2017
Earlier work this paper cites.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Variational dropout sparsifies deep neural networks
Dmitry Molchanov, Arsenii Ashukha, and Dmitry P. Vetrov · 2017
Earlier work this paper cites.
Concrete dropout
Yarin Gal, Jiri Hron, and Alex Kendall · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Adversarial dropout for supervised and semi-supervised learning
Sungrae Park, Jun-Keon Park, Su-Jin Shin, and Il-Chul Moon · 2018
Cited alongside, same era.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
DARTS: differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2019
Cited alongside, same era.
Code summarization with structure-induced transformer
Hongqiu Wu, Hai Zhao, and Min Zhang · 2020
Later among the works it cites.
Edropout: Energy-based dropout and pruning of deep neural networks
Hojjat Salehinejad and Shahrokh Valaee · 2020
Later among the works it cites.
Learnable bernoulli dropout for bayesian deep learning
Shahin Boluki, Randy Ardywibowo, Siamak Zamani Dadaneh, Mingyuan Zhou, and Xiaoning Qian · 2020
Later among the works it cites.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin · 2020
Later among the works it cites.
Scheduled drophead: A regularization method for transformer models
Wangchunshu Zhou, Tao Ge, Furu Wei, Ming Zhou, and Ke Xu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ARM: augment-reinforce-merge gradient for stochastic binary networks
Mingzhang Yin and Mingyuan Zhou · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Cited alongside, same era.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Cited alongside, same era.
ELECTRA: pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning · 2020
Cited alongside, same era.
Hard-coded gaussian attention for neural machine translation
Weiqiu You, Simeng Sun, and Mohit Iyyer · 2020
Cited alongside, same era.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Cited alongside, same era.
Zewei Sun, Shujian Huang, Xinyu Dai, and Jiajun Chen · 2020
Later among the works it cites.
How does selective mechanism improve self-attention networks?
Xinwei Geng, Longyue Wang, Xing Wang, Bing Qin, Ting Liu, and Zhaopeng Tu · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Later among the works it cites.
{DEBERTA}: {DECODING}-{enhanced} {bert} {with} {disentangled} {attention}
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Closest in time.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas · 2021
Closest in time.
Autodropout: Learning dropout patterns to regularize deep networks
Hieu Pham and Quoc V. Le · 2021
Closest in time.
Contextual dropout: An efficient sample-dependent dropout module
Xinjie Fan, Shujian Zhang, Korawat Tanwisuth, Xiaoning Qian, and Mingyuan Zhou · 2021
Closest in time.
Unidrop: A simple yet effective technique to improve transformer without extra cost
Zhen Wu, Lijun Wu, Meng Qi, Yingce Xia, Shufang Xie, Tao Qin, Xinyu Dai, and Tie-Yan Liu · 2021
Closest in time.