Fetching the paper…
Reading the bibliography…
Dropout is a powerful and widely used technique to regularize the training of deep neural networks.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1995
Earlier work this paper cites.
Manual and automatic evaluation of summaries
Chin-Yew Lin and Eduard Hovy · 2002
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
The difficulty of training deep architectures and the effect of unsupervised pre-training
Dumitru Erhan, Pierre-Antoine Manzagol, Yoshua Bengio, Samy Bengio, and Pascal Vincent · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky et al · 2009
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Adaptive dropout for training deep neural networks
Lei Jimmy Ba and Brendan Frey · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Fast dropout training
Sida Wang and Christopher Manning · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Analyzing noise in autoencoders and deep networks
Ben Poole, Jascha Sohl-Dickstein, and Surya Ganguli · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomáš Kočiskỳ, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Towards dropout training for convolutional neural networks
Haibing Wu and Xiaodong Gu · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Shakeout: A new regularized deep neural network training scheme
Guoliang Kang, Jun Li, and Dacheng Tao · 2016
Earlier work this paper cites.
Dropout with expectation-linear regularization
Xuezhe Ma, Yingkai Gao, Zhiting Hu, Yaoliang Yu, Yuntian Deng, and Eduard Hovy · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Weight normalization: a simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P Kingma · 2016
Earlier work this paper cites.
Recurrent dropout without memory loss
Stanislau Semeniuta, Aliaksei Severyn, and Erhardt Barth · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Learning structured sparsity in deep neural networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov · 2017
Cited alongside, same era.
Structured bayesian pruning via log-normal multiplicative noise
Kirill Neklyudov, Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Later among the works it cites.
Dropattention: A regularization method for fully-connected self-attention networks
Lin Zehui, Pengfei Liu, Luyao Huang, Junkun Chen, Xipeng Qiu, and Xuanjing Huang · 2019
Later among the works it cites.
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma · 2019
Later among the works it cites.
Muse: Parallel multi-scale attention for sequence to sequence learning
Guangxiang Zhao, Xu Sun, Jingjing Xu, Zhiyuan Zhang, and Liangchen Luo · 2019
Later among the works it cites.
Incorporating bert into neural machine translation
Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tieyan Liu · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2018
Cited alongside, same era.
Born again neural networks
Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Cited alongside, same era.
Orthogonal weight normalization: Solution to optimization over multiple dependent stiefel manifolds in deep neural networks
Lei Huang, Xianglong Liu, Bo Lang, Adams Yu, Yongliang Wang, and Bo Li · 2018
Cited alongside, same era.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Cited alongside, same era.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli · 2018
Cited alongside, same era.
A call for clarity in reporting bleu scores
Matt Post · 2018
Cited alongside, same era.
Later among the works it cites.
Better fine-tuning by reducing representational collapse
Armen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal, Luke Zettlemoyer, and Sonal Gupta · 2020
Later among the works it cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Later among the works it cites.
Very deep transformers for neural machine translation
Xiaodong Liu, Kevin Duh, Liyuan Liu, and Jianfeng Gao · 2020
Later among the works it cites.
Self-distillation amplifies regularization in hilbert space
Hossein Mobahi, Mehrdad Farajtabar, and Peter L Bartlett · 2020
Later among the works it cites.
A survey of regularization strategies for deep models
Reza Moradi, Reza Berangi, and Behrouz Minaei · 2020
Later among the works it cites.
Data diversification: A simple strategy for neural machine translation
Xuan-Phi Nguyen, Shafiq Joty, Kui Wu, and Ai Ti Aw · 2020
Later among the works it cites.
Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training
Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou · 2020
Later among the works it cites.
Dinghan Shen, Mingzhi Zheng, Yelong Shen, Yanru Qu, and Weizhu Chen · 2020
Later among the works it cites.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu · 2020
Later among the works it cites.
Scheduled drophead: A regularization method for transformer models
Wangchunshu Zhou, Tao Ge, Furu Wei, Ming Zhou, and Ke Xu · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Seed: Self-supervised distillation for visual representation
Zhiyuan Fang, Jianfeng Wang, Lijuan Wang, Lei Zhang, Yezhou Yang, and Zicheng Liu · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Simcse: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen · 2021
Closest in time.
Mixkd: Towards efficient distillation of large-scale language models
Kevin J Liang, Weituo Hao, Dinghan Shen, Yufan Zhou, Weizhu Chen, Changyou Chen, and Lawrence Carin · 2021
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Closest in time.
Autodropout: Learning dropout patterns to regularize deep networks
Hieu Pham and Quoc V Le · 2021
Closest in time.
Not all attention is all you need
Hongqiu Wu, Hai Zhao, and Min Zhang · 2021
Closest in time.
Rethinking soft labels for knowledge distillation: A bias-variance tradeoff perspective
Helong Zhou, Liangchen Song, Jiajie Chen, Ye Zhou, Guoli Wang, Junsong Yuan, and Qian Zhang · 2021
Closest in time.