Fetching the paper…
Reading the bibliography…
One of the generally accepted views of modern deep learning is that increasing the number of parameters usually leads to better quality.
Neural network ensembles
Hansen, L. and Salamon, P · 1990
Earlier work this paper cites.
Neural network ensembles, cross validation and active learning
Krogh, A. and Vedelsby, J · 1994
Earlier work this paper cites.
Analysis of decision boundaries in linearly combined neural classifiers
Tumer, K. and Ghosh, J · 1996
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
The elements of statistical learning: Data mining, inference, and prediction
Hastie, T., Tibshirani, R., Friedman, J., and Franklin, J · 2004
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
Cettolo, M., Niehues, J., Stuker, S., Bentivogli, L., and Federico, M · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Why M heads are better than one: Training a diverse ensemble of deep networks
Lee, S., Purushwalkam, S., Cogswell, M., Crandall, D. J., and Batra, D · 2015
Earlier work this paper cites.
Tensorizing neural networks
Novikov, A., Podoprikhin, D., Osokin, A., and Vetrov, D · 2015
Cited alongside, same era.
Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. J · 2016
Cited alongside, same era.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Cited alongside, same era.
Inception-v4, Inception-ResNet and the impact of residual connections on learning
Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. A · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
The power of ensembles for active learning in image classification
Beluch, W. H., Genewein, T., Nürnberger, A., and Köhler, J. M · 2018
Later among the works it cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Later among the works it cites.
GPipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Chen, D., Lee, H., Ngiam, J., Le, Q. V., and Chen, Z · 2018
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., Klingner, J., Shah, A., Johnson, M., Liu, X., Łukasz Kaiser, Gouws, S., Kato, Y., Kudo, T., Kazawa, H., Stevens, K., Kurian, G., Patil, N., Wang, W., Young, C., Smith, J., Riesa, J., Rudnick, A., Vinyals, O., Corrado, G., Hughes, M., and Dean, J · 2016
Cited alongside, same era.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Cited alongside, same era.
Training deeper neural machine translation models with transparent attention
Bapna, A., Chen, M., Firat, O., Cao, Y., and Wu, Y · 2018
Cited alongside, same era.
Gao, Y., Cai, Z., Chen, Y., Chen, W., Yang, K., Sun, C., and Yao, C · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Later among the works it cites.
Very deep self-attention networks for end-to-end speech recognition
Pham, N., Nguyen, T., Niehues, J., Müller, M., and Waibel, A · 2019
Later among the works it cites.
Can you trust your model's uncertainty? evaluating predictive uncertainty under dataset shift
Snoek, J., Ovadia, Y., Fertig, E., Lakshminarayanan, B., Nowozin, S., Sculley, D., Dillon, J., Ren, J., and Nado, Z · 2019
Later among the works it cites.
EfficientNet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q · 2019
Later among the works it cites.
Learning deep transformer models for machine translation
Wang, Q., Li, B., Xiao, T., Zhu, J., Li, C., Wong, D. F., and Chao, L. S · 2019
Later among the works it cites.
Pitfalls of in-domain uncertainty estimation and ensembling in deep learning
Ashukha, A., Lyzhov, A., Molchanov, D., and Vetrov, D · 2020
Closest in time.