Fetching the paper…
Reading the bibliography…
Selecting suitable architecture parameters and training hyperparameters is essential for enhancing machine learning (ML) model performance.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 1908
Earlier work this paper cites.
Deep ensembles: A loss landscape perspective
Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan. 2019 · 1912
Earlier work this paper cites.
Orthogonal distance regression
Paul T Boggs and Janet E Rogers. 1990 · 1990
Earlier work this paper cites.
PAC-Bayesian model averaging. In Annual Conference on Computational Learning Theory . 164–170
David A McAllester. 1999 · 1999
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Power-law distributions in empirical data
Aaron Clauset, Cosma Rohilla Shalizi, and Mark EJ Newman. 2009 · 2009
Earlier work this paper cites.
Dinghan Shen, Mingzhi Zheng, Yelong Shen, Yanru Qu, and Weizhu Chen. 2020 · 2009
Earlier work this paper cites.
Powerlaw: a Python package for analysis of heavy-tailed distributions
Jeff Alstott, Ed Bullmore, and Dietmar Plenz. 2014 · 2014
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation. In Proceedings of the ninth workshop on statistical machine translation . 12–58
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, et al · 2014
Earlier work this paper cites.
Norm-based capacity control in neural networks. In Conference on Learning Theory . PMLR, 1376–1401
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro. 2015 · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
A boundary tilting persepective on the phenomenon of adversarial examples
Thomas Tanay and Lewis Griffin. 2016 · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter Bartlett, Dylan Foster, and Matus Telgarsky. 2017 · 2017
Earlier work this paper cites.
Deep learning for natural language processing: advantages and challenges
Hang Li. 2017 · 2017
Earlier work this paper cites.
Charles H Martin and Michael W Mahoney. 2017 · 2017
Earlier work this paper cites.
Exploring Generalization in Deep Learning
Behnam Neyshabur, Srinadh Bhojanapalli, David Mcallester, and Nati Srebro. 2017 · 2017
Earlier work this paper cites.
Pac-bayesian margin bounds for convolutional neural networks
Konstantinos Pitas, Mike Davies, and Pierre Vandergheynst. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Understanding Back-Translation at Scale. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing . 489–500
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018 · 2018
Cited alongside, same era.
Large Margin Deep Networks for Classification
Gamaleldin Elsayed, Dilip Krishnan, Hossein Mobahi, Kevin Regan, and Samy Bengio. 2018 · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of DNNs. In Conference on Neural Information Processing Systems . 8803–8812
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry Vetrov, and Andrew Gordon Wilson. 2018 · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization. In Conference on Uncertainty in Artificial Intelligence . 876–885
P Izmailov, AG Wilson, D Podoprikhin, D Vetrov, and T Garipov. 2018 · 2018
Cited alongside, same era.
Predicting the Generalization Gap in Deep Networks with Margin Distributions. In International Conference on Learning Representations
Neural Architecture Search on ImageNet in Four GPU Hours: A Theoretically Inspired Perspective. In International Conference on Learning Representations
Wuyang Chen, Xinyu Gong, and Zhangyang Wang. 2020 · 2020
Later among the works it cites.
Funnel-transformer: Filtering out sequential redundancy for efficient language processing
Zihang Dai, Guokun Lai, Yiming Yang, and Quoc Le. 2020 · 2020
Later among the works it cites.
In search of robust measures of generalization
Gintare Karolina Dziugaite, Alexandre Drouin, Brady Neal, Nitarshan Rajkumar, Ethan Caballero, Linbo Wang, Ioannis Mitliagkas, and Daniel M Roy. 2020 · 2020
Later among the works it cites.
Does learning require memorization? a short tale about a long tail. In Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing . 954–959
Vitaly Feldman. 2020 · 2020
Later among the works it cites.
Sharpness-aware Minimization for Efficiently Improving Generalization. In International Conference on Learning Representations
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yiding Jiang, Dilip Krishnan, Hossein Mobahi, and Samy Bengio. 2018 · 2018
Cited alongside, same era.
A PAC-Bayesian Approach to Spectrally-Normalized Margin Bounds for Neural Networks. In International Conference on Learning Representations
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro. 2018 · 2018
Cited alongside, same era.
Scaling Neural Machine Translation. In Proceedings of the Third Conference on Machine Translation: Research Papers . 1–9
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) . 1112–1122
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
HAWQ: Hessian aware quantization of neural networks with mixed-precision. In IEEE/CVF International Conference on Computer Vision . 293–302
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2019 · 2019
Cited alongside, same era.
Fantastic Generalization Measures and Where to Find Them. In International Conference on Learning Representations
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 4171–4186
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
ALBERT: A lite BERT for self-supervised learning of language representations. In International Conference on Learning Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Neurips 2020 competition: Predicting generalization in deep learning
Yiding Jiang, Pierre Foret, Scott Yak, Daniel M Roy, Hossein Mobahi, Gintare Karolina Dziugaite, Samy Bengio, Suriya Gunasekar, Isabelle Guyon, and Behnam Neyshabur. 2020 · 2020
Later among the works it cites.
FlauBERT: Unsupervised Language Model Pre-training for French. In Proceedings of the 12th Language Resources and Evaluation Conference . 2479–2490
Hang Le, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoit Crabbé, Laurent Besacier, and Didier Schwab. 2020 · 2020
Later among the works it cites.
Understanding the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han. 2020 · 2020
Later among the works it cites.
Heavy-tailed Universality predicts trends in test accuracies for very large pre-trained deep neural networks. In SIAM International Conference on Data Mining . SIAM, 505–513
Charles H Martin and Michael W Mahoney. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Later among the works it cites.
Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Association for Computational Linguistics, 38–45
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Later among the works it cites.
On layer normalization in the transformer architecture. In International Conference on Machine Learning . PMLR, 10524–10533
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. 2020 · 2020
Later among the works it cites.
Boundary thickness and robustness in learning models
Yaoqing Yang, Rajiv Khanna, Yaodong Yu, Amir Gholami, Kurt Keutzer, Joseph E Gonzalez, Kannan Ramchandran, and Michael W Mahoney. 2020 · 2020
Later among the works it cites.
DIALOGPT: Large-Scale Generative Pre-training for Conversational Response Generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations . 270–278
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and William B Dolan. 2020 · 2020
Later among the works it cites.
Implicit Self-Regularization in Deep Neural Networks: Evidence from Random Matrix Theory and Implications for Learning
Charles H Martin and Michael W Mahoney. 2021a · 2021
Later among the works it cites.
Predicting trends in the quality of state-of-the-art neural networks without access to training or testing data
Charles H Martin, Tongsu Serena Peng, and Michael W Mahoney. 2021 · 2021
Later among the works it cites.
Taxonomizing local versus global structure in neural network loss landscapes. In Thirty-Fifth Conference on Neural Information Processing Systems
Yaoqing Yang, Liam Hodgkinson, Ryan Theisen, Joe Zou, Joseph E Gonzalez, Kannan Ramchandran, and Michael W Mahoney. 2021 · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Test accuracy vs. generalization gap: model selection in NLP without accessing training or testing data. In proceedings of the 29th ACM SIGKDD international conference on knowledge discovery & data mining
Yaoqing Yang, Ryan Theisen, Liam Hodgkinson, Joseph E Gonzalez, Kannan Ramchandran, Charles H Martin, and Michael W Mahoney. 2023 · 2023
Closest in time.