Fetching the paper…
Reading the bibliography…
The performance of fine-tuning pre-trained language models largely depends on the hyperparameter configuration.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2002
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
James S Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay. 2011 · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio. 2012 · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams. 2012 · 2012
Earlier work this paper cites.
Almost optimal exploration in multi-armed bandits
Zohar Karnin, Tomer Koren, and Oren Somekh. 2013 · 2013
Earlier work this paper cites.
Multi-task bayesian optimization
Kevin Swersky, Jasper Snoek, and Ryan P Adams. 2013 · 2013
Earlier work this paper cites.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals. 2014 · 2014
Earlier work this paper cites.
Efficient hyper-parameter optimization for nlp applications
Lidan Wang, Minwei Feng, Bowen Zhou, Bing Xiang, and Sridhar Mahadevan. 2015 · 2015
Earlier work this paper cites.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al. 2017 · 2017
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar. 2017 · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
Samuel L. Smith, Pieter-Jan Kindermans, and Quoc V. Le. 2017 · 2017
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L. Smith and Quoc V. Le. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
Bayesian hyperparameter optimization: overfitting, ensembles and conditional spaces
Julien-Charles Lévesque. 2018 · 2018
Cited alongside, same era.
Tune: A research platform for distributed model selection and training
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Later among the works it cites.
To tune or not to tune? adapting pretrained representations to diverse tasks
Matthew E. Peters, Sebastian Ruder, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Later among the works it cites.
Funnel-transformer: Filtering out sequential redundancy for efficient language processing
Zihang Dai, Guokun Lai, Yiming Yang, and Quoc Le. 2020 · 2020
Later among the works it cites.
SMART: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization
Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Richard Liaw, Eric Liang, Robert Nishihara, Philipp Moritz, Joseph E Gonzalez, and Ion Stoica. 2018 · 2018
Cited alongside, same era.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Cited alongside, same era.
Sentence encoders on STILTs: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Optuna: A next-generation hyperparameter optimization framework
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Show your work: Improved reporting of experimental results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Amog Kamsetty. 2020 · 2020
Later among the works it cites.
A system for massively parallel hyperparameter tuning
Liam Li, Kevin Jamieson, Afshin Rostamizadeh, Ekaterina Gonina, Jonathan Ben-tzur, Moritz Hardt, Benjamin Recht, and Ameet Talwalkar. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Understanding and robustifying differentiable architecture search
Arber Zela, Thomas Elsken, Tonmoy Saikia, Yassine Marrakchi, Thomas Brox, and Frank Hutter. 2020 · 2020
Later among the works it cites.
Reproducible and efficient benchmarks for hyperparameter optimization of neural machine translation systems
Xuan Zhang and Kevin Duh. 2020 · 2020
Later among the works it cites.
DEBERTA: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Closest in time.
Economic hyperparameter optimization with blended search strategy
Chi Wang, Qingyun Wu, Silu Huang, and Amin Saied. 2021 · 2021
Closest in time.
Revisiting few-sample BERT fine-tuning
Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q Weinberger, and Yoav Artzi. 2021 · 2021
Closest in time.