Fetching the paper…
Reading the bibliography…
Extremely large pre-trained language models (PTMs) such as GPT-3 are usually released as a service.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Support vector machines for multi-class pattern recognition
Weston, J. and Watkins, C · 1999
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Hansen, N. and Ostermeier, A · 2001
Earlier work this paper cites.
Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (CMA-ES)
Hansen, N., Müller, S. D., and Koumoutsakos, P · 2003
Earlier work this paper cites.
Optimization by direct search: New perspectives on some classical and modern methods
Kolda, T. G., Lewis, R. M., and Torczon, V · 2003
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
Introduction to Derivative-Free Optimization
Conn, A. R., Scheinberg, K., and Vicente, L. N · 2009
Earlier work this paper cites.
Practical Bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Derivative-free optimization: A review of algorithms and comparison of software implementations
Rios, L. M. and Sahinidis, N. V · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J. J., and LeCun, Y · 2015
Earlier work this paper cites.
The CMA evolution strategy: A tutorial
Hansen, N · 2016
Earlier work this paper cites.
Derivative-free optimization of high-dimensional non-convex functions by sequential random embeddings
Qian, H., Hu, Y., and Yu, Y · 2016
Earlier work this paper cites.
Taking the human out of the loop: A review of Bayesian optimization
Shahriari, B., Swersky, K., Wang, Z., Adams, R. P., and de Freitas, N · 2016
Earlier work this paper cites.
Bayesian optimization in a billion dimensions via random embeddings
Wang, Z., Hutter, F., Zoghi, M., Matheson, D., and de Freitas, N · 2016
Earlier work this paper cites.
Sequential classification-based optimization for direct policy search
Hu, Y.-Q., Qian, H., and Yu, Y · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I · 2017
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes
Li, C., Farkhoor, H., Liu, R., and Yosinski, J · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. R · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2021
Later among the works it cites.
Making pre-trained language models better few-shot learners
Gao, T., Fisch, A., and Chen, D · 2021
Later among the works it cites.
PPT: pre-trained prompt tuning for few-shot learning
Gu, Y., Han, X., Liu, Z., and Huang, M · 2021
Later among the works it cites.
WARP: word-level adversarial reprogramming
Hambardzumyan, K., Khachatrian, H., and May, J · 2021
Later among the works it cites.
Towards a unified view of parameter-efficient transfer learning
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
To tune or not to tune? adapting pretrained representations to diverse tasks
Peters, M. E., Ruder, S., and Smith, N. A · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Re-examining linear embeddings for high-dimensional Bayesian optimization
Letham, B., Calandra, R., Rai, A., and Bakshy, E · 2020
Cited alongside, same era.
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Later among the works it cites.
True few-shot learning with language models
Perez, E., Kiela, D., and Cho, K · 2021
Later among the works it cites.
Learning how to ask: Querying lms with mixtures of soft prompts
Qin, G. and Eisner, J · 2021
Later among the works it cites.
Exploring low-dimensional intrinsic task subspace via prompt tuning
Qin, Y., Wang, X., Su, Y., Lin, Y., Ding, N., Liu, Z., Li, J., Hou, L., Li, P., Sun, M., and Zhou, J · 2021
Later among the works it cites.
Exploiting cloze-questions for few-shot text classification and natural language inference
Schick, T. and Schütze, H · 2021
Later among the works it cites.
It’s not just size that matters: Small language models are also few-shot learners
Schick, T. and Schütze, H · 2021
Later among the works it cites.
ERNIE 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation
Sun, Y., Wang, S., Feng, S., Ding, S., Pang, C., Shang, J., Liu, J., Chen, X., Zhao, Y., Lu, Y., Liu, W., Wu, Z., Gong, W., Liang, J., Shang, Z., Sun, P., Liu, W., Ouyang, X., Yu, D., Tian, H., Wu, H., and Wang, H · 2021
Later among the works it cites.
Yuan 1.0: Large-scale pre-trained language model in zero-shot and few-shot learning
Wu, S., Zhao, X., Yu, T., Zhang, R., Shen, C., Liu, H., Li, F., Zhu, H., Luo, J., Xu, L., and Zhang, X · 2021
Later among the works it cites.
Zeng, W., Ren, X., Su, T., Wang, H., Liao, Y., Wang, Z., Jiang, X., Yang, Z., Wang, K., Zhang, X., Li, C., Gong, Z., Yao, Y., Huang, X., Wang, J., Yu, J., Guo, Q., Yu, Y., Zhang, Y., Wang, J., Tao, H., Yan, D., Yi, Z., Peng, F., Jiang, F., Zhang, H., Deng, L., Zhang, Y., Lin, Z., Zhang, C., Zhang, S., Guo, M., Gu, S., Fan, G., Wang, Y., Jin, X., Liu, Q., and Tian, Y · 2021
Later among the works it cites.
Revisiting few-sample BERT fine-tuning
Zhang, T., Wu, F., Katiyar, A., Weinberger, K. Q., and Artzi, Y · 2021
Later among the works it cites.
Factual probing is [MASK]: learning vs. learning to recall
Zhong, Z., Friedman, D., and Chen, D · 2021
Later among the works it cites.
Paradigm shift in natural language processing
Sun, T., Liu, X., Qiu, X., and Huang, X · 2022
Closest in time.