Fetching the paper…
Reading the bibliography…
It has become standard to solve NLP tasks by fine-tuning pre-trained language models (LMs), especially in low-data settings.
Extensions of lipschitz mappings into a hilbert space
Johnson, W. B · 1984
Earlier work this paper cites.
Building a question answering test collection
Voorhees, E. M. and Tice, D. M · 2000
Earlier work this paper cites.
Mining and summarizing customer reviews
Hu, M. and Liu, B · 2004
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Pang, B. and Lee, L · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Pang, B. and Lee, L · 2005
Earlier work this paper cites.
Annotating expressions of opinions and emotions in language
Wiebe, J., Wilson, T., and Cardie, C · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Bar Haim, R., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., Magnini, B., and Szpektor, I · 2006
Earlier work this paper cites.
Tensor programs ii: Neural tangent kernel for any architecture
Yang, G · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B · 2007
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2008
Earlier work this paper cites.
The fifth PASCAL recognizing textual entailment challenge
Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D · 2009
Earlier work this paper cites.
Tensor programs iii: Neural matrix laws
Yang, G · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Measuring the intrinsic dimension of objective landscapes
Li, C., Farkhoor, H., Liu, R., and Yosinski, J · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Earlier work this paper cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks, 2018
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z., Li, Y., and Liang, Y · 2019
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
A theoretical analysis of contrastive unsupervised representation learning
Saunshi, N., Plevrakis, O., Arora, S., Khodak, M., and Khandeparkar, H · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
Wide feedforward or recurrent neural networks of any architecture are gaussian processes
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Later among the works it cites.
Predicting what you already know helps: Provable self-supervised learning
Lee, J. D., Lei, Q., Saunshi, N., and ZHUO, J · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Later among the works it cites.
Fast adaptation with linearized neural networks
Maddox, W., Tang, S., Moreno, P., Gordon Wilson, A., and Damianou, A · 2021
Later among the works it cites.
A mathematical exploration of why language models help solve downstream tasks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang, G · 2019
Cited alongside, same era.
Harnessing the power of infinitely wide deep nets on small-data tasks
Arora, S., Du, S. S., Li, Z., Salakhutdinov, R., Wang, R., and Yu, D · 2020
Cited alongside, same era.
Electra: Pre-training text encoders as discriminators rather than generators
Clark, K., Luong, M.-T., Le, Q. V., and Manning, C. D · 2020
Cited alongside, same era.
SpanBERT: Improving pre-training by representing and predicting spans
Joshi, M., Chen, D., Liu, Y., Weld, D. S., Zettlemoyer, L., and Levy, O · 2020
Cited alongside, same era.
Understanding the difficulty of training transformers
Liu, L., Liu, X., Gao, J., Chen, W., and Han, J · 2020
Cited alongside, same era.
Gradients as features for deep representation learning
Mu, F., Liang, Y., and Li, Y · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al · 2020
Cited alongside, same era.
Saunshi, N., Malladi, S., and Arora, S · 2021
Later among the works it cites.
Exploiting cloze-questions for few-shot text classification and natural language inference
Schick, T. and Schütze, H · 2021
Later among the works it cites.
Self-supervised learning from a multi-view perspective
Tsai, Y.-H. H., Wu, Y., Salakhutdinov, R., and Morency, L.-P · 2021
Later among the works it cites.
Tensor programs iv: Feature learning in infinite-width neural networks
Yang, G. and Hu, E. J · 2021
Later among the works it cites.
Tensor programs iib: Architectural universality of neural tangent kernel training dynamics
Yang, G. and Littwin, E · 2021
Later among the works it cites.
BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Ben Zaken, E., Goldberg, Y., and Ravfogel, S · 2022
Closest in time.
Evolved optimizer for vision
Chen, X., Liang, C., Huang, D., Real, E., Liu, Y., Wang, K., Hsieh, C.-J., Lu, Y., and Le, Q. V · 2022
Closest in time.
Learning with asymmetric kernels: Least squares and feature interpretation, 2022
He, M., He, F., Shi, L., Huang, X., and Suykens, J. A. K · 2022
Closest in time.
Robust training of neural networks using scale invariant architectures
Li, Z., Bhojanapalli, S., Zaheer, M., Reddi, S., and Kumar, S · 2022
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G · 2022
Closest in time.
Cutting down on prompts and parameters: Simple few-shot learning with language models
Logan IV, R., Balazevic, I., Wallace, E., Petroni, F., Singh, S., and Riedel, S · 2022
Closest in time.
A qualitative study of the dynamic behavior for adaptive gradient algorithms
Ma, C., Wu, L., and E, W · 2022
Closest in time.
On the sdes and scaling rules for adaptive gradient algorithms, 2022
Malladi, S., Lyu, K., Panigrahi, A., and Arora, S · 2022
Closest in time.
Fast finite width neural tangent kernel
Novak, R., Sohl-Dickstein, J., and Schoenholz, S. S · 2022
Closest in time.
Understanding contrastive learning requires incorporating inductive biases
Saunshi, N., Ash, J., Goel, S., Misra, D., Zhang, C., Arora, S., Kakade, S., and Krishnamurthy, A · 2022
Closest in time.
More than a toy: Random matrix models predict how real-world neural representations generalize
Wei, A., Hu, W., and Steinhardt, J · 2022
Closest in time.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J · 2022
Closest in time.
Adaptive optimization in the $\infty$-width limit
Littwin, E. and Yang, G · 2023
Closest in time.