Fetching the paper…
Reading the bibliography…
The Transformer architecture has become prominent in developing large causal language models.
Bert rediscovers the classical nlp pipeline
Tenney, I., Das, D., and Pavlick, E. (2019) · 1905
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M. (2019) · 1911
Earlier work this paper cites.
GLU variants improve transformer
Shazeer, N. (2020) · 2002
Earlier work this paper cites.
Numerical Optimization
Nocedal, J. and Wright, S. (2006) · 2006
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L. (2010) · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E. (2010) · 2010
Earlier work this paper cites.
Topic modeling with contextualized word representation clusters
Thompson, L. and Mimno, D. (2020) · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. (2011) · 2011
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T. (2013) · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R. (2015) · 2015
Earlier work this paper cites.
A Kronecker-factored approximate Fisher matrix for convolution layers
Grosse, R. and Martens, J. (2016) · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R. (2017) · 2017
Earlier work this paper cites.
Optimization as a model for few-shot learning
Ravi, S. and Larochelle, H. (2017) · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Conneau, A., Kruszewski, G., Lample, G., Barrault, L., and Baroni, M. (2018) · 2018
Cited alongside, same era.
Bilevel programming for hyperparameter optimization and meta-learning
Franceschi, L., Frasconi, P., Salzo, S., Grazzi, R., and Pontil, M. (2018) · 2018
Cited alongside, same era.
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2019) · 2019
Cited alongside, same era.
Transformers as meta-learners for implicit neural representations
Chen, Y. and Wang, X. (2022) · 2022
Later among the works it cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A. (2022) · 2022
Later among the works it cites.
Locating and editing factual associations in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y. (2022) · 2022
Later among the works it cites.
Does bert rediscover a classical nlp pipeline?
Niu, J., Lu, W., and Penn, G. (2022) · 2022
Later among the works it cites.
Transductive decoupled variational inference for few-shot classification
Singh, A. and Jamali-Rad, H. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Cited alongside, same era.
Meta-curvature
Park, E. and Oliva, J. B. (2019) · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Cited alongside, same era.
Visualizing and measuring the geometry of bert
Reif, E., Yuan, A., Wattenberg, M., Viegas, F. B., Coenen, A., Pearce, A., and Kim, B. (2019) · 2019
Cited alongside, same era.
Root mean square layer normalization
Zhang, B. and Sennrich, R. (2019) · 2019
Cited alongside, same era.
Probing bert in hyperbolic spaces
Chen, B., Fu, Y., Xu, G., Xie, P., Tan, C., Chen, M., and Jing, L. (2021) · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Dirk Weissenborn, X. Z., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. (2021) · 2021
Cited alongside, same era.
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S. (2023) · 2023
Closest in time.
Why can GPT learn in-context? language models implicitly perform gradient descent as meta-optimizers
Dai, D., Sun, Y., Dong, L., Hao, Y., Ma, S., Sui, Z., and Wei, F. (2023) · 2023
Closest in time.
The dual form of neural networks revisited: Connecting test time predictions to training patterns via spotlights of attention
Irie, K., Csordás, R., and Schmidhuber, J. (2023) · 2023
Closest in time.
Summary of ChatGPT-related research and perspective towards the future of large language models
Liu, Y., Han, T., Ma, S., Zhang, J., Yang, Y., Tian, J., He, H., Li, A., He, M., Liu, Z., Wu, Z., Zhao, L., Zhu, D., Li, X., Qiang, N., Shen, D., Liu, T., and Ge, B. (2023) · 2023
Closest in time.
Emergent linear representations in world models of self-supervised sequence models
Nanda, N., Lee, A., and Wattenberg, M. (2023) · 2023
Closest in time.
OpenAI (2023) · 2023
Closest in time.
Why larger language models do in-context learning differently?
Shi, Z., Wei, J., Xu, Z., and Liang, Y. (2023) · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. (2023) · 2023
Closest in time.
Larger language models do in-context learning differently
Wei, J., Wei, J., Tay, Y., Tran, D., Webson, A., Lu, Y., Chen, X., Liu, H., Huang, D., Zhou, D., et al. (2023) · 2023
Closest in time.
Transformer-based causal language models perform clustering
Wu, X. and Varshney, L. R. (2024) · 2024
Closest in time.