Fetching the paper…
Reading the bibliography…
The transformer architecture by Vaswani et al.
Language Models are Few-Shot Learners, July 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Convolutional neural networks at constrained time cost
He, K. and Sun, J · 2015
Earlier work this paper cites.
Srivastava, R. K., Greff, K., and Schmidhuber, J · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., van der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Openwebtext corpus
Gokaslan, A. and Cohen, V · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Fast transformer decoding: One write-head is all you need
Shazeer, N · 2019
Cited alongside, same era.
OpenWebText2 dataset, as part of ‘the Pile: An 800gb dataset of diverse text for language modeling‘
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Cited alongside, same era.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Cited alongside, same era.
Depth-wise attention (dwatt): A layer fusion method for data-efficient classification
ElNokrashy, M., AlKhamissi, B., and Diab, M · 2022
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Later among the works it cites.
Will we run out of data? an analysis of the limits of scaling datasets in machine learning
Villalobos, P., Sevilla, J., Heim, L., Besiroglu, T., Hobbhahn, M., and Ho, A · 2022
Later among the works it cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 2020
Cited alongside, same era.
GLU variants improve transformer
Shazeer, N · 2020
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B. and Komatsuzaki, A · 2021
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
The impact of depth and width on transformer language model generalization
Petty, J., van Steenkiste, S., Dasgupta, I., Sha, F., Garrette, D., and Linzen, T
Cited in the paper.
The impact of depth and width on transformer language model generalization
Petty, J., van Steenkiste, S., Dasgupta, I., Sha, F., Garrette, D., and Linzen, T
Cited in the paper.
He, B. and Hofmann, T · 2023
Later among the works it cites.
Encouraging divergent thinking in large language models through multi-agent debate
Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Tu, Z., and Shi, S · 2023
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2023
Later among the works it cites.
Cotformer: More tokens with attention make up for less depth
Mohtashami, A., Pagliardini, M., and Jaggi, M · 2023
Later among the works it cites.
GPT-4 Technical Report, March 2023
OpenAI · 2023
Later among the works it cites.
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J · 2023
Later among the works it cites.