Fetching the paper…
Reading the bibliography…
The decoder-only Transformer architecture with causal masking and relative position encoding (RPE) has become the de facto choice in language modeling.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Strong data-processing inequalities for channels and bayesian networks, 2016
Polyanskiy, Y. and Wu, Y · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
On controllable sparse alternatives to softmax
Laha, A., Chemmengath, S. A., Agrawal, P., Khapra, M., Sankaranarayanan, K., and Ramaswamy, H. G · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language, 2019
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Kenton, J. D. M.-W. C. and Toutanova, L. K · 2019
Earlier work this paper cites.
Rethinking softmax cross-entropy loss for adversarial robustness
Pang, T., Xu, K., Dong, Y., Du, C., Chen, N., and Zhu, J · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Improving latent alignment in text summarization by generalizing the pointer generator
Shen, X., Zhao, Y., Su, H., and Klakow, D · 2019
Earlier work this paper cites.
Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering
Yadav, V., Bethard, S., and Surdeanu, M · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Rethinking positional encoding in language pre-training
Ke, G., He, D., and Liu, T.-Y · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Are transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S · 2020
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Your transformer may not be as powerful as you expect
Luo, S., Li, S., Zheng, S., Liu, T.-Y., Wang, L., and He, D · 2022
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N., and Lewis, M · 2022
Later among the works it cites.
Quantizable transformers: Removing outliers by helping attention heads do nothing
Bondarenko, Y., Nagel, M., and Blankevoort, T · 2023
Later among the works it cites.
Vision transformers need registers
Darcet, T., Oquab, M., Mairal, J., and Bojanowski, P · 2023
Later among the works it cites.
The minipile challenge for data-efficient language models, 2023
Kaddour, J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Park, S., Yun, C., Lee, J., and Shin, J · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2021
Cited alongside, same era.
Ast-transformer: Encoding abstract syntax trees efficiently for code summarization
Tang, Z., Li, C., Ge, J., Shen, X., Zhu, Z., and Luo, B · 2021
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
How much does attention actually attend? questioning the importance of attention in pretrained transformers, 2022
Hassid, M., Peng, H., Rotem, D., Kasai, J., Montero, I., Smith, N. A., and Schwartz, R · 2022
Cited alongside, same era.
Transformer quality in linear time
Hua, W., Dai, Z., Liu, H., and Le, Q · 2022
Cited alongside, same era.
The impact of positional encoding on length generalization in transformers, 2023
Kazemnejad, A., Padhi, I., Ramamurthy, K. N., Das, P., and Reddy, S · 2023
Later among the works it cites.
Provable memorization capacity of transformers
Kim, J., Kim, M., and Mozafari, B · 2023
Later among the works it cites.
Bidirectional language models are also few-shot learners
Patel, A., Li, B., Rasooli, M. S., Constant, N., Raffel, C., and Callison-Burch, C · 2023
Later among the works it cites.
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Later among the works it cites.