Fetching the paper…
Reading the bibliography…
After their successful debut in natural language processing, Transformer architectures are now becoming the de-facto standard in many domains.
Numerical operator calculus in higher dimensions
Beylkin, G. and Mohlenkamp, M. J · 2002
Earlier work this paper cites.
Multiresolution quantum chemistry in multiwavelet bases
Harrison, R. J., Fann, G. I., Yanai, T., and Beylkin, G · 2003
Earlier work this paper cites.
On the efficient evaluation of coalescence integrals in population balance models
Hackbusch, W · 2006
Earlier work this paper cites.
Multivariate regression and machine learning with sums of separable functions
Beylkin, G., Garcke, J., and Mohlenkamp, M. J · 2009
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2012
Earlier work this paper cites.
Japanese and korean voice search
Schuster, M. and Nakajima, K · 2012
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Inductive bias of deep convolutional networks through pooling geometry
Cohen, N. and Shashua, A · 2017
Earlier work this paper cites.
Analysis and design of convolutional networks via hierarchical tensor decompositions
Cohen, N., Sharir, O., Levine, Y., Tamari, R., Yakira, D., and Shashua, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Accurate de novo prediction of protein contact map by ultra-deep learning model
Wang, S., Sun, S., Li, Z., Zhang, R., and Xu, J · 2017
Earlier work this paper cites.
Benefits of depth for long-term memory of recurrent networks
Levine, Y., Sharir, O., Ziv, A., and Shashua, A · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
Breaking the softmax bottleneck: A high-rank RNN language model
Yang, Z., Dai, Z., Salakhutdinov, R., and Cohen, W. W · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Attention is not Explanation
Jain, S. and Wallace, B. C · 2019
Cited alongside, same era.
Quantum entanglement in deep learning architectures
Levine, Y., Sharir, O., Cohen, N., and Shashua, A · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
Rives, A., Meier, J., Sercu, T., Goyal, S., Lin, Z., Liu, J., Guo, D., Ott, M., Zitnick, C. L., Ma, J., and Fergus, R · 2019
Cited alongside, same era.
Analysing mathematical reasoning abilities of neural models
Saxton, D., Grefenstette, E., Hill, F., and Kohli, P · 2019
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2020
Later among the works it cites.
Limits to depth efficiencies of self-attention
Levine, Y., Wies, N., Sharir, O., Bata, H., and Shashua, A · 2020
Later among the works it cites.
Improving transformer models by reordering their sublayers
Press, O., Smith, N. A., and Levy, O · 2020
Later among the works it cites.
Learning to deceive with attention-based explanations
Pruthi, D., Gupta, M., Dhingra, B., Neubig, G., and Lipton, Z. C · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M · 2020
Cited alongside, same era.
Low-rank bottleneck in multi-head attention models
Bhojanapalli, S., Yun, C., Rawat, A. S., Reddi, S., and Kumar, S · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
On identifiability in transformers
Brunner, G., Liu, Y., Pascual, D., Richter, O., Ciaramita, M., and Wattenhofer, R · 2020
Cited alongside, same era.
Generative pretraining from pixels
Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I · 2020
Cited alongside, same era.
Richter, O. and Wattenhofer, R · 2020
Later among the works it cites.
Deep autoregressive models for the efficient variational simulation of many-body quantum systems
Sharir, O., Levine, Y., Wies, N., Carleo, G., and Shashua, A · 2020
Later among the works it cites.
Scaling autoregressive video models
Weissenborn, D., Täckström, O., and Uszkoreit, J · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Closest in time.
{PMI}-masking: Principled masking of correlated spans
Levine, Y., Lenz, B., Lieber, O., Abend, O., Leyton-Brown, K., Tennenholtz, M., and Shoham, Y · 2021
Closest in time.
Msa transformer
Rao, R., Liu, J., Verkuil, R., Meier, J., Canny, J. F., Abbeel, P., Sercu, T., and Rives, A · 2021
Closest in time.
Experimental t5 pre-trained model checkpoints
Shazeer, N · 2021
Closest in time.
mT5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2021
Closest in time.
Using the output embedding to improve language models
Press, O. and Wolf, L · 2025
Closest in time.