Fetching the paper…
Reading the bibliography…
Most neural networks utilize the same amount of compute for every example independent of the inherent complexity of the input.
Lingvo: a modular and scalable framework for sequence-to-sequence modeling
Shen, J., Nguyen, P., Wu, Y., Chen, Z., Chen, M. X., Jia, Y., Kannan, A., Sainath, T., Cao, Y., Chiu, C.-C., et al · 1902
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
Spall, J. C. et al · 1992
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A. C · 2013
Earlier work this paper cites.
Exponentially increasing the capacity-to-computation ratio for conditional computation in deep learning, 2014
Cho, K. and Bengio, Y · 2014
Earlier work this paper cites.
Low-rank approximations for conditional feedforward computation in deep neural networks
Davis, A. S. and Arel, I · 2014
Earlier work this paper cites.
Mixture of experts: A literature survey
Masoudnia, S. and Ebrahimpour, R · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks, 2016
Graves, A · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Earlier work this paper cites.
Chiu, C.-C. and Raffel, C · 2017
Cited alongside, same era.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Johnson, M., Schuster, M., Le, Q. V., Krikun, M., Wu, Y., Chen, Z., Thorat, N., Viégas, F., Wattenberg, M., Corrado, G., et al · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Cited alongside, same era.
Online and linear-time attention by enforcing monotonic alignments
Raffel, C., Luong, M.-T., Liu, P. J., Weiss, R. J., and Eck, D · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Monotonic infinite lookback attention for simultaneous machine translation
Arivazhagan, N., Cherry, C., Macherey, W., Chiu, C.-C., Yavuz, S., Pang, R., Li, W., and Raffel, C · 2019
Later among the works it cites.
Depth-adaptive transformer, 2019
Elbayad, M., Gu, J., Grave, E., and Auli, M · 2019
Later among the works it cites.
Reducing transformer depth on demand with structured dropout, 2019
Fan, A., Grave, E., and Joulin, A · 2019
Later among the works it cites.
Text repair model for neural machine translation
Freitag, M., Caswell, I., and Roy, S · 2019
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, L · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Cited alongside, same era.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Cited alongside, same era.
Tensor2tensor for neural machine translation
Vaswani, A., Bengio, S., Brevdo, E., Chollet, F., Gomez, A. N., Gouws, S., Jones, L., Kaiser, L., Kalchbrenner, N., Parmar, N., Sepassi, R., Shazeer, N., and Uszkoreit, J · 2018
Cited alongside, same era.
Conditional computation in neural networks for faster models
Bengio, E., Bacon, P.-L., Pineau, J., and Precup, D
Cited in the paper.
Large memory layers with product keys
Lample, G., Sablayrolles, A., Ranzato, M., Denoyer, L., and Jégou, H · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I · 2019
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Later among the works it cites.