Fetching the paper…
Reading the bibliography…
We break the linear link between the layer size and its inference cost by introducing the fast feedforward (FFF) architecture, a log-time alternative to feedforward networks.
Multidimensional binary search trees used for associative searching
Bentley, J. L. 1975 · 1975
Earlier work this paper cites.
A database for handwritten text recognition research
Hull, J. J. 1994 · 1994
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D.; Lee, H.; Xu, Y.; Chen, D.; Firat, O.; Huang, Y.; Krikun, M.; Shazeer, N.; and Chen, Z. 2020 · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A.; Hinton, G.; et al. 2009 · 2009
Earlier work this paper cites.
MNIST handwritten digit database
LeCun, Y.; Cortes, C.; and Burges, C. 2010 · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y.; Wang, T.; Coates, A.; Bissacco, A.; Wu, B.; and Ng, A. Y. 2011 · 2011
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y.; Léonard, N.; and Courville, A. 2013 · 2013
Cited alongside, same era.
Low-rank approximations for conditional feedforward computation in deep neural networks
Davis, A.; and Arel, I. 2013 · 2013
Cited alongside, same era.
Conditional computation in neural networks for faster models
Bengio, E.; Bacon, P.-L.; Pineau, J.; and Precup, D. 2015 · 2015
Cited alongside, same era.
Dynamic capacity networks
Almahairi, A.; Ballas, N.; Cooijmans, T.; Zheng, Y.; Larochelle, H.; and Courville, A. 2016 · 2016
Cited alongside, same era.
Decision forests, convolutional networks and the models in-between
Ioannou, Y.; Robertson, D.; Zikic, D.; Kontschieder, P.; Shotton, J.; Brown, M.; and Criminisi, A. 2016 · 2016
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Later among the works it cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H.; Rasul, K.; and Vollgraf, R. 2017 · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Later among the works it cites.
Rewriting a deep generative model
Bau, D.; Liu, S.; Wang, T.; Zhu, J.-Y.; and Torralba, A. 2020 · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W.; Zoph, B.; and Shazeer, N. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q.; Hinton, G.; and Dean, J. 2017 · 2017
Cited alongside, same era.