Fetching the paper…
Reading the bibliography…
Humans have the ability to adapt the type of information they use, the procedure they employ, and the amount of time they spend when solving problems.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Hierarchical modularity in human brain functional networks
Meunier, D., Lambiotte, R., Fornito, A., Ersche, K., and Bullmore, E. T · 2009
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2012
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. V · 2012
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Graves, A · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2016
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C., Shrivastava, A., Singh, S., and Gupta, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, Ł · 2018
Earlier work this paper cites.
Graph transformer networks
Yun, S., Jeong, M., Kim, R., Kang, J., and Kim, H. J · 2019
Earlier work this paper cites.
Transferring inductive biases through knowledge distillation
Abnar, S., Dehghani, M., and Zuidema, W · 2020
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Earlier work this paper cites.
Burtsev, M. S., Kuratov, Y., Peganov, A., and Sapunov, G. V · 2020
Earlier work this paper cites.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Hahn, M · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Big transfer (bit): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N · 2020
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Xue, F., Shi, Z., Wei, F., Lou, Y., Liu, Y., and You, Y · 2021
Later among the works it cites.
Palbert: Teaching albert to ponder
Balagansky, N. and Gavrilov, D · 2022
Later among the works it cites.
End-to-end algorithm synthesis with recurrent networks: Logical extrapolation without overthinking
Bansal, A., Schwarzschild, A., Borgnia, E., Emam, Z., Huang, F., Goldblum, M., and Goldstein, T · 2022
Later among the works it cites.
Better plain vit baselines for imagenet-1k
Beyer, L., Zhai, X., and Kolesnikov, A · 2022
Later among the works it cites.
Adaptive token sampling for efficient vision transformers
Fayyaz, M., Koohpayegani, S. A., Rezaei, F., and Gall, S. H. P. J · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The right tool for the job: Matching model and instance complexities
Schwartz, R., Stanovsky, G., Swayamdipta, S., Dodge, J., and Smith, N. A · 2020
Cited alongside, same era.
Exploring the limits of large scale pre-training
Abnar, S., Dehghani, M., Neyshabur, B., and Sedghi, H · 2021
Cited alongside, same era.
Banino, A., Balaguer, J., and Blundell, C · 2021
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021
Fedus, W., Zoph, B., and Shazeer, N · 2021
Cited alongside, same era.
Highly accurate protein structure prediction with alphafold
Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., et al · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Polyvit: Co-training vision transformers on images, videos and audio
Likhosherstov, V., Arnab, A., Choromanski, K., Lucic, M., Tay, Y., Weller, A., and Dehghani, M · 2021
Cited alongside, same era.
Later among the works it cites.
A review of sparse expert models in deep learning
Fedus, W., Dean, J., and Zoph, B · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
A generalist neural algorithmic learner
Ibarz, B., Kurin, V., Papamakarios, G., Nikiforou, K., Bennani, M., Csordás, R., Dudzik, A., Bošnjak, M., Vitvitskyi, A., Rubanova, Y., et al · 2022
Later among the works it cites.
Adavit: Adaptive vision transformers for efficient image recognition
Meng, L., Li, H., Chen, B.-C., Lan, S., Wu, Z., Jiang, Y.-G., and Lim, S.-N · 2022
Later among the works it cites.
Confident adaptive language modeling
Schuster, T., Fisch, A., Gupta, J., Dehghani, M., Bahri, D., Tran, V. Q., Tay, Y., and Metzler, D · 2022
Later among the works it cites.
Scaling laws vs model architectures: How does inductive bias influence scaling?
Tay, Y., Dehghani, M., Abnar, S., Chung, H. W., Fedus, W., Rao, J., Narang, S., Tran, V. Q., Yogatama, D., and Metzler, D · 2022
Later among the works it cites.
The clrs algorithmic reasoning benchmark
Veličković, P., Badia, A. P., Budden, D., Pascanu, R., Banino, A., Dashevskiy, M., Hadsell, R., and Blundell, C · 2022
Later among the works it cites.
A-ViT: Adaptive tokens for efficient vision transformer
Yin, H., Vahdat, A., Alvarez, J., Mallya, A., Kautz, J., and Molchanov, P · 2022
Later among the works it cites.
Retrieval-enhanced machine learning
Zamani, H., Diaz, F., Dehghani, M., Metzler, D., and Bendersky, M · 2022
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2022
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A., Caron, M., Geirhos, R., Alabdulmohsin, I., et al · 2023
Closest in time.
Kumar, M., Dehghani, M., and Houlsby, N · 2023
Closest in time.
L2 norm guided adaptive computation
Shemiranifar, M. and Dehghani, M · 2023
Closest in time.