Fetching the paper…
Reading the bibliography…
Self-attention architectures, which are rapidly pushing the frontier in natural language processing, demonstrate a surprising depth-inefficient behavior: previous works indicate that increasing the internal representation (network width) is just as useful as increasing the number of self-attention layers (network depth).
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 1910
Earlier work this paper cites.
Inequalities
Godfrey Harold Hardy, John Edensor Littlewood, and George Pólya · 1952
Earlier work this paper cites.
Numerical operator calculus in higher dimensions
Gregory Beylkin and Martin J Mohlenkamp · 2002
Earlier work this paper cites.
Multiresolution quantum chemistry in multiwavelet bases
Robert J Harrison, George I Fann, Takeshi Yanai, and Gregory Beylkin · 2003
Earlier work this paper cites.
The zero set of a polynomial
Richard Caron and Tim Traynor · 2005
Earlier work this paper cites.
On the efficient evaluation of coalescence integrals in population balance models
Wolfgang Hackbusch · 2006
Earlier work this paper cites.
Multivariate regression and machine learning with sums of separable functions
Gregory Beylkin, Jochen Garcke, and Martin J Mohlenkamp · 2009
Earlier work this paper cites.
Low-rank matrix approximation using point-wise operators
Arash Amini, Amin Karbasi, and Farokh Marvasti · 2012
Earlier work this paper cites.
Tensor spaces and numerical tensor calculus , volume 42
Wolfgang Hackbusch · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
On the complexity of neural network classifiers: A comparison between shallow and deep architectures
Monica Bianchini and Franco Scarselli · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata · 2016
Earlier work this paper cites.
On the expressive power of deep learning: A tensor analysis
Nadav Cohen, Or Sharir, and Amnon Shashua · 2016
Earlier work this paper cites.
A cheap linear attention mechanism with fast lookups and fixed-size representations
Alexandre de Brébisson and Pascal Vincent · 2016
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
Ankur P Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit · 2016
Cited alongside, same era.
Tractable generative convolutional arithmetic circuits
Or Sharir, Ronen Tamari, Nadav Cohen, and Amnon Shashua · 2016
Cited alongside, same era.
Residual networks behave like ensembles of relatively shallow networks
Expressive numbers of two or more hidden layer relu neural networks
K. Inoue · 2019
Later among the works it cites.
Attention is not explanation
Sarthak Jain and Byron C. Wallace · 2019
Later among the works it cites.
Quantum entanglement in deep learning architectures
Yoav Levine, Or Sharir, Nadav Cohen, and Amnon Shashua · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Later among the works it cites.
Improving transformer models by reordering their sublayers
Ofir Press, Noah A Smith, and Omer Levy · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andreas Veit, Michael J Wilber, and Serge Belongie · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Inductive bias of deep convolutional networks through pooling geometry
Nadav Cohen and Amnon Shashua · 2017
Cited alongside, same era.
Depth separation for neural networks
Amit Daniely · 2017
Cited alongside, same era.
A structured self-attentive sentence embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio · 2017
Cited alongside, same era.
The expressive power of neural networks: A view from the width
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Cited alongside, same era.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher · 2017
Cited alongside, same era.
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C Lipton · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V Le · 2019
Later among the works it cites.
Wider or deeper: Revisiting the resnet model for visual recognition
Zifeng Wu, Chunhua Shen, and Anton Van Den Hengel · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Closest in time.
On identifiability in transformers
Gino Brunner, Yang Liu, Damian Pascual Ortiz, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer · 2020
Closest in time.
Depth-width trade-offs for neural networks via topological entropy
Kaifeng Bu, Yaobo Zhang, and Qingxian Luo · 2020
Closest in time.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning · 2020
Closest in time.
Quasi-equivalence of width and depth of neural networks
Fenglei Fan, Rongjie Lai, and Ge Wang · 2020
Closest in time.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Closest in time.
Thao Nguyen, Maithra Raghu, and Simon Kornblith · 2020
Closest in time.
Normalized attention without probability cage
Oliver Richter and Roger Wattenhofer · 2020
Closest in time.
Turing-NLG: A 17-billion-parameter language model by microsoft
Corby Rosset · 2020
Closest in time.
Off-policy recommendation system without exploration
Chengwei Wang, Tengfei Zhou, Chen Chen, Tianlei Hu, and Gang Chen · 2020
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Pmi-masking: Principled masking of correlated spans
Yoav Levine, Barak Lenz, Opher Lieber, Omri Abend, Kevin Leyton-Brown, Moshe Tennenholtz, and Yoav Shoham · 2021
Closest in time.