Fetching the paper…
Reading the bibliography…
A technical note aiming to offer deeper intuition for the LayerNorm function common in deep neural networks.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Beyond the spectral theorem: Decomposing arbitrary functions of nondiagonalizable operators
P. M. Riechers and J. P. Crutchfield · 2018
Earlier work this paper cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Cited alongside, same era.
On the expressivity role of LayerNorm in transformers’ attention
Shaked Brody, Uri Alon, and Eran Yahav · 2023
Cited alongside, same era.
Traveling words: A geometric interpretation of transformers
Raul Molina · 2023
Later among the works it cites.
https://pytorch.org/docs/stable/generated/torch.nn.LayerNorm.html
LayerNorm — PyTorch 2.3 documentation · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…