Fetching the paper…
Reading the bibliography…
Transformers have significantly advanced the field of natural language processing, but comprehending their internal mechanisms remains a challenge.
Dispersion on a sphere
Ronald Aylmer Fisher · 1953
Earlier work this paper cites.
Laplacian eigenmaps and spectral techniques for embedding and clustering
Mikhail Belkin and Partha Niyogi · 2001
Earlier work this paper cites.
The corpus of contemporary american english as the first reliable monitor corpus of english
Mark Davies · 2010
Earlier work this paper cites.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Residual connections encourage iterative inference
Stanislaw Jastrzebski, Devansh Arpit, Nicolas Ballas, Vikas Verma, Tong Che, and Yoshua Bengio · 2017
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction
Leland McInnes, John Healy, and James Melville · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Cited alongside, same era.
Interpreting gpt: The logit lens
nostalgebraist · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Cited alongside, same era.
Initialization is critical for preserving global data structure in both t-sne and umap
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg · 2022
Later among the works it cites.
Re-examining layernorm
Eric Windsor · 2022
Later among the works it cites.
The singular value decompositions of transformer weight matrices are highly interpretable
Beren Millidge and Sid Black · 2022
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Closest in time.
Palm 2 technical report
Google · 2023
Closest in time.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dmitry Kobak and George C Linderman · 2021
Cited alongside, same era.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Cited alongside, same era.
Detrs with collaborative hybrid assignments training
Zhuofan Zong, Guanglu Song, and Yu Liu · 2022
Cited alongside, same era.
Analyzing transformers in embedding space
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant · 2022
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al · 2023
Closest in time.
On the expressivity role of layernorm in transformers’ attention
Shaked Brody, Uri Alon, and Eran Yahav · 2023
Closest in time.
Eliciting latent predictions from transformers with the tuned lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Closest in time.
Github - karpathy/nanogpt: The simplest, fastest repository for training/finetuning medium-sized gpts., 2023
Andrej Karpathy · 2023
Closest in time.