Fetching the paper…
Reading the bibliography…
Recent research suggests that the feed-forward module within Transformers can be viewed as a collection of key-value memories, where the keys learn to capture specific patterns from the input based on the training examples.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 1905
Earlier work this paper cites.
Investigating multilingual nmt representations at scale
Sneha Reddy Kudugunta, Ankur Bapna, Isaac Caswell, Naveen Arivazhagan, and Orhan Firat. 2019 · 1909
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019 · 1910
Earlier work this paper cites.
How language-neutral is multilingual bert?
Jindřich Libovický, Rudolf Rosa, and Alexander Fraser. 2019 · 1911
Earlier work this paper cites.
Language control in the bilingual brain
Jenny Crinion, Robert Turner, Alice Grogan, Takashi Hanakawa, Uta Noppeney, Joseph T Devlin, Toshihiko Aso, Shinichi Urayama, Hidenao Fukuyama, Katharine Stockton, et al. 2006 · 2006
Earlier work this paper cites.
Announcing czeng 2.0 parallel corpus with over 2 gigawords
Tom Kocmi, Martin Popel, and Ondrej Bojar. 2020 · 2007
Earlier work this paper cites.
Lexical processing in the bilingual brain: Evidence from grammatical/morphological deficits
Michele Miozzo, Albert Costa, Mireia Hernandez, and Brenda Rapp. 2010 · 2010
Earlier work this paper cites.
Speaking in multiple languages: Neural correlates of language proficiency in multilingual word production
Gerda Videsott, Bärbel Herrnberger, Klaus Hoenig, Edgar Schilly, Jo Grothe, Werner Wiater, Manfred Spitzer, and Markus Kiefer. 2010 · 2010
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
From zero to hero: On the limitations of zero-shot language transfer with multilingual transformers
Anne Lauscher, Vinit Ravishankar, Ivan Vulić, and Goran Glavaš. 2020 · 2020
Cited alongside, same era.
Rethinking the value of transformer components
Wenxuan Wang and Zhaopeng Tu. 2020 · 2020
Cited alongside, same era.
How linguistically fair are multilingual pre-trained language models?
Monojit Choudhury and Amit Deshpande. 2021 · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. 2022 · 2022
Later among the works it cites.
Large models are parsimonious learners: Activation sparsity in trained transformers
Zonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li, Ankit Singh Rawat, Sashank J Reddi, Ke Ye, Felix Chern, Felix Yu, Ruiqi Guo, et al. 2022 · 2022
Later among the works it cites.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Later among the works it cites.
Lifting the curse of multilinguality by pre-training modular transformers
Jonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, and Mikel Artetxe. 2022 · 2022
Later among the works it cites.
Kformer: Knowledge injection in transformer feed-forward layers
Yunzhi Yao, Shaohan Huang, Li Dong, Furu Wei, Huajun Chen, and Ningyu Zhang. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al. 2021 · 2021
Cited alongside, same era.
Analyzing the mono-and cross-lingual pretraining dynamics of multilingual language models
Terra Blevins, Hila Gonen, and Luke Zettlemoyer. 2022 · 2022
Cited alongside, same era.
When is bert multilingual? isolating crucial ingredients for cross-lingual transfer
Ameet Deshpande, Partha Talukdar, and Karthik Narasimhan. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Moefication: Transformer feed-forward layers are mixtures of experts
Zhengyan Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2022 · 2022
Later among the works it cites.
Eliciting latent predictions from transformers with the tuned lens
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023 · 2023
Closest in time.