Fetching the paper…
Reading the bibliography…
Feed-forward layers constitute two-thirds of a transformer model's parameters, yet their role in the network remains under-explored.
Augmenting self-attention with persistent memory
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin. 2019 · 1907
Earlier work this paper cites.
Transformer with depth-wise lstm
Hongfei Xu, Qiuhui Liu, Deyi Xiong, and Josef van Genabith. 2020 · 2007
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent. 2009 · 2009
Earlier work this paper cites.
End-to-end memory networks
S. Sukhbaatar, J. Weston, and R. Fergus. 2015 · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Understanding convolutional neural networks for text classification
Alon Jacovi, Oren Sar Shalom, and Yoav Goldberg. 2018 · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli. 2019 · 2019
Earlier work this paper cites.
What does BERT look at? An analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Cited alongside, same era.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Closest in time.
Analyzing individual neurons in pre-trained language models
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov. 2020 · 2020
Closest in time.
Explaining black box predictions and unveiling data artifacts through influence functions
Xiaochuang Han, Byron C. Wallace, and Yulia Tsvetkov. 2020 · 2020
Closest in time.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning
Milad Nasr, Reza Shokri, and Amir Houmansadr. 2019 · 2019
Cited alongside, same era.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Cited alongside, same era.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov. 2019 · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
Improving transformer models by reordering their sublayers
Ofir Press, Noah A. Smith, and Omer Levy. 2020 · 2020
Closest in time.
Tx-ray: Quantifying and explaining model-knowledge transfer in (un-) supervised nlp
Nils Rethmeier, Vageesh Kumar Saxena, and Isabelle Augenstein. 2020 · 2020
Closest in time.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020 · 2020
Closest in time.
Attention-based neural beamforming layers for multi-channel speech recognition
Bhargav Pulugundla, Yang Gao, Brian King, Gokce Keskin, Harish Mallidi, Minhua Wu, Jasha Droppo, and Roland Maas. 2021 · 2021
Closest in time.