An unsolvable problem of elementary number theory
Alonzo Church · 1936
Earlier work this paper cites.
Scheme: an interpreter for extended lambda calculus
Gerald Jay Sussman · 1975
Earlier work this paper cites.
Direct and indirect effects
Judea Pearl · 2001
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik Tjong Kim Sang and Fien De Meulder · 2003
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Dependency-based word embeddings
Omer Levy and Yoav Goldberg · 2014
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Word translation without parallel data
Original
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou · 2017
Earlier work this paper cites.
Distinguishing antonyms and synonyms in a pattern-based neural network
Kim Anh Nguyen, Sabine Schulte im Walde, and Ngoc Thang Vu · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Learning to make analogies by contrasting abstract relational structure
Felix Hill, Adam Santoro, David Barrett, Ari Morcos, and Timothy Lillicrap · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni · 2018
Earlier work this paper cites.
Context-aware neural machine translation learns anaphora resolution
Elena Voita, Pavel Serdyukov, Rico Sennrich, and Ivan Titov · 2018
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Earlier work this paper cites.
Do attention heads in BERT track syntactic dependencies?
Original
Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R Bowman · 2019
Earlier work this paper cites.
Attention is not explanation
Sarthak Jain and Byron C Wallace · 2019
Earlier work this paper cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky · 2019
Earlier work this paper cites.
Open sesame: Getting inside BERT’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank · 2019
Earlier work this paper cites.
Visualizing and measuring the geometry of BERT
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Attention is not only a weight: Analyzing transformers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui · 2020
Earlier work this paper cites.