Fetching the paper…
Reading the bibliography…
Recent studies on interpretability of attention distributions have led to notions of faithful and plausible explanations for a model's predictions.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 1905
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 1906
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov. 2019 · 1906
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 1908
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour. 1999 · 1999
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
A universal part-of-speech tagset
Slav Petrov, Dipanjan Das, and Ryan T. McDonald. 2011 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Pharmacovigilance from social media: mining adverse drug reaction mentions using sequence labeling with word embedding cluster features
Azadeh Nikfarjam, Abeed Sarker, Karen O’Connor, Rachel E. Ginn, and Graciela Gonzalez-Hernandez. 2015 · 2015
Cited alongside, same era.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, and Tomas Mikolov. 2015 · 2015
Cited alongside, same era.
Mimic-iii, a freely accessible critical care database
Alistair E. W. Johnson, Tom J. Pollard, Lu Shen, Li wei H. Lehman, Mengling Feng, Mohammad M. Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G. Mark. 2016 · 2016
Cited alongside, same era.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi S. Jaakkola. 2016 · 2016
Cited alongside, same era.
From softmax to sparsemax: A sparse model of attention and multi-label classification
André F. T. Martins and Ramón Fernández Astudillo. 2016 · 2016
Advances in pre-training distributed word representations
Tomas Mikolov, Edouard Grave, Piotr Bojanowski, Christian Puhrsch, and Armand Joulin. 2018 · 2018
Later among the works it cites.
Interpretable structure induction via sparse attention
Ben Peters, Vlad Niculae, and André F. T. Martins. 2018 · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
Exploring interpretable lstm neural networks over multi-variable data
Tian Guo, Tao Lin, and Nino Antulov-Fantulin. 2019 · 2019
Later among the works it cites.
Attention is not explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Later among the works it cites.
What does bert learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Diversity driven attention model for query-based abstractive summarization
Preksha Nema, Mitesh M. Khapra, Anirban Laha, and Balaraman Ravindran. 2017 · 2017
Cited alongside, same era.
A regularized framework for sparse and structured neural attention
Vlad Niculae and Mathieu Blondel. 2017 · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Deriving machine attention from human rationales
Yujia Bao, Shiyu Chang, Mo Yu, and Regina Barzilay. 2018 · 2018
Cited alongside, same era.
Towards understanding the geometry of knowledge graph embeddings
Chandrahas, Aditya Sharma, and Partha P. Talukdar. 2018 · 2018
Cited alongside, same era.
Sparse and constrained attention for neural machine translation
Chaitanya Malaviya, Pedro Ferreira, and André F. T. Martins. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Selective attention for context-aware neural machine translation
Sameen Maruf, André F. T. Martins, and Gholamreza Haffari. 2019 · 2019
Later among the works it cites.
Re-evaluating adem: A deeper look at scoring dialogue responses
Ananya Sai, Mithun Das Gupta, Mitesh M. Khapra, and Mukundhan Srinivasan. 2019 · 2019
Later among the works it cites.
Is attention interpretable?
Sofia Serrano and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Ssn: Learning sparse switchable normalization via sparsestmax
Wenqi Shao, Tianjian Meng, Jingyu Li, Ruimao Zhang, Yudian Li, Xiaogang Wang, and Ping Luo. 2019 · 2019
Later among the works it cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
Empirical study of transformer’s attention mechanism via the lens of kernel
Yao-Hung Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.