Fetching the paper…
Reading the bibliography…
Attention mechanisms have seen wide adoption in neural NLP models.
Deriving machine attention from human rationales
Yujia Bao, Shiyu Chang, Mo Yu, and Regina Barzilay. 2018 · 1913
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Pharmacovigilance from social media: mining adverse drug reaction mentions using sequence labeling with word embedding cluster features
Azadeh Nikfarjam, Abeed Sarker, Karen O’Connor, Rachel Ginn, and Graciela Gonzalez. 2015 · 2015
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart van Merrië\parnboer, Armand Joulin, and Tomas Mikolov. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Retain: An interpretable predictive model for healthcare using reverse time attention mechanism
Edward Choi, Mohammad Taha Bahadori, Jimeng Sun, Joshua Kulas, Andy Schuetz, and Walter Stewart. 2016 · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, H Lehman Li-wei, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016 · 2016
Cited alongside, same era.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Cited alongside, same era.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016 · 2016
Cited alongside, same era.
The mythos of model interpretability
Zachary C Lipton. 2016 · 2016
Cited alongside, same era.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Andre Martins and Ramon Astudillo. 2016 · 2016
Cited alongside, same era.
Yoon Kim, Carl Denton, Luong Hoang, and Alexander M Rush. 2017 · 2017
Later among the works it cites.
Interpretable neural models for natural language processing
Tao Lei et al. 2017 · 2017
Later among the works it cites.
Right for the right reasons: training differentiable models by constraining their explanations
Andrew Slavin Ross, Michael C Hughes, and Finale Doshi-Velez. 2017 · 2017
Later among the works it cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
An interpretable knowledge transfer model for knowledge base completion
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human versus machine attention in document classification: A dataset with crowdsourced annotations
Nikolaos Pappas and Andrei Popescu-Belis. 2016 · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
Ankur Parikh, Oscar Tä\parckströ\parm, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Cited alongside, same era.
Bidirectional attention flow for machine comprehension
Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2016 · 2016
Cited alongside, same era.
Dynamic coattention networks for question answering
Caiming Xiong, Victor Zhong, and Richard Socher. 2016 · 2016
Cited alongside, same era.
Rationale-augmented convolutional neural networks for text classification
Ye Zhang, Iain Marshall, and Byron C Wallace. 2016 · 2016
Cited alongside, same era.
A causal framework for explaining the predictions of black-box sequence-to-sequence models
David Alvarez-Melis and Tommi Jaakkola. 2017 · 2017
Cited alongside, same era.
Qizhe Xie, Xuezhe Ma, Zihang Dai, and Eduard Hovy. 2017 · 2017
Later among the works it cites.
Pathologies of neural models make interpretation difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Pedro Rodriguez, Mohit Iyyer, and Jordan Boyd-Graber. 2018 · 2018
Later among the works it cites.
Interpreting recurrent and attention-based neural models: a case study on natural language inference
Reza Ghaeini, Xiaoli Fern, and Prasad Tadepalli. 2018 · 2018
Later among the works it cites.
Explainable prediction of medical codes from clinical text
James Mullenbach, Sarah Wiegreffe, Jon Duke, Jimeng Sun, and Jacob Eisenstein. 2018 · 2018
Later among the works it cites.
Interpretable structure induction via sparse attention
Ben Peters, Vlad Niculae, and André\par FT Martins. 2018 · 2018
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.