Fetching the paper…
Reading the bibliography…
Attention-based models are successful when trained on large amounts of data.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003 · 2003
Earlier work this paper cites.
Using “annotator rationales” to improve machine learning for text categorization
Omar Zaidan, Jason Eisner, and Christine Piatko. 2007 · 2007
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Latent aspect rating analysis on review text data: a rating regression approach
Hongning Wang, Yue Lu, and Chengxiang Zhai. 2010 · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011 · 2011
Earlier work this paper cites.
Domain adaptation for large-scale sentiment classification: A deep learning approach
Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011 · 2011
Earlier work this paper cites.
Extensions of recurrent neural network language model
Tomáš Mikolov, Stefan Kombrink, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2011 · 2011
Earlier work this paper cites.
Marginalized denoising autoencoders for domain adaptation
Minmin Chen, Zhixiang Xu, Kilian Weinberger, and Fei Sha. 2012 · 2012
Earlier work this paper cites.
Learning attitudes and attributes from multi-aspect reviews
Julian McAuley, Jure Leskovec, and Dan Jurafsky. 2012 · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Alexander M Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2016 · 2016
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Later among the works it cites.
Reading wikipedia to answer open-domain questions
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017 · 2017
Later among the works it cites.
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loic Barrault, and Antoine Bordes. 2017 · 2017
Later among the works it cites.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville. 2017 · 2017
Later among the works it cites.
Supervised attention for sequence-to-sequence constituency parsing
Hidetaka Kamigaito, Katsuhiko Hayashi, Tsutomu Hirao, Hiroya Takamura, Manabu Okumura, and Masaaki Nagata. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Cited alongside, same era.
Neural machine translation with supervised attention
Lemao Liu, Masao Utiyama, Andrew Finch, and Eiichiro Sumita. 2016 · 2016
Cited alongside, same era.
Hierarchical attention networks for document classification
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016 · 2016
Cited alongside, same era.
Rationale-augmented convolutional neural networks for text classification
Ye Zhang, Iain Marshall, and Byron C Wallace. 2016 · 2016
Cited alongside, same era.
Bi-transferring deep neural networks for domain adaptation
Guangyou Zhou, Zhiwen Xie, Jimmy Xiangji Huang, and Tingting He. 2016 · 2016
Cited alongside, same era.
Martin Arjovsky, Soumith Chintala, and Léon Bottou. 2017 · 2017
Cited alongside, same era.
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017 · 2017
Later among the works it cites.
Exploiting argument information to improve event detection via supervised attention mechanisms
Shulin Liu, Yubo Chen, Kang Liu, and Jun Zhao. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Aspect-augmented adversarial networks for domain adaptation
Yuan Zhang, Regina Barzilay, and Tommi Jaakkola. 2017 · 2017
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Closest in time.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky. 2016 · 2030
Closest in time.