Fetching the paper…
Reading the bibliography…
Attention is a powerful and ubiquitous mechanism for allowing neural models to focus on particular salient pieces of information by taking their weighted average when making predictions.
compare-mt: A tool for holistic comparison of language generation systems
Graham Neubig, Zi-Yi Dou, Junjie Hu, Paul Michel, Danish Pruthi, and Xinyi Wang · 1903
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S. Denker, and Sara A. Solla · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David G. Stork · 1993
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst · 2007
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Report on the 11 th iwslt evaluation campaign , iwslt 2014
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Earlier work this paper cites.
Auto-sizing neural networks: With applications to n-gram language models
Kenton Murray and David Chiang · 2015
Earlier work this paper cites.
Layer-wise relevance propagation for neural networks with local renormalization layers
Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, Klaus-Robert Müller, and Wojciech Samek · 2016
Cited alongside, same era.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush · 2016
Cited alongside, same era.
Pruning filters for efficient convnets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit · 2016
Cited alongside, same era.
Compression of neural machine translation models via pruning
MTNT: A testbed for machine translation of noisy text
Paul Michel and Graham Neubig · 2018
Later among the works it cites.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli · 2018
Later among the works it cites.
Improving language understanding with unsupervised learning
Alec Radford, Karthik Narasimhan, Time Salimans, and Ilya Sutskever · 2018
Later among the works it cites.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato and Jörg Tiedemann · 2018
Later among the works it cites.
Disan: Directional self-attention network for rnn/cnn-free language understanding
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, Shirui Pan, and Chengqi Zhang · 2018
Later among the works it cites.
Linguistically-informed self-attention for semantic role labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abigail See, Minh-Thang Luong, and Christopher D. Manning · 2016
Cited alongside, same era.
Weighted transformer network for machine translation
Karim Ahmed, Nitish Shirish Keskar, and Richard Socher · 2017
Cited alongside, same era.
Structured pruning of deep convolutional neural networks
Sajid Anwar, Kyuyeon Hwang, and Wonyong Sung · 2017
Cited alongside, same era.
Pruning convolutional neural networks for resource efficient inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz · 2017
Cited alongside, same era.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher · 2017
Cited alongside, same era.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Later among the works it cites.
Why self-attention? a targeted evaluation of neural machine translation architectures
Gongbo Tang, Mathias Müller, Annette Rios, and Rico Sennrich · 2018
Later among the works it cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Closest in time.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Titov Ivan · 2019
Closest in time.