Fetching the paper…
Reading the bibliography…
Neural attention has become central to many state-of-the-art models in natural language processing and related domains.
Simple Statistical Gradient-following Algorithms for Connectionist Reinforcement Learning
Ronald J. Williams · 1992
Earlier work this paper cites.
The Mathematics of Statistical Machine Translation: Parameter Estimation
Peter F Brown, Vincent J Della Pietra, Stephen A Della Pietra, and Robert L Mercer · 1993
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F. Brown, Vincent J. Della Pietra, Stephen A. Della Pietra, and Robert L. Mercer · 1993
Earlier work this paper cites.
HMM-based Word Alignment in Statistical Translation
Stephan Vogel, Hermann Ney, and Christoph Tillmann · 1996
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, et al · 2007
Earlier work this paper cites.
A Simple, Fast, and Effective Reparameterization of IBM Model 2
Chris Dyer, Victor Chahuneau, and Noah A. Smith · 2013
Earlier work this paper cites.
Stochastic variational inference
Matthew D Hoffman, David M Blei, Chong Wang, and John Paisley · 2013
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stuker, Luisa Bentivogli, and Marcello Federico · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Neural Variational Inference and Learning in Belief Networks
Andriy Mnih and Karol Gregor · 2014
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Black Box Variational Inference
Rajesh Ranganath, Sean Gerrish, and David M. Blei · 2014
Earlier work this paper cites.
Stochastic Backpropagation and Approximate Inference in Deep Generative Models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Earlier work this paper cites.
Multiple Object Recognition with Visual Attention
Jimmy Ba, Volodymyr Mnih, and Koray Kavukcuoglu · 2015
Earlier work this paper cites.
Learning Wake-Sleep Recurrent Attention Models
Jimmy Ba, Ruslan R Salakhutdinov, Roger B Grosse, and Brendan J Frey · 2015
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Importance Weighted Autoencoders
Yuri Burda, Roger Grosse, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals · 2015
Earlier work this paper cites.
Describing Multimedia Content using Attention-based Encoder-Decoder Networks
Kyunghyun Cho, Aaron Courville, and Yoshua Bengio · 2015
Earlier work this paper cites.
Attention-Based Models for Speech Recognition
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
A Recurrent Latent Variable Model for Sequential Data
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron Courville, and Yoshua Bengio · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Earlier work this paper cites.
Recurrent Models of Visual Attention
Volodymyr Mnih, Nicola Heess, Alex Graves, and Koray Kavukcuoglu · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
A Neural Attention Model for Abstractive Sentence Summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston · 2015
Earlier work this paper cites.
End-To-End Memory Networks
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus · 2015
Earlier work this paper cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Incorporating Structural Alignment Biases into an Attentional Neural Translation Model
Trevor Cohn, Cong Duy Vu Hoang, Ekaterina Vymolova, Kaisheng Yao, Chris Dyer, and Gholamreza Haffari · 2016
Earlier work this paper cites.
Sequential Neural Models with Stochastic Layers
Marco Fraccaro, Soren Kaae Sonderby, Ulrich Paquet, and Ole Winther · 2016
Cited alongside, same era.
Incorporating Copying Mechanism in Sequence-to-Sequence Learning
Jiatao Gu, Zhengdong Lu, Hang Li, and Victor OK Li · 2016
Cited alongside, same era.
Dynamic Neural Turing Machine with Soft and Hard Addressing Schemes
Caglar Gulcehre, Sarath Chandar, Kyunghyun Cho, and Yoshua Bengio · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Cited alongside, same era.
Structured Inference Networks for Nonlinear State Space Models
Rahul G. Krishnan, Uri Shalit, and David Sontag · 2017
Later among the works it cites.
Learning Structured Text Representations
Yang Liu and Mirella Lapata · 2017
Later among the works it cites.
Dropout with Expectation-linear Regularization
Xuezhe Ma, Yingkai Gao, Zhiting Hu, Yaoliang Yu, Yuntian Deng, and Eduard Hovy · 2017
Later among the works it cites.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh · 2017
Later among the works it cites.
A Regularized Framework for Sparse and Structured Neural Attention
Vlad Niculae and Mathieu Blondel · 2017
Later among the works it cites.
Online and Linear-Time Attention by Enforcing Monotonic Alignments
Colin Raffel, Minh-Thang Luong, Peter J Liu, Ron J Weiss, and Douglas Eck · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tao Lei, Regina Barzilay, and Tommi Jaakkola · 2016
Cited alongside, same era.
From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label Classification
André F. T. Martins and Ramón Fernandez Astudillo · 2016
Cited alongside, same era.
Variational Inference for Monte Carlo Objectives
Andriy Mnih and Danilo J. Rezende · 2016
Cited alongside, same era.
Variational inference for monte carlo objectives
Andriy Mnih and Danilo J Rezende · 2016
Cited alongside, same era.
Iterative Refinement for Machine Translation
Roman Novak, Michael Auli, and David Grangier · 2016
Cited alongside, same era.
Reasoning about Entailment with Neural Attention
Tim Rocktäschel, Edward Grefenstette, Karl Moritz Hermann, Tomas Kocisky, and Phil Blunsom · 2016
Cited alongside, same era.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Later among the works it cites.
A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues
Iulian Vlad Serban, Alessandro Sordoni, Laurent Charlin Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2017
Later among the works it cites.
Classification of Radiology Reports Using Neural Attention Models
Bonggun Shin, Falgun H Chokshi, Timothy Lee, and Jinho D Choi · 2017
Later among the works it cites.
Autoencoding Variational Inference for Topic Models
Akash Srivastava and Charles Sutton · 2017
Later among the works it cites.
REBAR: Low-variance, Unbiased Gradient Estimates for Discrete Latent Variable Models
George Tucker, Andriy Mnih, Chris J. Maddison, Dieterich Lawson, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
The Neural Noisy Channel
Lei Yu, Phil Blunsom, Chris Dyer, Edward Grefenstette, and Tomas Kocisky · 2017
Later among the works it cites.
Structured Attentions for Visual Question Answering
Chen Zhu, Yanpeng Zhao, Shuaiyi Huang, Kewei Tu, and Yi Ma · 2017
Later among the works it cites.
Bottom-up and Top-Down Attention for Image Captioning and Visual Question Answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Closest in time.
Classical Structured Prediction Losses for Sequence to Sequence Learning
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato · 2018
Closest in time.
Backpropagation through the Void: Optimizing control variates for black-box gradient estimation
Will Grathwohl, Dami Choi, Yuhuai Wu, Geoffrey Roeder, and David Duvenaud · 2018
Closest in time.
Towards neural phrase-based machine translation
Po-Sen Huang, Chong Wang, Sitao Huang, Dengyong Zhou, and Li Deng · 2018
Closest in time.
Pathwise Derivatives Beyond the Reparameterization Trick
Martin Jankowiak and Fritz Obermeyer · 2018
Closest in time.
Semi-amortized variational autoencoders
Yoon Kim, Sam Wiseman, Andrew C Miller, David Sontag, and Alexander M Rush · 2018
Closest in time.
On the Challenges of Learning with Inference Networks on Sparse, High-dimensional Data
Rahul G. Krishnan, Dawen Liang, and Matthew Hoffman · 2018
Closest in time.
Learning Hard Alignments in Variational Inference
Dieterich Lawson, Chung-Cheng Chiu, George Tucker, Colin Raffel, Kevin Swersky, and Navdeep Jaitly · 2018
Closest in time.
Deterministic Non-Autoregressive Neural Sequence Modeling by Iterative Refinement
Jason Lee, Elman Mansimov, and Kyunghyun Cho · 2018
Closest in time.
Differentiable Dynamic Programming for Structured Prediction and Attention
Arthur Mensch and Mathieu Blondel · 2018
Closest in time.
SparseMAP: Differentiable Sparse Structured Inference
Vlad Niculae, André F. T. Martins, Mathieu Blondel, and Claire Cardie · 2018
Closest in time.
A Stochastic Decoder for Neural Machine Translation
Philip Schulz, Wilker Aziz, and Trevor Cohn · 2018
Closest in time.
Surprisingly Easy Hard-Attention for Sequence to Sequence Learning
Shiv Shankar, Siddhant Garg, and Sunita Sarawagi · 2018
Closest in time.
Variational Recurrent Neural Machine Translation
Jinsong Su, Shan Wu, Deyi Xiong, Yaojie Lu, Xianpei Han, and Biao Zhang · 2018
Closest in time.
Hard Non-Monotonic Attention for Character-Level Transduction
Shijie Wu, Pamela Shapiro, and Ryan Cotterell · 2018
Closest in time.