Fetching the paper…
Reading the bibliography…
We demonstrate that a continuous relaxation of the argmax operation can be used to create a differentiable approximation to greedy decoding for sequence-to-sequence (seq2seq) models.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik F Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Learning as search optimization: Approximate large margin methods for structured prediction
Hal Daumé III and Daniel Marcu. 2005 · 2005
Earlier work this paper cites.
Search-based structured prediction
Hal Daumé, John Langford, and Daniel Marcu. 2009 · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey J Gordon, and Drew Bagnell. 2011 · 2011
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign, iwslt 2014
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Cited alongside, same era.
Approximation-aware dependency parsing by belief propagation
Matthew R. Gormley, Mark Dredze, and Jason Eisner. 2015 · 2015
Cited alongside, same era.
Learning to transduce with unbounded memory
Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Alexander M Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Cited alongside, same era.
Building end-to-end dialogue systems using generative hierarchical neural network models
Iulian V Serban, Alessandro Sordoni, Yoshua Bengio, Aaron Courville, and Joelle Pineau. 2015 · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2016 · 2016
Later among the works it cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016 · 2016
Later among the works it cites.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M Rush. 2016 · 2016
Later among the works it cites.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio. 2017 · 2017
Closest in time.
Decoding as continuous optimization in neural machine translation
Cong Duy Vu Hoang, Gholamreza Haffari, and Trevor Cohn. 2017 · 2017
Closest in time.
The concrete distribution: A continuous relaxation of discrete random variables
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C Courville, Ruslan Salakhutdinov, Richard S Zemel, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Globally normalized transition-based neural networks
Daniel Andor, Chris Alberti, David Weiss, Aliaksei Severyn, Alessandro Presta, Kuzman Ganchev, Slav Petrov, and Michael Collins. 2016 · 2016
Cited alongside, same era.
Chris J Maddison, Andriy Mnih, and Yee Whye Teh. 2017 · 2017
Closest in time.