Fetching the paper…
Reading the bibliography…
Sequence-to-sequence translation methods based on generation with a side-conditioned language model have recently shown promising results in several tasks.
“An efficient gradient-based algorithm for online training of recurrent network trajectories,”
R. Williams and J. Peng, · 1990
Earlier work this paper cites.
“Long Short-Term Memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Bidirectional recurrent neural networks,”
M. Schuster and K. Paliwal, · 1997
Earlier work this paper cites.
“Learning to forget: Continual prediction with LSTM,”
F. A. Gers, J. Schmidhuber, and F. Cummins, · 1999
Earlier work this paper cites.
“Bi-directional conversion beween graphemes and phonemes using a joint n-gram model,”
L. Galescu and J. F. Allen, · 2001
Earlier work this paper cites.
“A neural probabilistic language model,”
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin, · 2003
Earlier work this paper cites.
“Conditional and joint models for grapheme-to-phoneme conversion,”
S. Chen, · 2003
Earlier work this paper cites.
“Applying many-to-many alignments and hidden markov models to letter-to-phoneme conversion,”
S. Jiampojamarn, G. Kondrak, and T. Sherif, · 2007
Earlier work this paper cites.
“Joint-sequence models for grapheme-to-phoneme conversion,”
M. Bisani and H. Ney, · 2008
Earlier work this paper cites.
“Strategies for training large scale neural network language models,”
T. Mikolov, A. Deoras, D. Povey, L. Burget, and J. Cernocky, · 2011
Earlier work this paper cites.
“Continuous space translation models with neural networks,”
L. H. Son, A. Allauzen, and F. Yvon, · 2012
Earlier work this paper cites.
“Context dependent recurrent neural network language model,”
T. Mikolov and G. Zweig, · 2012
Cited alongside, same era.
“Statistical language models based on neural networks,”
T. Mikolov, · 2012
Cited alongside, same era.
“Joint language and translation modeling with recurrent neural networks.,”
M. Auli, M. Galley, C. Quirk, and G. Zweig, · 2013
Cited alongside, same era.
“Recurrent continuous translation nodels,”
N. Kalchbrenner and P. Blunsom, · 2013
Cited alongside, same era.
“Recurrent neural networks for language understanding,”
K. Yao, G. Zweig, M. Hwang, Y. Shi, and Dong Yu, · 2013
Cited alongside, same era.
“Generating sequences with recurrent neural networks,”
A. Graves, · 2013
Cited alongside, same era.
“Deep visual-semantic alignments for generating image descriptions,”
A. Karpathy and F.-F. Li, · 2014
Later among the works it cites.
“Show and tell: A neural image caption generator,”
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, · 2014
Later among the works it cites.
“Long-term recurrent convolutional networks for visual recognition and description,”
J. Donahue, L. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, · 2014
Later among the works it cites.
“Learning phrase representation using RNN encoder-decoder for statistical machine translation,”
K. Cho, B. v. Merrienboer, C. Gulcchre, F. Bougares, H. Schwenk, and Y. Bengjo, · 2014
Later among the works it cites.
“Neural machine translation by jointly learning to align and translate,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Investigation of recurrent-neural-network architectures and learning methods for language understanding,”
G. Mesnil, X. He, L. Deng, and Y. Bengio, · 2013
Cited alongside, same era.
“Fast and robust neural network joint models for statistical machine translation,”
J. Devlin, R. Zbib, Z. Huang, T. Lamar, R. Schwartz, and J. Makhoul, · 2014
Cited alongside, same era.
“Sequence to sequence learning with neural networks,”
H. Sutskever, O. Vinyals, and Q. V. Le, · 2014
Cited alongside, same era.
“Translation modeling with bidirectional recurrent neural networks,”
M. Sundermeyer, T. Alkhouli, J. Wuebker, and H. Ney, · 2014
Cited alongside, same era.
“From captions to visual concepts and back,”
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, L. Zitnick, and G. Zweig, · 2014
Cited alongside, same era.
“A maximum entropy approach to natural language processing,”
A. Berger, S. Della Pietra, and V. Della Pietra,
Cited in the paper.
D. Bahdanau, K. Cho, and Y. Bengio, · 2014
Later among the works it cites.
“Spoken language understanding using long short-term memory neural networks,”
K. Yao, B. Peng, Y. Zhang, D. Yu, G. Zweig, and Y. Shi, · 2014
Later among the works it cites.
H. Sak, A. Senior, and F. Beaufays, · 2014
Later among the works it cites.
“An introduction to computational networks and the computational network toolkit,”
D. Yu, A. Eversole, M. Seltzer, K. Yao, Z. Huang, B. Guenter, O. Kuchaiev, Y. Zhang, F. Seide, H. Wang, J. Droppo, G. Zweig, C. Rossbach, J. Currey, J. Gao, A. May, B. Peng, A. Stolcke, and M. Slaney, · 2014
Later among the works it cites.
“Encoding linear models as weighted finite-state transducers,”
K. Wu and · 2014
Later among the works it cites.
“Grammar as a foreign language,”
O. Vinyals, L. Kaiser, T. Koo, S. Petrov, I. Sutskever, and G. Hinton, · 2015
Closest in time.
“Grapheme-to-phoneme conversion using long short-term memory recurrent neural networks,”
K. Rao, F. Peng, H. Sak, and F. Beaufays, · 2015
Closest in time.