Fetching the paper…
Reading the bibliography…
We introduce a novel schema for sequence to sequence learning with a Deep Q-Network (DQN), which decodes the output sequence iteratively.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
Q-learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
A. W. Moore and C. G. Atkeson · 1993
Earlier work this paper cites.
Chapter 25 - serial order: A parallel distributed processing approach
M. I. Jordan · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
C.-Y. Lin and F. J. Och · 2004
Earlier work this paper cites.
Reinforcement learning of local shape in the game of go
D. Silver, R. Sutton, and M. Müller · 2007
Earlier work this paper cites.
Reading to learn: Constructing features from semantic abstracts
J. Eisenstein, J. Clarke, D. Goldwasser, and D. Roth · 2009
Earlier work this paper cites.
High-level reinforcement learning in strategy games
C. Amato and G. Shani · 2010
Earlier work this paper cites.
Reading between the lines: Learning to map high-level instructions to commands
S. R. K. Branavan, L. S. Zettlemoyer, and R. Barzilay · 2010
Earlier work this paper cites.
Toward understanding natural language directions
T. Kollar, S. Tellex, D. Roy, and N. Roy · 2010
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Cited alongside, same era.
Learning to follow navigational directions
A. Vogel and D. Jurafsky · 2010
Cited alongside, same era.
Learning to win by reading manuals in a monte-carlo framework
S. R. K. Branavan, D. Silver, and R. Barzilay · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Lstm neural networks for language modeling
M. Sundermeyer, R. Schlüter, and H. Ney · 2012
Cited alongside, same era.
Weakly supervised learning of semantic parsers for mapping instructions to actions
Y. Artzi and L. Zettlemoyer · 2013
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, Ç. Gülçehre, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Later among the works it cites.
Recurrent models of visual attention
V. Mnih, N. Heess, A. Graves, and K. Kavukcuoglu · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Speech recognition with deep recurrent neural networks
A. Graves, A. Mohamed, and G. E. Hinton · 2013
Cited alongside, same era.
Learning to parse natural language commands to a robot control system
C. Matuszek, E. Herbst, L. Zettlemoyer, and D. Fox · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Learning to win by reading manuals in a monte-carlo framework
S. R. K. Branavan, D. Silver, and R. Barzilay · 2014
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2014
Cited alongside, same era.
K. Gregor, I. Danihelka, A. Graves, and D. Wierstra · 2015
Closest in time.
Deep recurrent q-learning for partially observable mdps
M. J. Hausknecht and P. Stone · 2015
Closest in time.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Closest in time.
Language understanding for text-based games using deep reinforcement learning
K. Narasimhan, T. Kulkarni, and R. Barzilay · 2015
Closest in time.
Action-conditional video prediction using deep networks in atari games
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. P. Singh · 2015
Closest in time.
Weakly supervised memory networks
S. Sukhbaatar, A. Szlam, J. Weston, and R. Fergus · 2015
Closest in time.
Reinforcement learning neural turing machines
W. Zaremba and I. Sutskever · 2015
Closest in time.