Fetching the paper…
Reading the bibliography…
Recurrent Neural Networks (RNNs) with attention mechanisms have obtained state-of-the-art results for many sequence processing tasks.
Untersuchungen zu dynamischen neuronalen Netzen. Diploma thesis, Institut für Informatik, Lehrstuhl Prof. Brauer, Technische Universität München, 1991
Hochreiter, S · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P · 1994
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
El Hihi, S. and Bengio, Y · 1995
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Hochreiter, S · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
A clockwork rnn
Koutnik, J., Greff, K., Gomez, F., and Schmidhuber, J · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Le, Q. V., Jaitly, N., and Hinton, G. E · 2015
Earlier work this paper cites.
A hierarchical recurrent encoder-decoder for generative context-aware query suggestion
Sordoni, A., Bengio, Y., Vahabi, H., Lioma, C., Grue Simonsen, J., and Nie, J.-Y · 2015
Earlier work this paper cites.
Srivastava, R. K., Greff, K., and Schmidhuber, J · 2015
Earlier work this paper cites.
Depth-gated recurrent neural networks
Yao, K., Cohn, T., Vylomova, K., Duh, K., and Dyer, C · 2015
Cited alongside, same era.
Learning efficient algorithms with hierarchical attentive memory
Andrychowicz, M. and Kurach, K · 2016
Cited alongside, same era.
End-to-end attention-based large vocabulary speech recognition
Bahdanau, D., Chorowski, J., Serdyuk, D., Brakel, P., and Bengio, Y · 2016
Cited alongside, same era.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
Chan, W., Jaitly, N., Le, Q., and Vinyals, O · 2016
Cited alongside, same era.
Hierarchical multiscale recurrent neural networks
Chung, J., Ahn, S., and Bengio, Y · 2016
Cited alongside, same era.
Machine comprehension using match-lstm and answer pointer
Wang, S. and Jiang, J · 2016
Later among the works it cites.
Hierarchical attention networks for document classification
Yang, Z., Yang, D., Dyer, C., He, X., Smola, A., and Hovy, E · 2016
Later among the works it cites.
Boosting image captioning with attributes
Yao, T., Pan, Y., Li, Y., Qiu, Z., and Mei, T · 2016
Later among the works it cites.
Skip rnn: Learning to skip state updates in recurrent neural networks
Campos, V., Jou, B., Giró-i Nieto, X., Torres, J., and Chang, S.-F · 2017
Later among the works it cites.
Reading wikipedia to answer open-domain questions
Chen, D., Fisch, A., Weston, J., and Bordes, A · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cui, Y., Chen, Z., Wei, S., Wang, S., Liu, T., and Hu, G · 2016
Cited alongside, same era.
Adaptive computation time for recurrent neural networks
Graves, A · 2016
Cited alongside, same era.
Gulcehre, C., Ahn, S., Nallapati, R., Zhou, B., and Bengio, Y · 2016
Cited alongside, same era.
Text understanding with the attention sum reader network
Kadlec, R., Schmid, M., Bajgar, O., and Kleindienst, J · 2016
Cited alongside, same era.
Ask me anything: Dynamic memory networks for natural language processing
Kumar, A., Irsoy, O., Ondruska, P., Iyyer, M., Bradbury, J., Gulrajani, I., Zhong, V., Paulus, R., and Socher, R · 2016
Cited alongside, same era.
Key-value memory networks for directly reading documents
Miller, A., Fisch, A., Dodge, J., Karimi, A.-H., Bordes, A., and Weston, J · 2016
Cited alongside, same era.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Nallapati, R., Zhou, B., dos Santos, C. N., Gülçehre, Ç., and Xiang, B · 2016
Cited alongside, same era.
Later among the works it cites.
Searchqa: A new q&a dataset augmented with context from a search engine
Dunn, M., Sagun, L., Higgins, M., Guney, U., Cirik, V., and Cho, K · 2017
Later among the works it cites.
LSTM encoder-decoder architecture with attention mechanism for machine comprehension
Higgins, B. and Nho, E · 2017
Later among the works it cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Lu, J., Xiong, C., Parikh, D., and Socher, R · 2017
Later among the works it cites.
Reasonet: Learning to stop reading in machine comprehension
Shen, Y., Huang, P.-S., Gao, J., and Chen, W · 2017
Later among the works it cites.
S-net: From answer extraction to answer generation for machine reading comprehension
Tan, C., Wei, F., Yang, N., Lv, W., and Zhou, M · 2017
Later among the works it cites.
Making neural qa as simple as possible but not simpler
Weissenborn, D., Wiese, G., and Seiffe, L · 2017
Later among the works it cites.
Ask the right questions: Active question reformulation with reinforcement learning
Buck, C., Bulian, J., Ciaramita, M., Gajewski, W., Gesmundo, A., Houlsby, N., and Wang., W · 2018
Closest in time.
A question-focused multi-factor attention network for question answering
Kundu, S. and Ng, H. T · 2018
Closest in time.