Fetching the paper…
Reading the bibliography…
Neural networks trained with backpropagation often struggle to identify classes that have been observed a small number of times.
The psychology of language
Zipf, George K · 1935
Earlier work this paper cites.
The organization of behavior: A neurophysiological approach, 1949
Hebb, Donald O · 1949
Earlier work this paper cites.
Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters
Bridle, John S · 1990
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
Kneser, Reinhard and Ney, Hermann · 1995
Earlier work this paper cites.
Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory
McClelland, James L, McNaughton, Bruce L, and O’reilly, Randall C · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Lifelong learning algorithms
Thrun, Sebastian · 1998
Earlier work this paper cites.
Classes for fast maximum entropy training
Goodman, Joshua · 2001
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, Sepp, Younger, A Steven, and Conwell, Peter R · 2001
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, Ronan and Weston, Jason · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, Xavier and Bengio, Yoshua · 2010
Earlier work this paper cites.
English gigaword fifth edition ldc2011t07. dvd
Parker, Robert, Graff, David, Kong, Junbo, Chen, Ke, and Maeda, Kazuaki · 2011
Earlier work this paper cites.
Lstm neural networks for language modeling
Sundermeyer, Martin, Schlüter, Ralf, and Ney, Hermann · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
Graves, Alex, Wayne, Greg, and Danihelka, Ivo · 2014
Cited alongside, same era.
A convolutional neural network for modelling sentences
Kalchbrenner, Nal, Grefenstette, Edward, and Blunsom, Phil · 2014
Cited alongside, same era.
Strategies for training large vocabulary neural language models
Chen, Welin, Grangier, David, and Auli, Michael · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
Lake, Brenden M, Salakhutdinov, Ruslan, and Tenenbaum, Joshua B · 2015
Scaling memory-augmented neural networks with sparse reads and writes
Rae, Jack, Hunt, Jonathan J, Danihelka, Ivo, Harley, Timothy, Senior, Andrew W, Wayne, Gregory, Graves, Alex, and Lillicrap, Tim · 2016
Later among the works it cites.
One-shot learning with memory-augmented neural networks
Santoro, Adam, Bartunov, Sergey, Botvinick, Matthew, Wierstra, Daan, and Lillicrap, Timothy · 2016
Later among the works it cites.
Matching networks for one shot learning
Vinyals, Oriol, Blundell, Charles, Lillicrap, Tim, Wierstra, Daan, et al · 2016
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, Chelsea, Abbeel, Pieter, and Levine, Sergey · 2017
Later among the works it cites.
Unbounded cache model for online language modeling with open vocabulary
Grave, Edouard, Cisse, Moustapha M, and Joulin, Armand · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end memory networks
Sukhbaatar, Sainbayar, Weston, Jason, Fergus, Rob, et al · 2015
Cited alongside, same era.
Pointer networks
Vinyals, Oriol, Fortunato, Meire, and Jaitly, Navdeep · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, Marcin, Denil, Misha, Gomez, Sergio, Hoffman, Matthew W, Pfau, David, Schaul, Tom, and de Freitas, Nando · 2016
Cited alongside, same era.
Language modeling with gated convolutional networks
Dauphin, Yann N, Fan, Angela, Auli, Michael, and Grangier, David · 2016
Cited alongside, same era.
Hybrid computing using a neural network with dynamic external memory
Graves, Alex, Wayne, Greg, Reynolds, Malcolm, Harley, Tim, Danihelka, Ivo, Grabska-Barwińska, Agnieszka, Colmenarejo, Sergio Gómez, Grefenstette, Edward, Ramalho, Tiago, Agapiou, John, et al · 2016
Cited alongside, same era.
Gulcehre, Caglar, Ahn, Sungjin, Nallapati, Ramesh, Zhou, Bowen, and Bengio, Yoshua · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Jozefowicz, Rafal, Vinyals, Oriol, Schuster, Mike, Shazeer, Noam, and Wu, Yonghui · 2016
Cited alongside, same era.
Search engine guided non-parametric neural machine translation
Gu, Jiatao, Wang, Yong, Cho, Kyunghyun, and Li, Victor OK · 2017
Later among the works it cites.
Learning to remember rare events
Kaiser, Lukasz, Nachum, Ofir, Roy, Aurko, and Bengio, Samy · 2017
Later among the works it cites.
Learning to create and reuse words in open-vocabulary neural language modeling
Kawakami, Kazuya, Dyer, Chris, and Blunsom, Phil · 2017
Later among the works it cites.
On the state of the art of evaluation in neural language models
Melis, Gábor, Dyer, Chris, and Blunsom, Phil · 2017
Later among the works it cites.
Breaking the softmax bottleneck: a high-rank rnn language model
Yang, Zhilin, Dai, Zihang, Salakhutdinov, Ruslan, and Cohen, William W · 2017
Later among the works it cites.
Convolutional sequence modeling revisited, 2018
Bai, Shaojie, Kolter, J. Zico, and Koltun, Vladlen · 2018
Closest in time.
Memory-based parameter adaptation
Sprechmann, Pablo, Jayakumar, Siddhant, Rae, W. Jack, Pritzel, Alexander, Puigdomenech, Adria Badia, Uria, Benigno, Vinyals, Oriol, Hassabis, Demis, Pascanu, Razvan, and Blundell, Charles · 2018
Closest in time.
Deep meta-learning: Learning to learn in the concept space
Zhou, Fengwei, Wu, Bin, and Li, Zhenguo · 2018
Closest in time.