Fetching the paper…
Reading the bibliography…
Deep neural networks have excelled on a wide range of problems, from vision to language and game playing.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, Michael and Cohen, Neal J · 1989
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Marcus, Mitchell P, Marcinkiewicz, Mary Ann, and Santorini, Beatrice · 1993
Earlier work this paper cites.
Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory
McClelland, James L, McNaughton, Bruce L, and O’reilly, Randall C · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, Robert M · 1999
Earlier work this paper cites.
Local regression and likelihood
Loader, Clive · 2006
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Earlier work this paper cites.
Goodfellow, Ian J, Warde-Farley, David, Mirza, Mehdi, Courville, Aaron, and Bengio, Yoshua · 2013
Earlier work this paper cites.
Unsupervised learning of invariant representations with low sample complexity: the magic of sensory cortex or a new framework for machine learning?
Anselmi, Fabio, Leibo, Joel Z, Rosasco, Lorenzo, Mutch, Jim, Tacchetti, Andrea, and Poggio, Tomaso · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Earlier work this paper cites.
Weston, Jason, Chopra, Sumit, and Bordes, Antoine · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, Geoffrey, Vinyals, Oriol, and Dean, Jeff · 2015
Earlier work this paper cites.
Approximate hubel-wiesel modules and the data structures of neural computation
Leibo, Joel Z, Cornebise, Julien, Gómez, Sergio, and Hassabis, Demis · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, Olga, Deng, Jia, Su, Hao, Krause, Jonathan, Satheesh, Sanjeev, Ma, Sean, Huang, Zhiheng, Karpathy, Andrej, Khosla, Aditya, Bernstein, Michael, et al · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Zhang, Xiang, Zhao, Junbo, and LeCun, Yann · 2015
Cited alongside, same era.
Using fast weights to attend to the recent past
Ba, Jimmy, Hinton, Geoffrey E, Mnih, Volodymyr, Leibo, Joel Z, and Ionescu, Catalin · 2016
Cited alongside, same era.
Blundell, Charles, Uria, Benigno, Pritzel, Alexander, Li, Yazhe, Ruderman, Avraham, Leibo, Joel Z, Rae, Jack, Wierstra, Daan, and Hassabis, Demis · 2016
Cited alongside, same era.
Active long term memory networks
Furlanello, Tommaso, Zhao, Jiaping, Saxe, Andrew M, Itti, Laurent, and Tjan, Bosco S · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Yonghui, Schuster, Mike, Chen, Zhifeng, Le, Quoc V, Norouzi, Mohammad, Macherey, Wolfgang, Krikun, Maxim, Cao, Yuan, Gao, Qin, Macherey, Klaus, et al · 2016
Later among the works it cites.
Gradual learning of deep recurrent neural networks
Aharoni, Ziv, Rattner, Gal, and Permuter, Haim · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, Chelsea, Abbeel, Pieter, and Levine, Sergey · 2017
Later among the works it cites.
Bayesian recurrent neural networks
Fortunato, Meire, Blundell, Charles, and Vinyals, Oriol · 2017
Later among the works it cites.
Search engine guided non-parametric neural machine translation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving neural language models with a continuous cache
Grave, Edouard, Joulin, Armand, and Usunier, Nicolas · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2016
Cited alongside, same era.
What learning systems do intelligent agents need? complementary learning systems theory updated
Kumaran, Dharshan, Hassabis, Demis, and McClelland, James L · 2016
Cited alongside, same era.
One sentence one model for neural machine translation
Li, Xiaoqing, Zhang, Jiajun, and Zong, Chengqing · 2016
Cited alongside, same era.
Learning Without Forgetting , pp. 614–629
Li, Zhizhong and Hoiem, Derek · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Merity, Stephen, Xiong, Caiming, Bradbury, James, and Socher, Richard · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Oord, Aaron van den, Dieleman, Sander, Zen, Heiga, Simonyan, Karen, Vinyals, Oriol, Graves, Alex, Kalchbrenner, Nal, Senior, Andrew, and Kavukcuoglu, Koray · 2016
Cited alongside, same era.
Gu, Jiatao, Wang, Yong, Cho, Kyunghyun, and Li, Victor OK · 2017
Later among the works it cites.
Learning to remember rare events
Kaiser, Łukasz, Nachum, Ofir, Roy, Aurko, and Bengio, Samy · 2017
Later among the works it cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, James, Pascanu, Razvan, Rabinowitz, Neil, Veness, Joel, Desjardins, Guillaume, Rusu, Andrei A, Milan, Kieran, Quan, John, Ramalho, Tiago, Grabska-Barwinska, Agnieszka, et al · 2017
Later among the works it cites.
Dynamic evaluation of neural sequence models
Krause, Ben, Kahembwe, Emmanuel, Murray, Iain, and Renals, Steve · 2017
Later among the works it cites.
Gradient episodic memory for continuum learning
Lopez-Paz, David and Ranzato, Marc’Aurelio · 2017
Later among the works it cites.
On the state of the art of evaluation in neural language models
Melis, Gábor, Dyer, Chris, and Blunsom, Phil · 2017
Later among the works it cites.
Regularizing and optimizing lstm language models
Merity, Stephen, Keskar, Nitish Shirish, and Socher, Richard · 2017
Later among the works it cites.
Munkhdalai, Tsendsuren and Yu, Hong · 2017
Later among the works it cites.
Neural episodic control
Pritzel, Alexander, Uria, Benigno, Srinivasan, Sriram, Puigdomènech, Adrià, Vinyals, Oriol, Hassabis, Demis, Wierstra, Daan, and Blundell, Charles · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, David, Schrittwieser, Julian, Simonyan, Karen, Antonoglou, Ioannis, Huang, Aja, Guez, Arthur, Hubert, Thomas, Baker, Lucas, Lai, Matthew, Bolton, Adrian, Chen, Yutian Chen, Lillicrap, Timothy, Hui, Fan Hui, Sifre, Laurent, van den Driessche, George, Graepel, Thore, and Hassabis, Demis · 2017
Later among the works it cites.
Prototypical networks for few-shot learning
Snell, Jake, Swersky, Kevin, and Zemel, Richard S · 2017
Later among the works it cites.