Fetching the paper…
Reading the bibliography…
Despite being the workhorse of deep learning, the backpropagation algorithm is no panacea.
Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences
P. J. Werbos · 1974
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Learning process in an asymmetric threshold network
Yann Le Cun · 1986
Earlier work this paper cites.
Competitive learning: From interactive activation to adaptive resonance
Stephen Grossberg · 1987
Earlier work this paper cites.
The recent excitement about neural networks
Francis Crick · 1989
Earlier work this paper cites.
Contrastive hebbian learning in the continuous hopfield model
Javier R Movellan · 1991
Earlier work this paper cites.
Biologically plausible error-driven learning using local activation differences: The generalized recirculation algorithm
Randall C O’Reilly · 1996
Earlier work this paper cites.
A logarithmic neural network architecture for unbounded non-linear function approximation
J Wesley Hines · 1996
Earlier work this paper cites.
A new model for learning in graph domains
Marco Gori, Gabriele Monfardini, and Franco Scarselli · 2005
Earlier work this paper cites.
Spike timing–dependent plasticity: a hebbian learning rule
Natalia Caporale and Yang Dan · 2008
Earlier work this paper cites.
The graph neural network model
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini · 2008
Earlier work this paper cites.
Collective classification in network data
Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Deep boltzmann machines
Ruslan Salakhutdinov and Geoffrey Hinton · 2009
Earlier work this paper cites.
Factorization machines
Steffen Rendle · 2010
Earlier work this paper cites.
Ad click prediction: a view from the trenches
H Brendan McMahan, Gary Holt, David Sculley, Michael Young, Dietmar Ebner, Julian Grady, Lan Nie, Todd Phillips, Eugene Davydov, Daniel Golovin, et al · 2013
Earlier work this paper cites.
How auto-encoders could provide credit assignment in deep networks via target propagation
Yoshua Bengio · 2014
Earlier work this paper cites.
Spectral networks and locally connected networks on graphs
Joan Bruna, Wojciech Zaremba, Arthur Szlam, and Yann Lecun · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Difference target propagation
Dong-Hyun Lee, Saizheng Zhang, Asja Fischer, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep convolutional networks on graph-structured data
Mikael Henaff, Joan Bruna, and Yann LeCun · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Toward an integration of deep learning and neuroscience
Adam H Marblestone, Greg Wayne, and Konrad P Kording · 2016
Earlier work this paper cites.
Random synaptic feedback weights support error backpropagation for deep learning
Timothy P Lillicrap, Daniel Cownden, Douglas B Tweed, and Colin J Akerman · 2016
Earlier work this paper cites.
Direct feedback alignment provides learning in deep neural networks
Arild Nøkland · 2016
Earlier work this paper cites.
How important is weight symmetry in backpropagation?
Qianli Liao, Joel Z Leibo, and Tomaso Poggio · 2016
Earlier work this paper cites.
Wide & deep learning for recommender systems
Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al · 2016
Earlier work this paper cites.
Convolutional neural networks on graphs with fast localized spectral filtering
Michaël Defferrard, Xavier Bresson, and Pierre Vandergheynst · 2016
Earlier work this paper cites.
Gated graph sequence neural networks
Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard Zemel · 2016
Earlier work this paper cites.
Variational graph auto-encoders
Thomas N Kipf and Max Welling · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Cited alongside, same era.
Decoupled neural interfaces using synthetic gradients
Max Jaderberg, Wojciech Marian Czarnecki, Simon Osindero, Oriol Vinyals, Alex Graves, David Silver, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Sobolev training for neural networks
Wojciech M Czarnecki, Simon Osindero, Max Jaderberg, Grzegorz Swirszcz, and Razvan Pascanu · 2017
Cited alongside, same era.
Soft 3d reconstruction for view synthesis
Eric Penner and Li Zhang · 2017
Cited alongside, same era.
Deepfm: a factorization-machine based neural network for ctr prediction
Deepvoxels: Learning persistent 3d feature embeddings
Vincent Sitzmann, Justus Thies, Felix Heide, Matthias Nießner, Gordon Wetzstein, and Michael Zollhofer · 2019
Later among the works it cites.
Neural volumes: Learning dynamic renderable volumes from images
Stephen Lombardi, Tomas Simon, Jason Saragih, Gabriel Schwartz, Andreas Lehrmann, and Yaser Sheikh · 2019
Later among the works it cites.
Scene representation networks: Continuous 3d-structure-aware neural scene representations
Vincent Sitzmann, Michael Zollhöfer, and Gordon Wetzstein · 2019
Later among the works it cites.
Are we really making much progress? a worrying analysis of recent neural recommendation approaches
Maurizio Ferrari Dacrema, Paolo Cremonesi, and Dietmar Jannach · 2019
Later among the works it cites.
Just jump: Dynamic neighborhood aggregation in graph neural networks
Matthias Fey · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He · 2017
Cited alongside, same era.
Deep & cross network for ad click predictions
Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang · 2017
Cited alongside, same era.
Geometric deep learning: going beyond euclidean data
Michael M Bronstein, Joan Bruna, Yann LeCun, Arthur Szlam, and Pierre Vandergheynst · 2017
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Cited alongside, same era.
Assessing the scalability of biologically-motivated deep learning algorithms and architectures
Sergey Bartunov, Adam Santoro, Blake Richards, Luke Marris, Geoffrey E Hinton, and Timothy Lillicrap · 2018
Cited alongside, same era.
Matthias Fey and Jan E. Lenssen · 2019
Later among the works it cites.
How powerful are graph neural networks?
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka · 2019
Later among the works it cites.
Gpu accelerated t-distributed stochastic neighbor embedding
David M Chan, Roshan Rao, Forrest Huang, and John F Canny · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Ali Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Later among the works it cites.
On the morality of artificial intelligence
Alexandra Luccioni and Yoshua Bengio · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Backpropagation and the brain
Timothy P Lillicrap, Adam Santoro, Luke Marris, Colin J Akerman, and Geoffrey Hinton · 2020
Closest in time.
Learning to solve the credit assignment problem
Benjamin James Lansdell, Prashanth Ravi Prakash, and Konrad Paul Kording · 2020
Closest in time.
Two routes to scalable credit assignment without weight symmetry
Daniel Kunin, Aran Nayebi, Javier Sagastuy-Brena, Surya Ganguli, Jonathan M. Bloom, and Daniel L. K. Yamins · 2020
Closest in time.
Light-in-the-loop: using a photonics co-processor for scalable training of neural networks, 2020
Julien Launay, Iacopo Poli, Kilian Müller, Igor Carron, Laurent Daudet, Florent Krzakala, and Sylvain Gigan · 2020
Closest in time.
Bottom-Up and Top-Down Neuromorphic Processor Design: Unveiling Roads to Embedded Cognition
Charlotte Frenkel · 2020
Closest in time.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Closest in time.
Implicit neural representations with periodic activation functions
Vincent Sitzmann, Julien N.P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein · 2020
Closest in time.
Adaptive factorization network: Learning adaptive-order feature interactions
Weiyu Cheng, Yanyan Shen, and Linpeng Huang · 2020
Closest in time.
Kaggle contest dataset is now available for academic use!
Criteo · 2020
Closest in time.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Closest in time.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Closest in time.
Common Crawl
The Common Crawl Team · 2020
Closest in time.
Attention in psychology, neuroscience, and machine learning
Grace W Lindsay · 2020
Closest in time.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng · 2020
Closest in time.
Fixed encoder self-attention patterns in transformer-based machine translation
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann · 2020
Closest in time.