Fetching the paper…
Reading the bibliography…
We study how neural networks trained by gradient descent extrapolate, i.e., what they learn outside the support of the training distribution.
On a routing problem
Richard Bellman · 1958
Earlier work this paper cites.
Dynamic programming
Richard Bellman · 1966
Earlier work this paper cites.
The arbitrage theory of capital asset pricing
Stephen A Ross · 1976
Earlier work this paper cites.
The relationship between return and market value of common stocks
Rolf W Banz · 1981
Earlier work this paper cites.
A theory of the learnable
Leslie G Valiant · 1984
Earlier work this paper cites.
Nonlinear signal processing using neural networks: Prediction and system modelling
Alan Lapedes and Robert Farber · 1987
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
On the approximate realization of continuous mappings by neural networks
Ken-Ichi Funahashi · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Extrapolation and interpolation in neural network classifiers
Etienne Barnard and LFA Wessels · 1992
Earlier work this paper cites.
Extrapolation limitations of multilayer feedforward neural networks
Pamela J Haley and DONALD Soloway · 1992
Earlier work this paper cites.
Kolmogorov’s theorem and multilayer neural networks
Vera Kurkova · 1992
Earlier work this paper cites.
Common risk factors in the returns on stocks and bonds
Eugene F Fama and Kenneth R French · 1993
Earlier work this paper cites.
Improvement of learning in recurrent networks by substituting the sigmoid activation function
JM Sopena and R Alquezar · 1994
Earlier work this paper cites.
On the properties of periodic perceptrons
David B McCaughan · 1997
Earlier work this paper cites.
Learning bounds for domain adaptation
John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman · 2008
Earlier work this paper cites.
Domain adaptation: Learning bounds and algorithms
Yishay Mansour, Mehryar Mohri, and Afshin Rostamizadeh · 2009
Earlier work this paper cites.
The graph neural network model
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini · 2009
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan · 2010
Earlier work this paper cites.
Distributionally robust optimization and its tractable approximations
Joel Goh and Melvyn Sim · 2010
Earlier work this paper cites.
The nature of statistical learning theory
Vladimir Vapnik · 2013
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky · 2016
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Earlier work this paper cites.
Causal inference by using invariant prediction: identification and confidence intervals
Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen · 2016
Earlier work this paper cites.
Neural message passing for quantum chemistry
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl · 2017
Earlier work this paper cites.
Inferring and executing programs for visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Judy Hoffman, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Earlier work this paper cites.
Visual interaction networks: Learning a physics simulator from video
Nicholas Watters, Daniel Zoran, Theophane Weber, Peter Battaglia, Razvan Pascanu, and Andrea Tacchetti · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Understanding deep neural networks with rectified linear units
Raman Arora, Amitabh Basu, Poorya Mianjy, and Anirbit Mukherjee · 2018
Cited alongside, same era.
Relational inductive biases, deep learning, and graph networks
Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al · 2018
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Cited alongside, same era.
The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision
Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B. Tenenbaum, and Jiajun Wu · 2019
Later among the works it cites.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Later among the works it cites.
Analysing mathematical reasoning abilities of neural models
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli · 2019
Later among the works it cites.
Distributionally robust optimization and generalization in kernel methods
Matthew Staib and Stefanie Jegelka · 2019
Later among the works it cites.
Gradient dynamics of shallow univariate relu networks
Francis Williams, Matthew Trager, Daniele Panozzo, Claudio Silva, Denis Zorin, and Joan Bruna · 2019
Later among the works it cites.
Beto, bentz, becas: The surprising cross-lingual effectiveness of bert
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the power of over-parametrization in neural networks with quadratic activation
Simon S. Du and Jason D. Lee · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Gradient Descent Quantizes ReLU Network Features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Invariant models for causal transfer learning
Mateo Rojas-Carulla, Bernhard Schölkopf, Richard Turner, and Jonas Peters · 2018
Cited alongside, same era.
Measuring abstract reasoning in neural networks
Adam Santoro, Felix Hill, David Barrett, Ari Morcos, and Timothy Lillicrap · 2018
Cited alongside, same era.
Shijie Wu and Mark Dredze · 2019
Later among the works it cites.
How powerful are graph neural networks?
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka · 2019
Later among the works it cites.
Are girls neko or shōjo? cross-lingual alignment of non-isomorphic embeddings with iterative normalization
Mozhi Zhang, Keyulu Xu, Ken-ichi Kawarabayashi, Stefanie Jegelka, and Jordan Boyd-Graber · 2019
Later among the works it cites.
On learning invariant representations for domain adaptation
Han Zhao, Remi Tachet Des Combes, Kun Zhang, and Geoffrey Gordon · 2019
Later among the works it cites.
Harnessing the power of infinitely wide deep nets on small-data tasks
Sanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov, Ruosong Wang, and Dingli Yu · 2020
Closest in time.
Generalization of two-layer neural networks: An asymptotic viewpoint
Jimmy Ba, Murat Erdogdu, Taiji Suzuki, Denny Wu, and Tianzong Zhang · 2020
Closest in time.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Closest in time.
Principal neighbourhood aggregation for graph nets
Gabriele Corso, Luca Cavalleri, Dominique Beaini, Pietro Liò, and Petar Veličković · 2020
Closest in time.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Closest in time.
Pretrained transformers improve out-of-distribution robustness
Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song · 2020
Closest in time.
Strategies for pre-training graph neural networks
Weihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik, Percy Liang, Vijay Pande, and Jure Leskovec · 2020
Closest in time.
Deep learning for symbolic mathematics
Guillaume Lample and François Charton · 2020
Closest in time.
Neural arithmetic units
Andreas Madsen and Alexander Rosenberg Johansen · 2020
Closest in time.
Neural tangents: Fast and easy infinite neural networks in python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2020
Closest in time.
Distributionally robust neural networks
Shiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, and Percy Liang · 2020
Closest in time.
Neural execution of graph algorithms
Petar Velickovic, Rex Ying, Matilde Padovano, Raia Hadsell, and Charles Blundell · 2020
Closest in time.
Learning representations that support extrapolation
Taylor Webb, Zachary Dulberg, Steven Frankland, Alexander Petrov, Randall O’Reilly, and Jonathan Cohen · 2020
Closest in time.
What can neural networks reason about?
Keyulu Xu, Jingling Li, Mozhi Zhang, Simon S. Du, Ken ichi Kawarabayashi, and Stefanie Jegelka · 2020
Closest in time.
Interactive refinement of cross-lingual word embeddings
Michelle Yuan, Mozhi Zhang, Benjamin Van Durme, Leah Findlater, and Jordan Boyd-Graber · 2020
Closest in time.
Empirical or invariant risk minimization? a sample complexity perspective
Kartik Ahuja, Jun Wang, Amit Dhurandhar, Karthikeyan Shanmugam, and Kush R. Varshney · 2021
Closest in time.
Optimal rates for averaged stochastic gradient descent under neural tangent kernel regime
Atsushi Nitanda and Taiji Suzuki · 2021
Closest in time.
The risks of invariant risk minimization
Elan Rosenfeld, Pradeep Kumar Ravikumar, and Andrej Risteski · 2021
Closest in time.
Domain generalization with mixstyle
Kaiyang Zhou, Yongxin Yang, Yu Qiao, and Tao Xiang · 2021
Closest in time.