Fetching the paper…
Reading the bibliography…
Deep learning is usually described as an experiment-driven field under continuous criticizes of lacking theoretical foundations.
A logical calculus of the ideas immanent in nervous activity
W. S. McCulloch and W. Pitts · 1943
Earlier work this paper cites.
The Organization of Behavior: A Neuropsychological Theory
D. O. Hebb · 1949
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
F. Rosenblatt · 1958
Earlier work this paper cites.
The sizes of compact subsets of hilbert space and continuity of Gaussian processes
R. M. Dudley · 1967
Earlier work this paper cites.
Monte Carlo sampling methods using Markov chains and their applications
W. K. Hastings · 1970
Earlier work this paper cites.
Artificial intelligence methods and systems for medical consultation
C. A. Kulikowski · 1980
Earlier work this paper cites.
Stochastic relaxation, Gibbs distributions, and the Bayesian restoration of images
S. Geman and D. Geman · 1984
Earlier work this paper cites.
Learning while searching in constraint-satisfaction problems
R. Dechter · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Hybrid Monte Carlo
S. Duane, A. D. Kennedy, B. J. Pendleton, and D. Roweth · 1987
Earlier work this paper cites.
What size net gives valid generalization?
E. B. Baum and D. Haussler · 1989
Earlier work this paper cites.
Learnability and the Vapnik-Chervonenkis dimension
A. Blumer, A. Ehrenfeucht, D. Haussler, and M. K. Warmuth · 1989
Earlier work this paper cites.
On the generalization ability of neural network classifiers
M. T. Musavi, K. H. Chan, D. M. Hummels, and K. Kalantri · 1994
Earlier work this paper cites.
Bounding the Vapnik-Chervonenkis dimension of concept classes parameterized by real numbers
P. W. Goldberg and M. R. Jerrum · 1995
Earlier work this paper cites.
Sphere packing numbers for subsets of the boolean n n -cube with bounded Vapnik-Chervonenkis dimension
D. Haussler · 1995
Earlier work this paper cites.
Bayesian LEARNING FOR NEURAL NETWORKS
R. M. Neal · 1995
Earlier work this paper cites.
The VC dimension and pseudodimension of two-layer neural networks with discrete inputs
P. L. Bartlett and R. C. Williamson · 1996
Earlier work this paper cites.
Priors for infinite networks
R. M. Neal · 1996
Earlier work this paper cites.
Computing with infinite networks
C. Williams · 1996
Earlier work this paper cites.
Bounds for the computational power and learning complexity of analog neural nets
W. Maass · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
R. E. Schapire, Y. Freund, P. Bartlett, and W. S. Lee · 1998
Earlier work this paper cites.
Generalization performance of support vector machines and other pattern classifiers
P. Bartlett and J. Shawe-Taylor · 1999
Earlier work this paper cites.
Almost linear VC dimension bounds for piecewise polynomial networks
P. L. Bartlett, V. Maiorov, and R. Meir · 1999
Earlier work this paper cites.
PAC-Bayesian model averaging
D. A. McAllester · 1999
Earlier work this paper cites.
Some PAC-Bayesian theorems
D. A. McAllester · 1999
Earlier work this paper cites.
Multi-Valued and Universal Binary Neurons: Theory, Learning and Applications
I. Aizenberg, N. N. Aizenberg, and J. P. Vandewalle · 2000
Earlier work this paper cites.
Rademacher processes and bounding the risk of function learning
V. Koltchinskii and D. Panchenko · 2000
Earlier work this paper cites.
Rademacher penalties and structural risk minimization
V. Koltchinskii · 2001
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
V. Koltchinskii and D. Panchenko · 2002
Earlier work this paper cites.
Max-margin Markov networks
B. Taskar, C. Guestrin, and D. Koller · 2004
Earlier work this paper cites.
Local rademacher complexities
P. L. Bartlett, O. Bousquet, and S. Mendelson · 2005
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle · 2006
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
G. E. Hinton, S. Osindero, and Y.-W. Teh · 2006
Earlier work this paper cites.
Efficient learning of sparse representations with an energy-based model
M. Ranzato, C. Poultney, S. Chopra, and Y. Cun · 2006
Earlier work this paper cites.
On-road vehicle detection: A review
Z. Sun, G. Bebis, and R. Miller · 2006
Earlier work this paper cites.
Estimation of Dependences based on Empirical Data
V. Vapnik · 2006
Earlier work this paper cites.
Hierarchical Gaussian process latent variable models
N. D. Lawrence and A. J. Moore · 2007
Earlier work this paper cites.
Online gradient descent learning algorithms
Y. Ying and M. Pontil · 2008
Earlier work this paper cites.
Support vector machines for predicting the admission decision of a candidate to the school of physical education and sports at cukurova university
M. Acikkar and M. F. Akay · 2009
Earlier work this paper cites.
Neural Network Learning: Theoretical Foundations
M. Anthony and P. L. Bartlett · 2009
Earlier work this paper cites.
Evaluating the predictive validity of the compas risk and needs assessment system
T. Brennan, W. Dieterich, and B. Ehret · 2009
Earlier work this paper cites.
Building classifiers with independency constraints
T. Calders, F. Kamiran, and M. Pechenizkiy · 2009
Earlier work this paper cites.
Consumer credit-risk models via machine-learning algorithms
A. E. Khandani, A. J. Kim, and A. W. Lo · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
M. Welling and Y. W. Teh · 2011
Earlier work this paper cites.
Sparse algorithms are not stable: A no-free-lunch theorem
H. Xu, C. Caramanis, and S. Mannor · 2011
Earlier work this paper cites.
Bayesian posterior sampling via stochastic gradient Fisher scoring
S. Ahn, A. Korattikara, and M. Welling · 2012
Earlier work this paper cites.
Fairness through awareness
C. Dwork, M. Hardt, T. Pitassi, O. Reingold, and R. Zemel · 2012
Earlier work this paper cites.
Application of machine learning algorithms to an online recruitment system
E. Faliagka, K. Ramantas, A. Tsakalidis, and G. Tzimas · 2012
Earlier work this paper cites.
Data preprocessing techniques for classification without discrimination
F. Kamiran and T. Calders · 2012
Earlier work this paper cites.
Fairness-aware classifier with prejudice remover regularizer
T. Kamishima, S. Akaho, H. Asoh, and J. Sakuma · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
A neural autoregressive topic model
H. Larochelle and S. Lauly · 2012
Earlier work this paper cites.
Evasion attacks against machine learning at test time
B. Biggio, I. Corona, D. Maiorca, B. Nelson, N. Šrndić, P. Laskov, G. Giacinto, and F. Roli · 2013
Earlier work this paper cites.
Deep Gaussian processes
A. Damianou and N. D. Lawrence · 2013
Earlier work this paper cites.
It’s not privacy, and it’s not fair
C. Dwork and D. K. Mulligan · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
A. Graves, A.-r. Mohamed, and G. Hinton · 2013
Earlier work this paper cites.
Stochastic variational inference
M. D. Hoffman, D. M. Blei, C. Wang, and J. Paisley · 2013
Earlier work this paper cites.
Tighter pac-bayes bounds through distribution-dependent priors
G. Lever, F. Laviolette, and J. Shawe-Taylor · 2013
Earlier work this paper cites.
Stochastic gradient Riemannian Langevin dynamics on the probability simplex
S. Patterson and Y. W. Teh · 2013
Earlier work this paper cites.
The Nature of Statistical Learning Theory
V. Vapnik · 2013
Earlier work this paper cites.
Learning fair representations
R. Zemel, Y. Wu, K. Swersky, T. Pitassi, and C. Dwork · 2013
Earlier work this paper cites.
Learning polynomials with neural networks
A. Andoni, R. Panigrahy, G. Valiant, and L. Zhang · 2014
Earlier work this paper cites.
Training neural networks to predict student academic performance: A comparison of cuckoo search and gravitational search algorithms
J.-F. Chen and Q. H. Do · 2014
Earlier work this paper cites.
Stochastic gradient Hamiltonian Monte Carlo
T. Chen, E. Fox, and C. Guestrin · 2014
Earlier work this paper cites.
Deep learning: methods and applications
L. Deng and D. Yu · 2014
Earlier work this paper cites.
Bayesian sampling using stochastic gradient thermostats
N. Ding, Y. Fang, R. Babbush, C. Chen, R. D. Skeel, and H. Neven · 2014
Earlier work this paper cites.
Avoiding pathologies in very deep networks
D. Duvenaud, O. Rippel, R. Adams, and Z. Ghahramani · 2014
Earlier work this paper cites.
The algorithmic foundations of differential privacy
C. Dwork and A. Roth · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2014
Earlier work this paper cites.
Nested variational compression in deep Gaussian processes
J. Hensman and N. D. Lawrence · 2014
Earlier work this paper cites.
Auto-encoding variational Bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Earlier work this paper cites.
Preserving statistical validity in adaptive data analysis
C. Dwork, V. Feldman, M. Hardt, T. Pitassi, O. Reingold, and A. L. Roth · 2015
Earlier work this paper cites.
Certifying and removing disparate impact
M. Feldman, S. A. Friedler, J. Moeller, C. Scheidegger, and S. Venkatasubramanian · 2015
Earlier work this paper cites.
Steps toward deep kernel methods from infinite neural networks
T. Hazan and T. Jaakkola · 2015
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Earlier work this paper cites.
The variational fair autoencoder
C. Louizos, K. Swersky, Y. Li, M. Welling, and R. Zemel · 2015
Earlier work this paper cites.
A complete recipe for stochastic gradient mcmc
Y.-A. Ma, T. Chen, and E. Fox · 2015
Cited alongside, same era.
Norm-based capacity control in neural networks
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Cited alongside, same era.
Deep neural networks are easily fooled: High confidence predictions for unrecognizable images
A. Nguyen, J. Yosinski, and J. Clune · 2015
Cited alongside, same era.
On the generalization properties of differential privacy
K. Nissim and U. Stemmer · 2015
Cited alongside, same era.
Deep learning with differential privacy
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang · 2016
Cited alongside, same era.
Deep Gaussian processes for regression using approximate expectation propagation
T. Bui, D. Hernández-Lobato, J. Hernandez-Lobato, Y. Li, and R. Turner · 2016
The landscape of empirical risk for nonconvex losses
S. Mei, Y. Bai, and A. Montanari · 2018
Later among the works it cites.
The cost of fairness in binary classification
A. K. Menon and R. C. Williamson · 2018
Later among the works it cites.
Regularizing and optimizing lstm language models
S. Merity, N. S. Keskar, and R. Socher · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida · 2018
Later among the works it cites.
Foundations of machine learning
M. Mohri, A. Rostamizadeh, and A. Talwalkar · 2018
Later among the works it cites.
Generalization bounds of sgld for non-convex learning: Two theoretical viewpoints
W. Mou, L. Wang, X. Zhai, and K. Zheng · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Concentrated differential privacy: Simplifications, extensions, and lower bounds
M. Bun and T. Steinke · 2016
Cited alongside, same era.
An analysis of deep neural network models for practical applications
A. Canziani, A. Paszke, and E. Culurciello · 2016
Cited alongside, same era.
Statistical inference for model parameters in stochastic gradient descent
X. Chen, J. D. Lee, X. T. Tong, and Y. Zhang · 2016
Cited alongside, same era.
Group equivariant convolutional networks
T. Cohen and M. Welling · 2016
Cited alongside, same era.
Differential privacy as a mutual information constraint
P. Cuff and L. Yu · 2016
Cited alongside, same era.
Nonparametric stochastic approximation with large step-sizes
A. Dieuleveut and F. Bach · 2016
Cited alongside, same era.
V. Papyan · 2018
Later among the works it cites.
Generalization error bounds for noisy, iterative algorithms
A. Pensia, V. Jog, and P.-L. Loh · 2018
Later among the works it cites.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Later among the works it cites.
Certified defenses against adversarial examples
A. Raghunathan, J. Steinhardt, and P. Liang · 2018
Later among the works it cites.
Bayesian neural networks with weight sharing using Dirichlet processes
W. Roth and F. Pernkopf · 2018
Later among the works it cites.
Spurious local minima are common in two-layer ReLU neural networks
I. Safran and O. Shamir · 2018
Later among the works it cites.
Empirical analysis of the hessian of over-parametrized neural networks
L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bottou · 2018
Later among the works it cites.
Adversarially robust generalization requires more data
L. Schmidt, S. Santurkar, D. Tsipras, K. Talwar, and A. Madry · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, C. Ying, and Q. V. Le · 2018
Later among the works it cites.
A Bayesian perspective on generalization and stochastic gradient descent
S. L. Smith and Q. V. Le · 2018
Later among the works it cites.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
D. Soudry and E. Hoffer · 2018
Later among the works it cites.
Local optimality and generalization guarantees for the Langevin algorithm via empirical metastability
B. Tzen, T. Liang, and M. Raginsky · 2018
Later among the works it cites.
Qanet: Combining local convolution with global self-attention for reading comprehension
A. W. Yu, D. Dohan, M.-T. Luong, R. Zhao, K. Chen, M. Norouzi, and Q. V. Le · 2018
Later among the works it cites.
Advances in variational inference
C. Zhang, J. Bütepage, H. Kjellström, and S. Mandt · 2018
Later among the works it cites.
An information-theoretic view for deep learning
J. Zhang, T. Liu, and D. Tao · 2018
Later among the works it cites.
Empirical risk landscape analysis for understanding deep neural networks
P. Zhou and J. Feng · 2018
Later among the works it cites.
Understanding generalization and optimization performance of deep cnns
P. Zhou and J. Feng · 2018
Later among the works it cites.
Critical points of neural networks: Analytical forms and landscape properties
Y. Zhou and Y. Liang · 2018
Later among the works it cites.
Generalization error bounds with probabilistic guarantee for SGD in nonconvex optimization
Y. Zhou, Y. Liang, and H. Zhang · 2018
Later among the works it cites.
Can SGD learn recurrent neural networks with provable generalization?
Z. Allen-Zhu and Y. Li · 2019
Later among the works it cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Z. Allen-Zhu, Y. Li, and Y. Liang · 2019
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Later among the works it cites.
A state-of-the-art survey on deep learning theory and architectures
M. Z. Alom, T. M. Taha, C. Yakopcic, S. Westberg, P. Sidike, M. S. Nasrin, M. Hasan, B. C. Van Essen, A. A. Awwal, and V. K. Asari · 2019
Later among the works it cites.
S. Arora, S. S. Du, W. Hu, Z. Li, and R. Wang · 2019
Later among the works it cites.
Improved generalization bounds for robust learning
I. Attias, A. Kontorovich, and Y. Mansour · 2019
Later among the works it cites.
Lower bounds on adversarial robustness from optimal transport
A. N. Bhagoji, D. Cullina, and P. Mittal · 2019
Later among the works it cites.
Why do larger models generalize better? a theoretical perspective via the xor problem
A. Brutzkus and A. Globerson · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Y. Cao and Q. Gu · 2019
Later among the works it cites.
Capacity bounded differential privacy
K. Chaudhuri, J. Imola, and A. Machanavajjhala · 2019
Later among the works it cites.
Theoretical investigation of generalization bound for residual networks
H. Chen, Z. Mo, Z. Yang, and X. Wang · 2019
Later among the works it cites.
Fairness under unawareness: Assessing disparity when protected class is unobserved
J. Chen, N. Kallus, X. Mao, G. Svacha, and M. Udell · 2019
Later among the works it cites.
On generalization bounds of a family of recurrent neural networks
M. Chen, X. Li, and T. Zhao · 2019
Later among the works it cites.
How much over-parameterization is sufficient to learn deep ReLU networks?
Z. Chen, Y. Cao, D. Zou, and Q. Gu · 2019
Later among the works it cites.
A general theory of equivariant cnns on homogeneous spaces
T. S. Cohen, M. Geiger, and M. Weiler · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
S. Du, J. Lee, H. Li, L. Wang, and X. Zhai · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
S. S. Du, X. Zhai, B. Poczos, and A. Singh · 2019
Later among the works it cites.
Control batch size and learning rate to generalize well: Theoretical and empirical evidence
F. He, T. Liu, and D. Tao · 2019
Later among the works it cites.
Explaining landscape connectivity of low-cost solutions for multilayer nets
R. Kuditipudi, X. Wang, H. Lee, Y. Zhang, Z. Li, W. Hu, S. Arora, and R. Ge · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Later among the works it cites.
On generalization error bounds of noisy gradient methods for non-convex learning
J. Li, X. Luo, and M. Qiao · 2019
Later among the works it cites.
Generalization bounds for convolutional neural networks
S. Lin and J. Zhang · 2019
Later among the works it cites.
VC classes are adversarially robustly learnable, but only improperly
O. Montasser, S. Hanneke, and N. Srebro · 2019
Later among the works it cites.
Information-theoretic generalization bounds for sgld via data-dependent estimates
J. Negrea, M. Haghifam, G. K. Dziugaite, A. Khisti, and D. M. Roy · 2019
Later among the works it cites.
On connected sublevel sets in deep learning
Q. Nguyen · 2019
Later among the works it cites.
On the loss landscape of a class of deep neural networks with no bad local valleys
Q. Nguyen, M. C. Mukkamala, and M. Hein · 2019
Later among the works it cites.
Improved generalization bound of permutation invariant deep neural networks
A. Sannai and M. Imaizumi · 2019
Later among the works it cites.
Universal approximations of permutation invariant/equivariant functions by deep neural networks
A. Sannai, Y. Takai, and M. Cordonnier · 2019
Later among the works it cites.
Optimization for deep learning: Theory and algorithms
R. Sun · 2019
Later among the works it cites.
Theoretical analysis of adversarial learning: A minimax approach
Z. Tu, J. Zhang, and D. Tao · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
C. Wei, J. D. Lee, Q. Liu, and T. Ma · 2019
Later among the works it cites.
Y. Wen, K. Luk, M. Gazeau, G. Zhang, H. Chan, and J. Ba · 2019
Later among the works it cites.
Rademacher complexity for adversarially robust generalization
D. Yin, R. Kannan, and P. Bartlett · 2019
Later among the works it cites.
Small nonlinearities in activation functions create bad local minima in neural networks
C. Yun, S. Sra, and A. Jadbabaie · 2019
Later among the works it cites.
Deep neural networks with multi-branch architectures are intrinsically less non-convex
H. Zhang, J. Shao, and R. Salakhutdinov · 2019
Later among the works it cites.
Theoretically principled trade-off between robustness and accuracy
H. Zhang, Y. Yu, J. Jiao, E. P. Xing, L. E. Ghaoui, and M. I. Jordan · 2019
Later among the works it cites.
Distributionally adversarial attack
T. Zheng, C. Chen, and K. Ren · 2019
Later among the works it cites.
Tightening mutual information based bounds on generalization error
Y. Bu, S. Zou, and V. V. Veeravalli · 2020
Closest in time.
More data can expand the generalization gap between adversarially robust and standard models
L. Chen, Y. Min, M. Zhang, and A. Karbasi · 2020
Closest in time.
Efficient hyperparameter optimization by way of pac-bayes bound minimization
J. J. Cherian, A. G. Taube, R. T. McGibbon, P. Angelikopoulos, G. Blanc, M. Snarski, D. D. Richman, J. L. Klepeis, and D. E. Shaw · 2020
Closest in time.
Toward better generalization bounds with locally elastic stability
Z. Deng, H. He, and W. J. Su · 2020
Closest in time.
On the role of data in pac-bayes bounds
G. K. Dziugaite, K. Hsu, W. Gharbieh, and D. M. Roy · 2020
Closest in time.
W. E, C. Ma, S. Wojtowytsch, and L. Wu · 2020
Closest in time.
Truth or backpropaganda? An empirical investigation of deep learning theory
M. Goldblum, J. Geiping, A. Schwarzschild, M. Moeller, and T. Goldstein · 2020
Closest in time.
Deep learning for 3d point clouds: A survey
Y. Guo, H. Wang, Q. Hu, H. Liu, L. Liu, and M. Bennamoun · 2020
Closest in time.
M. Haghifam, J. Negrea, A. Khisti, D. M. Roy, and G. K. Dziugaite · 2020
Closest in time.
Why resnet works? residuals generalize
F. He, T. Liu, and D. Tao · 2020
Closest in time.
Piecewise linear activations substantially shape the loss surfaces of neural networks
F. He, B. Wang, and D. Tao · 2020
Closest in time.
Tighter generalization bounds for iterative differentially private learning algorithms
F. He, B. Wang, and D. Tao · 2020
Closest in time.
The local elasticity of neural networks
H. He and W. Su · 2020
Closest in time.
Generalization bounds via information density and conditional information density
F. Hellström and G. Durisi · 2020
Closest in time.
Normalizing flows: An introduction and review of current methods
I. Kobyzev, S. Prince, and M. Brubaker · 2020
Closest in time.
Y. Min, L. Chen, and A. Karbasi · 2020
Closest in time.
A survey of the usages of deep learning for natural language processing
D. W. Otter, J. R. Medina, and J. K. Kalita · 2020
Closest in time.
Adversarial risk via optimal transport and optimal couplings
M. S. Pydi and V. Jog · 2020
Closest in time.
Reasoning about generalization via conditional mutual information
T. Steinke and L. Zakynthinou · 2020
Closest in time.
Understanding generalization in recurrent neural networks
Z. Tu, F. He, and D. Tao · 2020
Closest in time.
The Interplay between Sampling and Optimization
C. Xiang · 2020
Closest in time.
Deep autoencoding topic model with scalable hybrid Bayesian inference
H. Zhang, B. Chen, Y. Cong, D. Guo, H. Liu, and M. Zhou · 2020
Closest in time.
Gradient descent optimizes over-parameterized deep ReLU networks
D. Zou, Y. Cao, D. Zhou, and Q. Gu · 2020
Closest in time.