Fetching the paper…
Reading the bibliography…
Deep equilibrium networks (DEQs) are a new class of models that eschews traditional depth in favor of finding the fixed point of a single nonlinear layer.
Praktische verfahren der gleichungsauflösung
von Mises, R. and Pollaczek-Geiringer, H · 1929
Earlier work this paper cites.
The numerical solution of parabolic and elliptic differential equations
Peaceman, D. W. and Rachford, Jr, H. H · 1955
Earlier work this paper cites.
Iterative procedures for nonlinear integral equations
Anderson, D. G · 1965
Earlier work this paper cites.
A class of methods for solving nonlinear simultaneous equations
Broyden, C. G · 1965
Earlier work this paper cites.
A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines
Hutchinson, M. F · 1989
Earlier work this paper cites.
Improving generalization performance using double backpropagation
Drucker, H. and Le Cun, Y · 1992
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn treebank
Marcus, M. P., Marcinkiewicz, M. A., and Santorini, B · 1993
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L., Li, K., and Li, F · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Randomized algorithms for estimating the trace of an implicit symmetric positive semi-definite matrix
Avron, H. and Toledo, S · 2011
Earlier work this paper cites.
The implicit function theorem: History, theory, and applications
Krantz, S. G. and Parks, H. R · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Improved bounds on sample size for implicit matrix trace estimators
Roosta-Khorasani, F. and Ascher, U · 2015
Earlier work this paper cites.
Ba, L. J., Kiros, R., and Hinton, G. E · 2016
Earlier work this paper cites.
The Cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B · 2016
Earlier work this paper cites.
A theoretically grounded application of dropout in recurrent neural networks
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Earlier work this paper cites.
OptNet: Differentiable optimization as a layer in neural networks
Amos, B. and Kolter, J. Z · 2017
Earlier work this paper cites.
Quasi-recurrent neural networks
Bradbury, J., Merity, S., Xiong, C., and Socher, R · 2017
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Cited alongside, same era.
Robust large margin deep neural networks
Sokolić, J., Giryes, R., Sapiro, G., and Rodrigues, M. R · 2017
Cited alongside, same era.
Fast estimation of tr(f(a)) via stochastic lanczos quadrature
Ubaru, S., Chen, J., and Saad, Y · 2017
Cited alongside, same era.
FFJORD: Free-form continuous dynamics for scalable reversible generative models
Grathwohl, W., Chen, R. T., Betterncourt, J., Sutskever, I., and Duvenaud, D · 2019
Later among the works it cites.
Robust learning with Jacobian regularization
Hoffman, J., Roberts, D. A., and Yaida, S · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Later among the works it cites.
Multiscale deep equilibrium models
Bai, S., Koltun, V., and Kolter, J. Z · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Multi-level residual networks from dynamical systems view
Chang, B., Meng, L., Haber, E., Tung, F., and Begert, D · 2018
Cited alongside, same era.
Neural ordinary differential equations
Chen, T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K · 2018
Cited alongside, same era.
Regularizing and optimizing LSTM language models
Merity, S., Keskar, N. S., and Socher, R · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Cited alongside, same era.
Sensitivity and generalization in neural networks: An empirical study
Novak, R., Bahri, Y., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J · 2018
Cited alongside, same era.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Deep implicit layers tutorial - neural ODEs, deep equilibirum models, and beyond
Duvenaud, D., Kolter, J. Z., and Johnson, M · 2020
Later among the works it cites.
Finlay, C., Jacobsen, J.-H., Nurbekyan, L., and Oberman, A. M · 2020
Later among the works it cites.
Learning differential equations that are easy to solve
Kelly, J., Bettencourt, J., Johnson, M. J., and Duvenaud, D · 2020
Later among the works it cites.
Stable and expressive recurrent vision models
Linsley, D., Ashok, A. K., Govindarajan, L. N., Liu, R., and Serre, T · 2020
Later among the works it cites.
Understanding the difficulty of training transformers
Liu, L., Liu, X., Gao, J., Chen, W., and Han, J · 2020
Later among the works it cites.
Hypersolvers: Toward fast continuous-depth models
Poli, M., Massaroli, S., Yamashita, A., Asama, H., and Park, J · 2020
Later among the works it cites.
Lipschitz bounded equilibrium networks
Revay, M., Wang, R., and Manchester, I. R · 2020
Later among the works it cites.
Monotone operator equilibrium networks
Winston, E. and Kolter, J. Z · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Later among the works it cites.
Dynamics of deep equilibrium linear models
Kawaguchi, K · 2021
Closest in time.
Implicit normalizing flows
Lu, C., Chen, J., Li, C., Wang, Q., and Zhu, J · 2021
Closest in time.
Hutch++: Optimal stochastic trace estimation
Meyer, R. A., Musco, C., Musco, C., and Woodruff, D. P · 2021
Closest in time.
Estimating Lipschitz constants of monotone deep equilibrium models
Pabbaraju, C., Winston, E., and Kolter, J. Z · 2021
Closest in time.