Fetching the paper…
Reading the bibliography…
Lipschitz constants of neural networks have been explored in various contexts in deep learning, such as provable adversarial robustness, estimating Wasserstein distance, stabilising training of GANs, and formulating invertible neural networks.
Praktische verfahren der gleichungsauflösung
Mises, R. and Pollaczek-Geiringer, H · 1929
Earlier work this paper cites.
Geometric Measure Theory
Federer, H · 1969
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Marcus, M. P., Marcinkiewicz, M. A., and Santorini, B · 1993
Earlier work this paper cites.
On the Lambert W function
Corless, R. M., Gonnet, G. H., Hare, D. E., Jeffrey, D. J., and Knuth, D. E · 1996
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous distributed systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., et al · 2016
Earlier work this paper cites.
Parseval networks: Improving robustness to adversarial examples
Cisse, M., Bojanowski, P., Grave, E., Dauphin, Y., and Usunier, N · 2017
Earlier work this paper cites.
Robust large margin deep neural networks
Sokolić, J., Giryes, R., Sapiro, G., and Rodrigues, M. R · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Training deeper neural machine translation models with transparent attention
Bapna, A., Chen, M. X., Firat, O., Cao, Y., and Wu, Y · 2018
Earlier work this paper cites.
Neural ordinary differential equations
Chen, R. T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D · 2018
Earlier work this paper cites.
Glow: Generative flow with invertible 1 × 1 1\times 1 convolutions
Kingma, D. P. and Dhariwal, P · 2018
Cited alongside, same era.
Spectral normalization for Generative Adversarial Networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Cited alongside, same era.
Lipschitz-margin training: Scalable certification of perturbation invariance for deep neural networks
Tsuzuku, Y., Sato, I., and Sugiyama, M · 2018
Cited alongside, same era.
Lipschitz regularity of deep neural networks: analysis and efficient estimation
Virmaux, A. and Scaman, K · 2018
Cited alongside, same era.
Non-local neural networks
Wang, X., Girshick, R., Gupta, A., and He, K · 2018
Cited alongside, same era.
Sorting out Lipschitz function approximation
Anil, C., Lucas, J., and Grosse, R · 2019
Cited alongside, same era.
Music Transformer
Huang, C.-Z. A., Vaswani, A., Uszkoreit, J., Simon, I., Hawthorne, C., Shazeer, N., Dai, A. M., Hoffman, M. D., Dinculescu, M., and Eck, D · 2019
Later among the works it cites.
Stand-alone self-attention in vision models
Parmar, N., Ramachandran, P., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J · 2019
Later among the works it cites.
Computational optimal transport
Peyré, G. and Cuturi, M · 2019
Later among the works it cites.
Transformer dissection: An unified understanding for Transformer’s attention via the lens of kernel
Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R · 2019
Later among the works it cites.
Learning deep transformer models for machine translation
Wang, Q., Li, B., Xiao, T., Zhu, J., Li, C., Wong, D. F., and Chao, L. S · 2019
Later among the works it cites.
Self-attention Generative Adversarial Networks
Zhang, H., Goodfellow, I., Metaxas, D., and Odena, A · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Invertible residual networks
Behrmann, J., Grathwohl, W., Chen, R. T. Q., Duvenaud, D., and Jacobsen, J.-H · 2019
Cited alongside, same era.
Residual flows for invertible generative modeling
Chen, R. T. Q., Behrmann, J., Duvenaud, D., and Jacobsen, J.-H · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Efficient and accurate estimation of Lipschitz constants for deep neural networks
Fazlyab, M., Robey, A., Hassani, H., Morari, M., and Pappas, G · 2019
Cited alongside, same era.
FFJORD: Free-form continuous dynamics for scalable reversible generative models
Grathwohl, W., Chen, R. T. Q., Betterncourt, J., Sutskever, I., and Duvenaud, D · 2019
Cited alongside, same era.
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Closest in time.
Lipschitz constant estimation of neural networks via sparse polynomial optimization
Latorre, F., Rolland, P., and Cevher, V · 2020
Closest in time.
Stabilizing Transformers for reinforcement learning
Parisotto, E., Song, H. F., Rae, J. W., Pascanu, R., Gulcehre, C., Jayakumar, S. M., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., Botvinick, M. M., Heess, N., and Hadsell, R · 2020
Closest in time.
A mathematical theory of attention
Vuckovic, J., Baratin, A., and Tachet des Combes, R · 2020
Closest in time.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Closest in time.