Fetching the paper…
Reading the bibliography…
Adversarial robustness is a key desirable property of neural networks.
On the pseudoinverse of a sum of symmetric matrices with applications to estimation
P. Kovanic · 1979
Earlier work this paper cites.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
M. I. Jordan and R. A. Jacobs · 1994
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction , volume 2
T. Hastie, R. Tibshirani, J. H. Friedman, and J. H. Friedman · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Twenty years of mixture of experts
S. E. Yuksel, J. N. Wilson, and P. D. Gader · 2012
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
D. Eigen, M. Ranzato, and I. Sutskever · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2015
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Cited alongside, same era.
Revisiting unreasonable effectiveness of data in deep learning era
C. Sun, A. Shrivastava, S. Singh, and A. Gupta · 2017
Cited alongside, same era.
Towards deep learning models resistant to adversarial attacks
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2018
Cited alongside, same era.
Are labels required for improving adversarial robustness?
J.-B. Alayrac, J. Uesato, P.-S. Huang, A. Fawzi, R. Stanforth, and P. Kohli · 2019
Cited alongside, same era.
On evaluating adversarial robustness
N. Carlini, A. Athalye, N. Papernot, W. Brendel, J. Rauber, D. Tsipras, I. Goodfellow, A. Madry, and A. Kurakin · 2019
Cited alongside, same era.
Intriguing properties of adversarial training at scale
C. Xie and A. Yuille · 2020
Later among the works it cites.
Sparse moes meet efficient ensembles
J. U. Allingham, F. Wenzel, Z. E. Mariet, B. Mustafa, J. Puigcerver, N. Houlsby, G. Jerfel, V. Fortuin, B. Lakshminarayanan, J. Snoek, D. Tran, C. R. Ruiz, and R. Jenatton · 2021
Later among the works it cites.
A universal law of robustness via isoperimetry
S. Bubeck and M. Sellke · 2021
Later among the works it cites.
Base layers: Simplifying training of large, sparse models
M. Lewis, S. Bhosale, T. Dettmers, N. Goyal, and L. Zettlemoyer · 2021
Later among the works it cites.
Carbon emissions and large neural network training
D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
F. Croce and M. Hein · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Uncovering the limits of adversarial training against norm-bounded adversarial examples
S. Gowal, C. Qin, J. Uesato, T. Mann, and P. Kohli · 2020
Cited alongside, same era.
GShard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2020
Cited alongside, same era.
Scaling Vision with Sparse Mixture of Experts
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. Susano Pinto, D. Keysers, and N. Houlsby · 2021
Later among the works it cites.
F. Xue, Z. Shi, F. Wei, Y. Lou, Y. Liu, and Y. You · 2021
Later among the works it cites.
Glam: Efficient scaling of language models with mixture-of-experts
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, et al · 2022
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2022
Closest in time.