Fetching the paper…
Reading the bibliography…
Fast linear transforms are ubiquitous in machine learning, including the discrete Fourier transform, discrete cosine transform, and other structured transformations such as convolutions.
Orthogonal Polynomials
Szegö, G · 1967
Earlier work this paper cites.
Displacement ranks of matrices and linear equations
Kailath, T., Kung, S.-Y., and Morf, M · 1979
Earlier work this paper cites.
A fast cosine transform in one and two dimensions
Makhoul, J · 1980
Earlier work this paper cites.
Random butterfly transformations with applications in computational linear algebra
Parker, D. S · 1995
Earlier work this paper cites.
Fast discrete polynomial transforms with applications to data analysis for distance transitive graphs
Driscoll, J. R., Healy, Jr., D. M., and Rockmore, D. N · 1997
Earlier work this paper cites.
Almost linear VC dimension bounds for piecewise polynomial networks
Bartlett, P. L., Maiorov, V., and Meir, R · 1999
Earlier work this paper cites.
Guest editors’ introduction: The top 10 algorithms
Dongarra, J. and Sullivan, F · 2000
Earlier work this paper cites.
Matrix-vector product for confluent cauchy-like matrices with application to confluent rational interpolation
Olshevsky, V. and Shokrollahi, M. A · 2000
Earlier work this paper cites.
Automatic generation of fast discrete signal transforms
Egner, S. and Püschel, M · 2001
Earlier work this paper cites.
Structured Matrices and Polynomials: Unified Superfast Algorithms
Pan, V. Y · 2001
Earlier work this paper cites.
Digital signal processing: principles algorithms and applications
Proakis, J. G · 2001
Earlier work this paper cites.
Symmetry-based matrix factorization
Egner, S. and Püschel, M · 2004
Earlier work this paper cites.
Semi-supervised learning by entropy minimization
Grandvalet, Y. and Bengio, Y · 2005
Earlier work this paper cites.
Sparse principal component analysis
Zou, H., Hastie, T., and Tibshirani, R · 2006
Earlier work this paper cites.
Algebraic signal processing theory
Puschel, M. and Moura, J. M · 2008
Earlier work this paper cites.
Supervised dictionary learning
Mairal, J., Ponce, J., Sapiro, G., Zisserman, A., and Bach, F. R · 2009
Earlier work this paper cites.
Algebraic signal processing theory: Cooley–Tukey type algorithms for real DFTs
Voronenko, Y. and Puschel, M · 2009
Earlier work this paper cites.
Robust principal component analysis?
Candès, E. J., Li, X., Ma, Y., and Wright, J · 2011
Cited alongside, same era.
An introduction to orthogonal polynomials
Chihara, T · 2011
Cited alongside, same era.
Algebraic complexity theory , volume 315
Bürgisser, P., Clausen, M., and Shokrollahi, M. A · 2013
Cited alongside, same era.
Predicting parameters in deep learning
Denil, M., Shakibi, B., Dinh, L., De Freitas, N., et al · 2013
Cited alongside, same era.
Fastfood - computing Hilbert space expansions in loglinear time
Le, Q., Sarlos, T., and Smola, A · 2013
Cited alongside, same era.
Neyshabur, B. and Panigrahy, R · 2013
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Later among the works it cites.
Flexible multilayer sparse approximations of matrices and applications
Le Magoarou, L. and Gribonval, R · 2016
Later among the works it cites.
Every matrix is a product of toeplitz matrices
Ye, K. and Lim, L.-H · 2016
Later among the works it cites.
Orthogonal random features
Yu, F. X. X., Suresh, A. T., Choromanski, K. M., Holtmann-Rice, D. N., and Kumar, S · 2016
Later among the works it cites.
CirCNN: accelerating and compressing deep neural networks using block-circulant weight matrices
Ding, C., Liao, S., Wang, Y., Li, Z., Liu, N., Zhuo, Y., Wang, C., Qian, X., Bai, Y., Yuan, G., et al · 2017
Later among the works it cites.
Nearly-tight VC-dimension bounds for piecewise linear neural networks
Harvey, N., Liaw, C., and Mehrabian, A · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Cited alongside, same era.
Speech and language processing , volume 3
Jurafsky, D. and Martin, J. H · 2014
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Cited alongside, same era.
Compressing neural networks with the hashing trick
Chen, W., Wilson, J., Tyree, S., Weinberger, K., and Chen, Y · 2015
Cited alongside, same era.
An exploration of parameter redundancy in deep networks with circulant projections
Cheng, Y., Yu, F. X., Feris, R. S., Kumar, S., Choudhary, A., and Chang, S.-F · 2015
Cited alongside, same era.
Chasing butterflies: In search of efficient dictionaries
Le Magoarou, L. and Gribonval, R · 2015
Cited alongside, same era.
Later among the works it cites.
Tunable efficient unitary neural networks (eunn) and their application to rnns
Jing, L., Shen, Y., Dubcek, T., Peurifoy, J., Skirlo, S., LeCun, Y., Tegmark, M., and Soljačić, M · 2017
Later among the works it cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Li, L., Jamieson, K., DeSalvo, G., Rostamizadeh, A., and Talwalkar, A · 2017
Later among the works it cites.
A two-pronged progress in structured dense matrix vector multiplication
De Sa, C., Gu, A., Puttagunta, R., Ré, C., and Rudra, A · 2018
Later among the works it cites.
Learning latent permutations with Gumbel-Sinkhorn networks
Mena, G., Belanger, D., Linderman, S., and Snoek, J · 2018
Later among the works it cites.
Quadrature-based features for kernel approximation
Munkhoeva, M., Kapushev, Y., Burnaev, E., and Oseledets, I · 2018
Later among the works it cites.
Learning compressed transforms with low displacement rank
Thomas, A., Gu, A., Dao, T., Rudra, A., and Ré, C · 2018
Later among the works it cites.
Neural arithmetic logic units
Trask, A., Hill, F., Reed, S. E., Rae, J., Dyer, C., and Blunsom, P · 2018
Later among the works it cites.
StrassenNets: Deep learning with a multiplication budget
Tschannen, M., Khanna, A., and Anandkumar, A · 2018
Later among the works it cites.
A semantic loss function for deep learning with symbolic knowledge
Xu, J., Zhang, Z., Friedman, T., Liang, Y., and Van den Broeck, G · 2018
Later among the works it cites.
Stochastic optimization of sorting networks via continuous relaxations
Grover, A., Wang, E., Zweig, A., and Ermon, S · 2019
Closest in time.
Learning neural PDE solvers with convergence guarantees
Hsieh, J.-T., Zhao, S., Eismann, S., Mirabella, L., and Ermon, S · 2019
Closest in time.