Fetching the paper…
Reading the bibliography…
Large neural networks excel in many domains, but they are expensive to train and fine-tune.
The approximation of one matrix by another of lower rank
Eckart, C. and Young, G · 1936
Earlier work this paper cites.
An algorithm for the machine calculation of complex fourier series
Cooley, J. W. and Tukey, J. W · 1965
Earlier work this paper cites.
A generalized solution of the orthogonal procrustes problem
Schönemann, P. H · 1966
Earlier work this paper cites.
Sparse matrices , volume 69
Tewarson, R. P · 1973
Earlier work this paper cites.
Displacement ranks of matrices and linear equations
Kailath, T., Kung, S.-Y., and Morf, M · 1979
Earlier work this paper cites.
FFTs in external or hierarchical memory
Bailey, D. H · 1990
Earlier work this paper cites.
Numerical methods for simultaneous diagonalization
Bunse-Gerstner, A., Byers, R., and Mehrmann, V · 1993
Earlier work this paper cites.
Random butterfly transformations with applications in computational linear algebra
Parker, D. S · 1995
Earlier work this paper cites.
Fast discrete polynomial transforms with applications to data analysis for distance transitive graphs
Driscoll, J. R., Healy Jr, D. M., and Rockmore, D. N · 1997
Earlier work this paper cites.
On a new class of structured matrices
Eidelman, Y. and Gohberg, I · 1999
Earlier work this paper cites.
Sense: sensitivity encoding for fast MRI
Pruessmann, K. P., Weiger, M., Scheidegger, M. B., and Boesiger, P · 1999
Earlier work this paper cites.
Spectral methods in MATLAB
Trefethen, L. N · 2000
Earlier work this paper cites.
Generalized autocalibrating partially parallel acquisitions (grappa)
Griswold, M. A., Jakob, P. M., Heidemann, R. M., Nittka, M., Jellus, V., Wang, J., Kiefer, B., and Haase, A · 2002
Earlier work this paper cites.
Picking winning tickets before training by preserving gradient flow
Wang, C., Zhang, G., and Grosse, R · 2002
Earlier work this paper cites.
Computed tomography: principles, design, artifacts, and recent advances , volume 114
Hsieh, J · 2003
Earlier work this paper cites.
Toeplitz and circulant matrices: A review
Gray, R. M · 2006
Earlier work this paper cites.
Sparse MRI: The application of compressed sensing for rapid mr imaging
Lustig, M., Donoho, D., and Pauly, J. M · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Structured matrices and polynomials: unified superfast algorithms
Pan, V. Y · 2012
Earlier work this paper cites.
Low-rank modeling of local k k -space neighborhoods (loraks) for constrained MRI
Haldar, J. P · 2013
Earlier work this paper cites.
Fastfood-computing hilbert space expansions in loglinear time
Le, Q., Sarlós, T., and Smola, A · 2013
Earlier work this paper cites.
Speech and language processing , volume 3
Jurafsky, D. and Martin, J. H · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Structured transforms for small-footprint deep learning
Sindhwani, V., Sainath, T., and Kumar, S · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Dsd: Dense-sparse-dense training for deep neural networks
Han, S., Pool, J., Narang, S., Mao, H., Gong, E., Tang, S., Elsen, E., Vajda, P., Paluri, M., Tran, J., et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Flexible multilayer sparse approximations of matrices and applications
Le Magoarou, L. and Gribonval, R · 2016
Earlier work this paper cites.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
ACDC: a structured efficient linear layer
Moczulski, M., Denil, M., Appleyard, J., and de Freitas, N · 2016
Earlier work this paper cites.
Beyond low rank+ sparse: Multiscale low rank matrix decomposition
Ong, F. and Lustig, M · 2016
Earlier work this paper cites.
Orthogonal random features
Yu, F. X., Suresh, A. T., Choromanski, K. M., Holtmann-Rice, D. N., and Kumar, S · 2016
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Dong, X., Chen, S., and Pan, S. J · 2017
Earlier work this paper cites.
GPU kernels for block-sparse weights
Gray, S., Radford, A., and Kingma, D. P · 2017
Cited alongside, same era.
Runtime neural pruning
Lin, J., Rao, Y., Lu, J., and Zhou, J · 2017
Cited alongside, same era.
Low-rank and adaptive sparse signal (lassi) models for highly accelerated dynamic imaging
Ravishankar, S., Moore, B. E., Nadakuditi, R. R., and Fessler, J. A · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Cited alongside, same era.
A two-pronged progress in structured dense matrix vector multiplication
De Sa, C., Gu, A., Puttagunta, R., Ré, C., and Rudra, A · 2018
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
Neural controlled differential equations for irregular time series
Kidger, P., Morrill, J., Foster, J., and Lyons, T · 2020
Later among the works it cites.
Deep-learning methods for parallel magnetic resonance imaging reconstruction: A survey of the current approaches, trends, and issues
Knoll, F., Hammernik, K., Zhang, C., Moeller, S., Pock, T., Sodickson, D. K., and Akcakaya, M · 2020
Later among the works it cites.
Fourier neural operator for parametric partial differential equations
Li, Z., Kovachki, N. B., Azizzadenesheli, K., Bhattacharya, K., Stuart, A., Anandkumar, A., et al · 2020
Later among the works it cites.
Finding trainable sparse networks through neural tangent transfer
Liu, T. and Zenke, F · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Cited alongside, same era.
Learning a variational network for reconstruction of accelerated MRI data
Hammernik, K., Klatzer, T., Kobler, E., Recht, M. P., Sodickson, D. K., Pock, T., and Knoll, F · 2018
Cited alongside, same era.
Deep generative adversarial neural networks for compressive sensing MRI
Mardani, M., Gong, E., Cheng, J. Y., Vasanawala, S. S., Zaharchuk, G., Xing, L., and Pauly, J. M · 2018
Cited alongside, same era.
Quadrature-based features for kernel approximation
Munkhoeva, M., Kapushev, Y., Burnaev, E., and Oseledets, I · 2018
Cited alongside, same era.
Learning compressed transforms with low displacement rank
Thomas, A., Gu, A., Dao, T., Rudra, A., and Ré, C · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Cited alongside, same era.
Later among the works it cites.
Mlperf training benchmark
Mattson, P., Cheng, C., Diamos, G., Coleman, C., Micikevicius, P., Patterson, D., Tang, H., Wei, G.-Y., Bailis, P., Bittorf, V., et al · 2020
Later among the works it cites.
Logarithmic pruning is all you need
Orseau, L., Hutter, M., and Rivasplata, O · 2020
Later among the works it cites.
Optimal lottery tickets via subsetsum: Logarithmic over-parameterization is sufficient
Pensia, A., Rajput, S., Nagle, A., Vishwakarma, H., and Papailiopoulos, D · 2020
Later among the works it cites.
Hypersolvers: Toward fast continuous-depth models
Poli, M., Massaroli, S., Yamashita, A., Asama, H., Park, J., et al · 2020
Later among the works it cites.
Universal differential equations for scientific machine learning
Rackauckas, C., Ma, Y., Martensen, J., Warner, C., Zubov, K., Supekar, R., Skinner, D., Ramadhan, A., and Edelman, A · 2020
Later among the works it cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Later among the works it cites.
Compressed sensing: From research to clinical practice with deep neural networks: Shortening scan times for magnetic resonance imaging
Sandino, C. M., Cheng, J. Y., Chen, F., Mardani, M., Pauly, J. M., and Vasanawala, S. S · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A. M · 2020
Later among the works it cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
Tanaka, H., Kunin, D., Yamins, D. L., and Ganguli, S · 2020
Later among the works it cites.
Butterfly transform: An efficient fft based neural architecture design
Vahid, K. A., Prabhu, A., Farhadi, A., and Rastegari, M · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Later among the works it cites.
Self-supervised physics-based deep learning MRI reconstruction without fully-sampled data
Yaman, B., Hosseini, S. A. H., Moeller, S., Ellermann, J., Uğurbil, K., and Akçakaya, M · 2020
Later among the works it cites.
Sparse linear networks with a fixed butterfly structure: theory and practice
Ailon, N., Leibovitch, O., and Nair, V · 2021
Later among the works it cites.
Accelerated MRI with un-trained neural networks
Darestani, M. Z. and Heckel, R · 2021
Later among the works it cites.
Measuring robustness in deep learning based compressive sensing
Darestani, M. Z., Chaudhari, A., and Heckel, R · 2021
Later among the works it cites.
A framework for few-shot language model evaluation, September 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2021
Later among the works it cites.
Top-KAST: Top-K always sparse training
Jayakumar, S. M., Pascanu, R., Rae, J. W., Osindero, S., and Elsen, E · 2021
Later among the works it cites.
Gotta go fast when generating data with score-based models
Jolicoeur-Martineau, A., Li, K., Piché-Taillefer, R., Kachman, T., and Mitliagkas, I · 2021
Later among the works it cites.
Sparse factorization of large square matrices
Khalitov, R., Yu, T., Cheng, L., and Yang, Z · 2021
Later among the works it cites.
Machine learning–accelerated computational fluid dynamics
Kochkov, D., Smith, J. A., Alieva, A., Wang, Q., Brenner, M. P., and Hoyer, S · 2021
Later among the works it cites.
Block pruning for faster transformers
Lagunas, F., Charlaix, E., Sanh, V., and Rush, A. M · 2021
Later among the works it cites.
Blind primed supervised (blips) learning for mr image reconstruction
Lahiri, A., Wang, G., Ravishankar, S., and Fessler, J. A · 2021
Later among the works it cites.
Deformable butterfly: A highly structured and sparse linear transform
Lin, R., Ran, J., Chiu, K. H., Chesi, G., and Wong, N · 2021
Later among the works it cites.
Differentiable multiple shooting layers
Massaroli, S., Poli, M., Sonoda, S., Suzuki, T., Park, J., Yamashita, A., and Asama, H · 2021
Later among the works it cites.
Ac/dc: Alternating compressed/decompressed training of deep neural networks
Peste, A., Iofinova, E., Vladu, A., and Alistarh, D · 2021
Later among the works it cites.
Mlp-Mixer: An all-mlp architecture for vision
Tolstikhin, I., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Keysers, D., Uszkoreit, J., Lucic, M., et al · 2021
Later among the works it cites.
Tokens-to-token ViT: Training vision transformers from scratch on imagenet
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Tay, F. E., Feng, J., and Yan, S · 2021
Later among the works it cites.
Calibrate before use: Improving few-shot performance of language models
Zhao, T. Z., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Later among the works it cites.
Pixelated butterfly: Simple and efficient sparse training for neural network models
Chen, B., Dao, T., Liang, K., Yang, J., Song, Z., Rudra, A., and Ré, C · 2022
Closest in time.