Fetching the paper…
Reading the bibliography…
While post-training model compression can greatly reduce the inference cost of a deep neural network, uncompressed training still consumes a huge amount of hardware resources, run-time and energy.
L. R. Tucker, “Some mathematical notes on three-mode factor analysis,” Psychometrika , vol. 31, no. 3, pp. 279–311, 1966
1966
Earlier work this paper cites.
J. D. Carroll and J.-J. Chang, “Analysis of individual differences in multidimensional scaling via an N-way generalization of “Eckart-Young” decomposition,” Psychometrika , vol. 35, no. 3, pp. 283–319, 1970
1970
Earlier work this paper cites.
S. J. Hanson and L. Y. Pratt, “Comparing biases for minimal network construction with back-propagation,” in NIPS , 1989, pp. 177–185
1989
Earlier work this paper cites.
Y. LeCun, J. S. Denker, and S. A. Solla, “Optimal brain damage,” in NIPS , 1990, pp. 598–605
1990
Earlier work this paper cites.
R. A. Harshman, M. E. Lundy et al. , “PARAFAC: Parallel factor analysis,” Computational Statistics and Data Analysis , vol. 18, no. 1, pp. 39–72, 1994
1994
Earlier work this paper cites.
M. I. Jordan, Z. Ghahramani, T. S. Jaakkola, and L. K. Saul, “An introduction to variational methods for graphical models,” Machine learning , vol. 37, no. 2, pp. 183–233, 1999
1999
Earlier work this paper cites.
T. G. Kolda and B. W. Bader, “Tensor decompositions and applications,” SIAM Review , vol. 51, no. 3, pp. 455–500, 2009
2009
Earlier work this paper cites.
C. M. Carvalho, N. G. Polson, and J. G. Scott, “The horseshoe estimator for sparse signals,” Biometrika , vol. 97, no. 2, pp. 465–480, 2010
2010
Earlier work this paper cites.
S. Gandy, B. Recht, and I. Yamada, “Tensor completion and low-n-rank tensor recovery via convex optimization,” Inverse Problems , vol. 27, no. 2, p. 025010, 2011
2011
Earlier work this paper cites.
I. V. Oseledets, “Tensor-train decomposition,” SIAM J. Sci. Computing , vol. 33, no. 5, pp. 2295–2317, 2011
2011
Earlier work this paper cites.
M. P. Wand, J. T. Ormerod, S. A. Padoan, R. Frühwirth et al. , “Mean field variational Bayes for elaborate distributions,” Bayesian Analysis , vol. 6, no. 4, pp. 847–900, 2011
2011
Earlier work this paper cites.
R. M. Neal, Bayesian learning for neural networks . Springer Science & Business Media, 2012, vol. 118
2012
Earlier work this paper cites.
T. N. Sainath, B. Kingsbury, V. Sindhwani, E. Arisoy, and B. Ramabhadran, “Low-rank matrix factorization for deep neural network training with high-dimensional output targets,” in ICASSP , 2013, pp. 6655–6659
2013
Earlier work this paper cites.
J. Xue, J. Li, and Y. Gong, “Restructuring of deep neural network acoustic models with singular value decomposition.” in Interspeech , 2013, pp. 2365–2369
2013
Earlier work this paper cites.
M. D. Hoffman, D. M. Blei, C. Wang, and J. Paisley, “Stochastic variational inference,” The Journal of Machine Learning Research , vol. 14, no. 1, pp. 1303–1347, 2013
2013
Earlier work this paper cites.
H. Zhou, L. Li, and H. Zhu, “Tensor regression with applications in neuroimaging data analysis,” Journal of the American Statistical Association , vol. 108, no. 502, pp. 540–552, 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
D. Goldfarb and Z. Qin, “Robust low-rank tensor recovery: Models and algorithms,” SIAM Journal on Matrix Analysis and Applications , vol. 35, no. 1, pp. 225–253, 2014
2014
Earlier work this paper cites.
S. Nakajima and M. Sugiyama, “Analysis of empirical MAP and empirical partially Bayes: Can they be alternatives to variational Bayes?” in Artificial Intelligence and Statistics , 2014, pp. 20–28
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep learning with limited numerical precision,” in International Conference on Machine Learning , 2015, pp. 1737–1746
2015
Cited alongside, same era.
A. Novikov, D. Podoprikhin, A. Osokin, and D. P. Vetrov, “Tensorizing neural networks,” in NIPS , 2015, pp. 442–450
2015
Cited alongside, same era.
Q. Zhao, L. Zhang, and A. Cichocki, “Bayesian CP factorization of incomplete tensors with automatic rank determination,” IEEE TPAMI , vol. 37, no. 9, pp. 1751–1763, 2015
C. Hawkins and Z. Zhang, “Variational Bayesian inference for robust streaming tensor factorization and completion,” in 2018 IEEE International Conference on Data Mining . IEEE, 2018, pp. 1446–1451
2018
Later among the works it cites.
2019
Later among the works it cites.
Y. Ma, R. Chen, W. Li, F. Shang, W. Yu, M. Cho, and B. Yu, “A unified approximation framework for compressing and accelerating deep neural networks,” in International Conf. on Tools with Artificial Intelligence , 2019, pp. 376–383
2019
Later among the works it cites.
C. Cui, C. Hawkins, and Z. Zhang, “Tensor methods for generating compact uncertainty quantification and deep learning models,” in Intl. Conf. Computer-Aided Design , 2019, pp. 1–6
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra, “Weight uncertainty in neural network,” in International Conference on Machine Learning . PMLR, 2015, pp. 1613–1622
2015
Cited alongside, same era.
2016
Cited alongside, same era.
W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” NIPS , vol. 29, pp. 2074–2082, 2016
2016
Cited alongside, same era.
Q. Zhao, G. Zhou, L. Zhang, A. Cichocki, and S.-I. Amari, “Bayesian robust tensor factorization for incomplete multiway data,” IEEE TNNLS , vol. 27, no. 4, pp. 736–748, 2016
2016
Cited alongside, same era.
Q. Liu and D. Wang, “Stein variational gradient descent: A general purpose Bayesian inference algorithm,” in NIPS , 2016, pp. 2378–2386
2016
Cited alongside, same era.
J. M. Alvarez and M. Salzmann, “Compression-aware training of deep networks,” in NIPS , 2017, pp. 856–867
2017
Cited alongside, same era.
K. Neklyudov, D. Molchanov, A. Ashukha, and D. P. Vetrov, “Structured Bayesian pruning via log-normal multiplicative noise,” in NIPS , 2017, pp. 6775–6784
2017
Cited alongside, same era.
X. Ma, P. Zhang, S. Zhang, N. Duan, Y. Hou, M. Zhou, and D. Song, “A tensorized transformer for language modeling,” in NIPS , 2019, pp. 2232–2242
2019
Later among the works it cites.
C. Deng, F. Sun, X. Qian, J. Lin, Z. Wang, and B. Yuan, “TIE: energy-efficient tensor train-based inference engine for deep neural network,” in ISCA , 2019, pp. 264–278
2019
Later among the works it cites.
K. Zhang, X. Zhang, and Z. Zhang, “Tucker tensor decomposition on FPGA,” in Proc. Intl. Conf. Computer-Aided Design , 2019, pp. 1–8
2019
Later among the works it cites.
E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy considerations for deep learning in NLP,” in Proc. Annual Meeting of the Association for Computational Linguistics , 2019, pp. 3645–3650
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
——, “Bayesian tensorized neural networks with automatic rank selection,” arXiv:1905.10478 , 2019
2019
Later among the works it cites.
P. Zhen, B. Liu, Y. Cheng, H.-B. Chen, and H. Yu, “Fast video facial expression recognition by deeply tensor-compressed lstm neural network on mobile device,” in Proceedings of the 4th ACM/IEEE Symposium on Edge Computing , 2019, pp. 298–300
2019
Later among the works it cites.
S. Ghosh, J. Yao, and F. Doshi-Velez, “Model selection in Bayesian neural networks via horseshoe priors,” Journal of Machine Learning Research , vol. 20, no. 182, pp. 1–46, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
X. Sun, N. Wang, C.-Y. Chen, J. Ni, A. Agrawal, X. Cui, S. Venkataramani, K. El Maghraoui, V. V. Srinivasan, and K. Gopalakrishnan, “Ultra-low precision 4-bit training of deep neural networks,” NIPS , vol. 33, 2020
2020
Closest in time.
O. Hrinchuk, V. Khrulkov, L. Mirvakhabova, E. Orlova, and I. Oseledets, “Tensorized embedding layers,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings , 2020, pp. 4847–4860
2020
Closest in time.
J. Kossaifi, A. Toisoul, A. Bulat, Y. Panagakis, T. M. Hospedales, and M. Pantic, “Factorized higher-order CNNs with an application to spatio-temporal emotion estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 6060–6069
2020
Closest in time.
D. Yang, W. Yu, H. Mu, and G. Yao, “Dynamic programming assisted quantization approaches for compressing normal and robust DNN models,” in Proc. Asia and South Pacific Design Automation Conference , 2021, pp. 351–357
2021
Closest in time.
K. Zhang, C. Hawkins, X. Zhang, C. Hao, and Z. Zhang, “On-FPGA training with ultra memory reduction: A low-precision tensor method,” ICLR Workshop of Hardware Aware Efficient Training , May 2021
2021
Closest in time.
A. Kolbeinsson, J. Kossaifi, Y. Panagakis, A. Bulat, A. Anandkumar, I. Tzoulaki, and P. M. Matthews, “Tensor dropout for robust learning,” IEEE Journal of Selected Topics in Signal Processing , vol. 15, no. 3, pp. 630–640, 2021
2021
Closest in time.