Fetching the paper…
Reading the bibliography…
Neural network models are widely used in solving many challenging problems, such as computer vision, personalized recommendation, and natural language processing.
F. L. Hitchcock, “The expression of a tensor or a polyadic as a sum of products,” J. of Mathematics and Physics , vol. 6, no. 1-4, pp. 164–189, 1927
1927
Earlier work this paper cites.
Hitchcock, Frank L., “Multiple invariants and generalized rank of a p-way matrix or tensor,” J. of Mathematics and Physics , vol. 7, no. 1-4, pp. 39–79, 1928
1928
Earlier work this paper cites.
J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi, “Modeling task relationships in multi-task learning with multi-gate mixture-of-experts,” in Proc. Int. Conf. Knowledge Discovery & Data Mining , 2018, p. 1930–1939
1939
Earlier work this paper cites.
L. R. Tucker, “Implications of factor analysis of three-way matrices for measurement of change,” in Problems in measuring change. , C. W. Harris, Ed. University of Wisconsin Press, 1963, pp. 122–137
1963
Earlier work this paper cites.
W. F. Tinney and J. W. Walker, “Direct solutions of sparse network equations by optimally ordered triangular factorization,” Proc. IEEE , vol. 55, no. 11, pp. 1801–1809, 1967
1967
Earlier work this paper cites.
J. D. Carroll and J.-J. Chang, “Analysis of individual differences in multidimensional scaling via an n-way generalization of “eckart-young” decomposition,” Psychometrika , vol. 35, pp. 283–319, 1970
1970
Earlier work this paper cites.
J. D. Carroll, S. Pruzansky, and J. B. Kruskal, “Candelinc: A general approach to multidimensional analysis of many-way arrays with linear constraints on parameters,” Psychometrika , vol. 45, pp. 3–24, 1980
1980
Earlier work this paper cites.
S. Winograd, Arithmetic Complexity of Computations . Siam, 1980, vol. 33
1980
Earlier work this paper cites.
H. Summala, “Risk control is not risk adjustment: The zero-risk theory of driver behaviour and its implications,” Ergonomics , pp. 491–506, 1988
1988
Earlier work this paper cites.
R. A. Harshman, Margaret, and E. Lundy, “Uniqueness proof for a family of models sharing features of tucker’s three-mode factor analysis and parafac/candecomp,” Psychometrika , 1996
1996
Earlier work this paper cites.
A. Degenne and M. Forsé, Introducing social networks . Sage, 1999
1999
Earlier work this paper cites.
D. J. Watts, Small Worlds: The Dynamics of Networks between Order and Randomness . Princeton University Press, 2004, vol. 9
2004
Earlier work this paper cites.
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE Trans. Neural Networks , vol. 20, no. 1, pp. 61–80, 2008
2008
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Proc. Conf. Neural Information Processing Systems , Dec. 2012, pp. 1097–1105
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” in Proc. Conf. Neural Information Processing Systems , 2015, pp. 1135–1143
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: surpassing human-level performance on ImagerNet classification,” in Proc. Int. Conf. Computer Vision , 2015, pp. 1026–1034
2015
Earlier work this paper cites.
K. Makantasis, K. Karantzalos, A. Doulamis, and N. Doulamis, “Deep supervised learning for hyperspectral data classification through convolutional neural networks,” in Int. Geoscience and Remote Sensing Symp. , 2015, pp. 4959–4962
2015
Earlier work this paper cites.
M. Abadi, P. Barham, J. Chen et al. , “TensorFlow: A system for large-scale machine learning,” in Proc. Symp. Operating Systems Design & Implementation , 2016, pp. 265–283
2016
Earlier work this paper cites.
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-neuron-free deep neural network computing,” in Proc. Int. Symp. Computer Architecture , 2016, pp. 1–13
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “EIE: efficient inference engine on compressed deep neural network,” in Proc. Int. Symp. Computer Architecture , 2016, pp. 243–254
2016
Earlier work this paper cites.
S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” in Proc. Int. Conf. Learning Representations , 2016
2016
Earlier work this paper cites.
A. Lavin and S. Gray, “Fast algorithms for convolutional neural networks,” in Proc. Conf. Computer Vision and Pattern Recognition , June 2016, pp. 4013–4021
2016
Earlier work this paper cites.
S. Teerapittayanon, B. McDanel, and H.-T. Kung, “BranchyNet: Fast inference via early exiting from deep neural networks,” in Proc. Int. Conf. Pattern Recognition , 2016, pp. 2464–2469
2016
Earlier work this paper cites.
S. Zhang, Z. Du, L. Zhang, H. Lan, S. Liu, L. Li, Q. Guo, T. Chen, and Y. Chen, “Cambricon-X: An accelerator for sparse neural networks,” in Proc. Int. Symp. Microarchitecture , 2016, pp. 1–12
2016
Earlier work this paper cites.
A. Ardakani, C. Condo, and W. J. Gross, “Activation pruning of deep convolutional neural networks,” in IEEE Global Conf. Signal and Information Processing (GlobalSIP) , 2017, pp. 1325–1329
2017
Earlier work this paper cites.
T. Bolukbasi, J. Wang, O. Dekel, and V. Saligrama, “Adaptive neural networks for efficient inference,” in Proc. Int. Conf. Machine Learning , 2017, pp. 527–536
2017
Earlier work this paper cites.
Y.-H. Chen, T. Krishna, J. S. Emer, and V. Sze, “Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks,” IEEE J. Solid-State Circuits , vol. 52, no. 1, pp. 127–138, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Figurnov, M. D. Collins, Y. Zhu, L. Zhang, J. Huang, D. P. Vetrov, and R. Salakhutdinov, “Spatially adaptive computation time for residual networks,” Proc. Conf. Computer Vision and Pattern Recognition , pp. 1790–1799, 2017
2017
Earlier work this paper cites.
S. Filippone, V. Cardellini, D. Barbieri, and A. Fanfarillo, “Sparse matrix-vector multiplication on GPGPUs,” ACM trans. Mathematical Software , vol. 43, no. 4, pp. 1–49, Mar. 2017
2017
Earlier work this paper cites.
S. Gray, A. Radford, and D. P. Kingma, “GPU kernels for block-sparse weights,” Technical report, OpenAI, Tech. Rep., 2017
2017
Cited alongside, same era.
S. Han, J. Kang, H. Mao, Y. Hu, X. Li, Y. Li, D. Xie, H. Luo, S. Yao, Y. Wang, H. Yang, and W. J. Dally, “ESE: Efficient speech recognition engine with sparse LSTM on FPGA,” in Proc. Int. Symp. Field-Programmable Gate Arrays , 2017, pp. 75–84
2017
Cited alongside, same era.
G. Huang, D. Chen, T. Li, F. Wu, L. van der Maaten, and K. Q. Weinberger, “Multi-scale dense networks for resource efficient image classification,” in Proc. Int. Conf. Learning Representations , 2017
2017
Cited alongside, same era.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Quantized neural networks: Training neural networks with low precision weights and activations,” J. Machine Learning Research , vol. 18, no. 1, p. 6869–6898, Jan. 2017
2017
Cited alongside, same era.
R. Teja Mullapudi, W. R. Mark, N. Shazeer, and K. Fatahalian, “HydraNets: Specialized dynamic architectures for efficient inference,” in Proc. Conf. Computer Vision and Pattern Recognition , 2018, pp. 8080–8089
2018
Later among the works it cites.
A. Veit and S. Belongie, “Convolutional networks with adaptive inference graphs,” in Proc. European Conf. Computer Vision. , 2018, pp. 3–18
2018
Later among the works it cites.
X. Wang, F. Yu, Z.-Y. Dou, T. Darrell, and J. E. Gonzalez, “SkipNet: Learning dynamic routing in convolutional networks,” in Proc. European Conf. Computer Vision. , 2018, pp. 409–424
2018
Later among the works it cites.
Z. Wu, T. Nagarajan, A. Kumar, S. Rennie, L. S. Davis, K. Grauman, and R. Feris, “Blockdrop: Dynamic inference paths in residual networks,” in Proc. Conf. Computer Vision and Pattern Recognition , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Lin, C. Sakr, Y. Kim, and N. Shanbhag, “PredictiveNet: an energy-efficient convolutional neural network via zero prediction,” in Proc. Int. Symp. Circuits & Systems , 2017, pp. 1–4
2017
Cited alongside, same era.
D. Neil, J. H. Lee, T. Delbruck, and S.-C. Liu, “Delta networks for optimized recurrent network computation,” in Proc. Int. Conf. Machine Learning , 2017, pp. 2584–2593
2017
Cited alongside, same era.
A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “SCNN: An accelerator for compressed-sparse convolutional neural networks,” in Proc. Int. Symp. Computer Architecture , 2017, pp. 27–40
2017
Cited alongside, same era.
H. Pratt, B. Williams, F. Coenen, and Y. Zheng, “FCNN: Fourier convolutional neural networks,” in Machine Learning and Knowledge Discovery in Databases , 2017, pp. 786–798
2017
Cited alongside, same era.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” in Proc. Conf. Neural Information Processing Systems , 2017, pp. 5099–5108
2017
Cited alongside, same era.
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in Proc. Int. Conf. Learning Representations , 2017
2017
Cited alongside, same era.
T. Zhang, S. Ye, K. Zhang, J. Tang, W. Wen, M. Fardad, and Y. Wang, “A systematic DNN weight pruning framework using alternating direction method of multipliers,” in Proc. European Conf. Computer Vision. , 2018, pp. 184–199
2018
Later among the works it cites.
X. Zhang, C. Xie, J. Wang, W. Zhang, and X. Fu, “Towards memory friendly long-short term memory networks (LSTMs) on mobile GPUs,” in Proc. Int. Symp. Microarchitecture , Oct 2018, pp. 162–174
2018
Later among the works it cites.
X. Zhou, Z. Du, Q. Guo, S. Liu, C. Liu, C. Wang, X. Zhou, L. Li, T. Chen, and Y. Chen, “Cambricon-S: Addressing irregularity in sparse neural networks through a cooperative software/hardware approach,” in Proc. Int. Symp. Microarchitecture , 2018, pp. 15–28
2018
Later among the works it cites.
J. Zhu, J. Jiang, X. Chen, and C.-Y. Tsui, “SparseNN: An energy-efficient neural network accelerator exploiting input and output sparsity,” in Proc. Design Automation & Test Europe Conf. , 2018, pp. 241–244
2018
Later among the works it cites.
T. Baltrusaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: a survey and taxonomy,” IEEE Trans. Pattern Analysis and Machine Intelligence , vol. 41, no. 2, p. 423–443, Feb. 2019
2019
Later among the works it cites.
S. Cao, L. Ma, W. Xiao, C. Zhang, Y. Liu, L. Zhang, L. Nie, and Z. Yang, “SeerNet: Predicting convolutional neural network feature-map sparsity through low-bit quantization,” in Proc. Conf. Computer Vision and Pattern Recognition , 2019, pp. 11 216–11 225
2019
Later among the works it cites.
C. Choy, J. Gwak, and S. Savarese, “4D spatio-temporal convnets: Minkowski convolutional neural networks,” in Proc. Conf. Computer Vision and Pattern Recognition , 2019, pp. 3075–3084
2019
Later among the works it cites.
X. Dai, H. Yin, and N. K. Jha, “Nest: A neural network synthesis tool based on a grow-and-prune paradigm,” in IEEE Trans. Computers , vol. 68, no. 10, Oct. 2019, pp. 1487–1497
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Frankle and M. Carbin, “The lottery ticket hypothesis: Training pruned neural networks,” in Proc. Int. Conf. Learning Representations , 2019
2019
Later among the works it cites.
A. Gondimalla, N. Chesnut, M. Thottethodi, and T. N. Vijaykumar, “Sparten: A sparse tensor accelerator for convolutional neural networks,” in Proc. Int. Symp. Microarchitecture , 2019, p. 151–165
2019
Later among the works it cites.
2019
Later among the works it cites.
W. Hua, Y. Zhou, C. De Sa, Z. Zhang, and G. E. Suh, “Boosting the performance of CNN accelerators with dynamic fine-grained channel gating,” in Proc. Int. Symp. Microarchitecture , 2019, p. 139–150
2019
Later among the works it cites.
W. Hua, Y. Zhou, C. M. De Sa, Z. Zhang, and G. E. Suh, “Channel gating neural networks,” in Proc. Conf. Neural Information Processing Systems , 2019, pp. 1884–1894
2019
Later among the works it cites.
Y. Kaya, S. Hong, and T. Dumitras, “Shallow-deep networks: Understanding and mitigating network overthinking,” in Proc. Int. Conf. Machine Learning , Jun 2019
2019
Later among the works it cites.
L. Liu, L. Deng, X. Hu, M. Zhu, G. Li, Y. Ding, and Y. Xie, “Dynamic sparse graph for efficient deep learning,” in Proc. Int. Conf. Learning Representations , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa et al. , “PyTorch: An imperative style, high-performance deep learning library,” in Proc. Conf. Neural Information Processing Systems , 2019, pp. 8026–8037
2019
Later among the works it cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI Blog , p. 9, 2019
2019
Later among the works it cites.
A. Ren, T. Zhang, S. Ye, J. Li, W. Xu, X. Qian, X. Lin, and Y. Wang, “ADMM-NN: An algorithm-hardware co-design framework of DNNs using alternating direction methods of multipliers,” in Proc. Int. Conf. Architectural Support for Programming Languages and Operating Systems , 2019, pp. 925–938
2019
Later among the works it cites.
J.-F. Zhang, C.-E. Lee, C. Liu, Y. S. Shao, S. W. Keckler, and Z. Zhang, “SNAP: A 1.67—21.55 TOPS/W sparse neural acceleration processor for unstructured sparse deep neural network inference in 16nm CMOS,” in Symp. VLSI Circuits , 2019, pp. C306–C307
2019
Later among the works it cites.
2019
Later among the works it cites.
H. Cai, C. Gan, and S. Han, “Once for all: Train one network and specialize it for efficient deployment,” in Proc. Int. Conf. Learning Representations , 2020
2020
Closest in time.
E. Qin, A. Samajdar, H. Kwon, V. Nadella, S. Srinivasan, D. Das, B. Kaul, and T. Krishna, “SIGMA: A sparse and irregular GEMM accelerator with flexible interconnects for DNN training,” in Proc. Int. Symp. High Performance Computer Architecture , 2020
2020
Closest in time.
2020
Closest in time.
W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” in Proc. Conf. Neural Information Processing Systems , 2016, pp. 2074–2082
2082
Closest in time.