Fetching the paper…
Reading the bibliography…
Network quantization has gained increasing attention with the rapid growth of large pre-trained language models~(PLMs).
R. J. Williams and D. Zipser, “A learning algorithm for continually running fully recurrent neural networks,” Neural computation , vol. 1, no. 2, pp. 270–280, 1989
1989
Earlier work this paper cites.
J. Dean and S. Ghemawat, “Mapreduce: simplified data processing on large clusters,” in Communications of the ACM , vol. 51, no. 1, 2008, pp. 107–113
2008
Earlier work this paper cites.
Q. Ho, J. Cipar, H. Cui, J. K. Kim, S. Lee, P. B. Gibbons, G. A. Gibson, G. R. Ganger, and E. P. Xing, “More effective distributed ml via a stale synchronous parallel parameter server,” in Advances in Neural Information Processing Systems , 2013, p. 1223
2013
Earlier work this paper cites.
M. Li, D. G. Andersen, A. J. Smola, and K. Yu, “Communication efficient distributed machine learning with the parameter server,” in Advances in Neural Information Processing Systems , vol. 27, 2014, pp. 19–27
2014
Earlier work this paper cites.
M. Courbariaux, Y. Bengio, and J.-P. David, “Binaryconnect: Training deep neural networks with binary weights during propagations,” in Advances in neural information processing systems , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “Fitnets: Hints for thin deep nets,” in International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
F. Li, B. Zhang, and B. Liu, “Ternary weight networks,” Preprint arXiv:1605.04711, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
L. Hou, Q. Yao, and J. T. Kwok, “Loss-aware binarization of deep networks,” in International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
Y. He, X. Zhang, and J. Sun, “Channel pruning for accelerating very deep neural networks,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 1398–1406
2017
Earlier work this paper cites.
L. Hou and J. T. Kwok, “Loss-aware weight quantization of deep networks,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
J.-H. Luo, H. Zhang, H.-Y. Zhou, C.-W. Xie, J. Wu, and W. Lin, “Thinet: pruning cnn filters for a thinner net,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 10, pp. 2525–2538, 2018
2018
Earlier work this paper cites.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, M. X. Chen, D. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu et al. , “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” in Advances in neural information processing systems , 2018
2018
Earlier work this paper cites.
Z. Jia, S. Lin, C. R. Qi, and A. Aiken, “Exploring hidden dimensions in accelerating convolutional neural networks,” in International Conference on Machine Learning , 2018, pp. 2274–2283
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in North American Chapter of the Association for Computational Linguistics , 2019
2019
Earlier work this paper cites.
P. Michel, O. Levy, and G. Neubig, “Are sixteen heads really better than one?” in Advances in Neural Information Processing Systems , 2019
2019
Earlier work this paper cites.
A. Fan, E. Grave, and A. Joulin, “Reducing transformer depth on demand with structured dropout,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Sun, Y. Cheng, Z. Gan, and J. Liu, “Patient knowledge distillation for bert model compression,” in Conference on Empirical Methods in Natural Language Processing , 2019
2019
Earlier work this paper cites.
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and Ł. Kaiser, “Universal transformers,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Nagel, M. v. Baalen, T. Blankevoort, and M. Welling, “Data-free quantization through weight equalization and bias correction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 1325–1334
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
S. K. Esser, J. L. McKinstry, D. Bablani, R. Appuswamy, and D. S. Modha, “Learned step size quantization,” in International Conference on Learning Representations , 2019
2019
Cited alongside, same era.
R. Zhao, Y. Hu, J. Dotzel, C. De Sa, and Z. Zhang, “Improving neural network quantization without retraining using outlier channel splitting,” in Proceedings of the International Conference on Machine Learning , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Liu, C. Shu, J. Wang, and C. Shen, “Structured knowledge distillation for dense prediction,” IEEE transactions on pattern analysis and machine intelligence , 2020
2020
Later among the works it cites.
D. Chen, Y. Li, M. Qiu, Z. Wang, B. Li, B. Ding, H. Deng, J. Huang, W. Lin, and J. Zhou, “Adabert: Task-adaptive bert compression with differentiable neural architecture search,” in International joint conference on artificial intelligence , 2020
2020
Later among the works it cites.
J. Tarnawski, A. Phanishayee, N. R. Devanur, D. Mahajan, and F. N. Paravecino, “Efficient algorithms for device placement of dnn graph operators,” in Advances in Neural Information Processing Systems , 2020
2020
Later among the works it cites.
J. H. Park, G. Yun, M. Y. Chang, N. T. Nguyen, S. Lee, J. Choi, S. H. Noh, and Y.-r. Choi, “Hetpipe: Enabling large { \{ DNN } \} training on (whimpy) heterogeneous { \{ GPU } \} clusters through integration of pipelined model parallelism and data parallelism,” in 2020 { \{ USENIX } \} Annual Technical Conference ( { \{ USENIX } \} { \{ ATC } \} 20) , 2020, pp. 307–321
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Liu, K. Simonyan, and Y. Yang, “Darts: Differentiable architecture search,” in Proceedings of the International Conference of Representation Learning , 2019
2019
Cited alongside, same era.
D. So, Q. Le, and C. Liang, “The evolved transformer,” in Proceedings of the International Conference on Machine Learning . PMLR, 2019, pp. 5877–5886
2019
Cited alongside, same era.
G. Bender, “Understanding and simplifying one-shot architecture search,” in Proceedings of the Proceedings of the International Conference on Machine Learning , 2019, pp. 549–558
2019
Cited alongside, same era.
D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia, “Pipedream: generalized pipeline parallelism for dnn training,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles , 2019, pp. 1–15
2019
Cited alongside, same era.
2019
Cited alongside, same era.
S. Chen, W. Wang, and S. J. Pan, “Deep neural network quantization via layer-wise optimization using limited training data,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2019, pp. 3329–3336
2019
Cited alongside, same era.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in Advances in Neural Information Processing Systems , 2020
2020
Cited alongside, same era.
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu, “Tinybert: Distilling bert for natural language understanding,” in Findings of Empirical Methods in Natural Language Processing , 2020
2020
Cited alongside, same era.
2020
Later among the works it cites.
L. Song, F. Chen, Y. Zhuo, X. Qian, H. Li, and Y. Chen, “Accpar: Tensor partitioning for heterogeneous deep learning accelerators,” in IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2020, pp. 342–355
2020
Later among the works it cites.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: Memory optimizations toward training trillion parameter models,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–16
2020
Later among the works it cites.
J. Fang, A. Shafiee, H. Abdel-Aziz, D. Thorsley, G. Georgiadis, and J. H. Hassoun, “Post-training piecewise linear quantization for deep neural networks,” in European Conference on Computer Vision , 2020, pp. 69–86
2020
Later among the works it cites.
P. Wang, Q. Chen, X. He, and J. Cheng, “Towards accurate post-training network quantization via bit-split and stitching,” in Proceedings of the International Conference on Machine Learning , 2020, pp. 9847–9856
2020
Later among the works it cites.
D. Zhou, M. Ye, C. Chen, T. Meng, M. Tan, X. Song, Q. Le, Q. Liu, and D. Schuurmans, “Go wide, then narrow: Efficient training of deep thin networks,” in Proceedings of the International Conference on Machine Learning , 2020, pp. 11 546–11 555
2020
Later among the works it cites.
H. Bai, J. Wu, I. King, and M. Lyu, “Few shot network compression via cross distillation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 04, 2020, pp. 3203–3210
2020
Later among the works it cites.
2020
Later among the works it cites.
Z. Huang, L. Hou, L. Shang, X. Jiang, X. Chen, and Q. Liu, “Ghostbert: Generate more features with cheap operations for bert,” in Annual Meeting of the Association for Computational Linguistics , 2021
2021
Closest in time.
H. Bai, W. Zhang, L. Hou, L. Shang, J. Jin, X. Jiang, Q. Liu, M. Lyu, and I. King, “Binarybert: Pushing the limit of bert quantization,” in Annual Meeting of the Association for Computational Linguistics , 2021
2021
Closest in time.
Y. Li, R. Gong, X. Tan, Y. Yang, P. Hu, Q. Zhang, F. Yu, W. Wang, and S. Gu, “Brecq: Pushing the limit of post-training quantization by block reconstruction,” in International Conference on Learning Representations , 2021
2021
Closest in time.
B. Zhuang, M. Tan, J. Liu, L. Liu, I. Reid, and C. Shen, “Effective training of convolutional neural networks with low-bitwidth weights and activations,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2021
2021
Closest in time.
S. Young, Z. Wang, D. Taubman, and B. Girod, “Transform quantization for cnn compression,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2021
2021
Closest in time.
I. Hubara, Y. Nahshan, Y. Hanani, R. Banner, and D. Soudry, “Improving post training neural quantization: Layer-wise calibration and integer programming,” in Proceedings of the International Conference on Machine Learning , 2021
2021
Closest in time.
J. Liu, B. Zhuang, Z. Zhuang, Y. Guo, J. Huang, J. Zhu, and M. Tan, “Discrimination-aware network pruning for deep model compression,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2021
2021
Closest in time.
G.-H. Wang, Y. Ge, and J. Wu, “Distilling knowledge by mimicking features,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2021
2021
Closest in time.
L. Zhang, C. Bao, and K. Ma, “Self-distillation: Towards efficient and compact neural networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2021
2021
Closest in time.
J. Xu, X. Tan, R. Luo, K. Song, J. Li, T. Qin, and T.-Y. Liu, “Nas-bert: Task-agnostic and adaptive-size bert compression with neural architecture search,” in Proceedings of the ACM SIGKDD international conference on knowledge discovery and data mining , 2021
2021
Closest in time.
Y. Yin, C. Chen, L. Shang, X. Jiang, X. Chen, and Q. Liu, “Autotinybert: Automatic hyper-parameter optimization for efficient pre-trained language models,” in Annual Meeting of the Association for Computational Linguistics , 2021
2021
Closest in time.
S. Fan, Y. Rong, C. Meng, Z. Cao, S. Wang, Z. Zheng, C. Wu, G. Long, J. Yang, L. Xia et al. , “Dapple: A pipelined data parallel approach for training large models,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2021, pp. 431–445
2021
Closest in time.
2021
Closest in time.