Fetching the paper…
Reading the bibliography…
Significant memory and computational requirements of large deep neural networks restrict their application on edge devices.
Improved knowledge distillation via teacher assistant: Bridging the gap between student and teacher
Mirzadeh, S., Farajtabar, M., Li, A., Levine, N., Matsukawa, A., and Ghasemzadeh, H. (2020) · 1902
Earlier work this paper cites.
Improved knowledge distillation via teacher assistant: Bridging the gap between student and teacher
Mirzadeh, S.-I., Farajtabar, M., Li, A., and Ghasemzadeh, H. (2019) · 1902
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Sun, S., Cheng, Y., Gan, Z., and Liu, J. (2019) · 1908
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q. (2019) · 1909
Earlier work this paper cites.
Distilled embedding: non-linear embedding factorization using knowledge distillation
Lioutas, V., Rashid, A., Kumar, K., Haidar, M. A., and Rezagholizadeh, M. (2019) · 1910
Earlier work this paper cites.
Fully quantized transformer for improved translation
Prato, G., Charlaix, E., and Rezagholizadeh, M. (2019) · 1910
Earlier work this paper cites.
Yolo nano: a highly compact you only look once convolutional neural network for object detection
Wong, A., Famuori, M., Shafiee, M. J., Li, F., Chwyl, B., and Chung, J. (2019) · 1910
Earlier work this paper cites.
Fully quantizing a simplified transformer for end-to-end speech recognition
Bie, A., Venkitesh, B., Monteiro, J., Haidar, M., Rezagholizadeh, M., et al. (2019) · 1911
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Statistical learning theory
Vapnik, V. (1998) · 1998
Earlier work this paper cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Sun, Z., Yu, H., Song, X., Liu, R., Yang, Y., and Zhou, D. (2020) · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C. (2005) · 2005
Earlier work this paper cites.
Knowledge distillation: A survey
Gou, J., Yu, B., Maybank, S. J., and Tao, D. (2020) · 2006
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D. (2009) · 2009
Cited alongside, same era.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al. (2009) · 2009
Cited alongside, same era.
Neural networks for machine learning, coursera
Hinton, G. (2012) · 2012
Cited alongside, same era.
The winograd schema challenge
Levesque, H., Davis, E., and Morgenstern, L. (2012) · 2012
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C. (2013) · 2013
Cited alongside, same era.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. R. (2017) · 2017
Later among the works it cites.
Quora question pairs
Chen, Z., Zhang, H., Zhang, X., and Zhao, L. (2018) · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Later among the works it cites.
Furlanello, T., Lipton, Z. C., Tschannen, M., Itti, L., and Anandkumar, A. (2018) · 2018
Later among the works it cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distilling the Knowledge in a Neural Network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Cited alongside, same era.
Unifying distillation and privileged information
Lopez-Paz, D., Bottou, L., Schölkopf, B., and Vapnik, V. (2015) · 2015
Cited alongside, same era.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
Chan, W., Jaitly, N., Le, Q., and Vinyals, O. (2016) · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016) · 2016
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L. (2017) · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017) · 2017
Cited alongside, same era.
Later among the works it cites.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T. (2018) · 2018
Later among the works it cites.
Tensor decomposition for compressing recurrent neural network
Tjandra, A., Sakti, S., and Nakamura, S. (2018) · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2018) · 2018
Later among the works it cites.
Deep mutual learning
Zhang, Y., Xiang, T., Hospedales, T. M., and Lu, H. (2018) · 2018
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R. (2019) · 2019
Later among the works it cites.
BAM! born-again multi-task networks for natural language understanding
Clark, K., Luong, M.-T., Khandelwal, U., Manning, C. D., and Le, Q. V. (2019) · 2019
Later among the works it cites.
Streaming end-to-end speech recognition for mobile devices
He, Y., Sainath, T. N., Prabhavalkar, R., McGraw, I., Alvarez, R., Zhao, D., Rybach, D., Kannan, A., Wu, Y., Pang, R., et al. (2019) · 2019
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R. (2019) · 2019
Later among the works it cites.