Fetching the paper…
Reading the bibliography…
Enhancing small language models for real-life application deployment is a significant challenge facing the research community.
Simonyan, K., Vedaldi, A., Zisserman, A.: Deep inside convolutional networks: visualising image classification models and saliency maps. In: Proceedings of the International Conference on Learning Representations (ICLR) (2014)
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Li, J., Chen, X., Hovy, E., Jurafsky, D.: Visualizing and understanding neural models in nlp. In: Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics (2016)
2016
Earlier work this paper cites.
Sundararajan, M., Taly, A., Yan, Q.: Axiomatic attribution for deep networks. In: International Conference on Machine Learning. pp. 3319–3328. PMLR (2017)
2017
Earlier work this paper cites.
Camburu, O.M., Rocktäschel, T., Lukasiewicz, T., Blunsom, P.: e-snli: Natural language inference with natural language explanations. In: Advances in Neural Information Processing Systems. vol. 31 (2018)
2018
Earlier work this paper cites.
Mishra, A., Marr, D.: Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy. In: International Conference on Learning Representations (2018)
2018
Earlier work this paper cites.
Tan, X., Ren, Y., He, D., Qin, T., Zhao, Z., Liu, T.Y.: Multilingual neural machine translation with knowledge distillation. In: International Conference on Learning Representations (2018)
2018
Earlier work this paper cites.
Cho, J.H., Hariharan, B.: On the efficacy of knowledge distillation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4794–4802 (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Rajani, N.F., McCann, B., Xiong, C., Socher, R.: Explain yourself! leveraging language models for commonsense reasoning. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. pp. 4932–4942 (2019)
2019
Earlier work this paper cites.
Talmor, A., Herzig, J., Lourie, N., Berant, J.: Commonsenseqa: A question answering challenge targeting commonsense knowledge. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). pp. 4149–4158 (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Brunner, G., Liu, Y., Pascual, D., Richter, O., Ciaramita, M., Wattenhofer, R.: On identifiability in transformers. In: 8th International Conference on Learning Representations (ICLR 2020)(virtual) (2020)
2020
Earlier work this paper cites.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M.e.a.: Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research 21
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Cited alongside, same era.
Patel, A., Bhattamishra, S., Goyal, N.: Are nlp models really able to solve simple math word problems? In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 2080–2094 (2021)
2021
Cited alongside, same era.
2022
Later among the works it cites.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E.e.a.: Chain-of-thought prompting elicits reasoning in large language models. In: Advances in Neural Information Processing Systems. vol. 35, pp. 24824–24837 (2022)
2022
Later among the works it cites.
West, P., Bhagavatula, C., Hessel, J., Hwang, J., Jiang, L., Le Bras, R.e.a.: Symbolic knowledge distillation: from general language models to commonsense models. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 4602–4625 (2022)
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A.W., Lester, B.e.a.: Finetuned language models are zero-shot learners. In: International Conference on Learning Representations (2021)
2021
Cited alongside, same era.
Beyer, L., Zhai, X., Royer, A., Markeeva, L., Anil, R., Kolesnikov, A.: Knowledge distillation: A good teacher is patient and consistent. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10925–10934 (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Hase, P., Bansal, M.: When can models learn from explanations? a formal framework for understanding the roles of explanation data. In: LNLS 2022. vol. 29 (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Sanh, V., Webson, A., Raffel, C., Bach, S.H., Sutawika, L., Alyafeai, Z.e.a.: Multitask prompted training enables zero-shot task generalization. In: ICLR 2022-Tenth International Conference on Learning Representations (2022)
2022
Cited alongside, same era.
2023
Later among the works it cites.
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A.e.a.: Palm: Scaling language modeling with pathways. Journal of Machine Learning Research 24
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.