Fetching the paper…
Reading the bibliography…
Large language models(LLMs) containing tens of billions of parameters (or even more) have demonstrated impressive capabilities in various NLP tasks.
Y. LeCun, J. Denker, and S. Solla, “Optimal brain damage,” Advances in neural information processing systems
1989
Earlier work this paper cites.
M. Marcus, B. Santorini, and M. A. Marcinkiewicz, “Building a large annotated corpus of english: The penn treebank,” 1993
1993
Earlier work this paper cites.
J. H. Friedman, “Greedy function approximation: a gradient boosting machine,” Annals of statistics
2001
Earlier work this paper cites.
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler, “Aligning books and movies: Towards story-like visual explanations by watching movies and reading books,” in Proceedings of the IEEE international conference on computer vision
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
R. Luo, F. Tian, T. Qin, E. Chen, and T.-Y. Liu, “Neural architecture optimization,” Advances in neural information processing systems
2018
Earlier work this paper cites.
P. Molchanov, A. Mallya, S. Tyree, I. Frosio, and J. Kautz, “Importance estimation for neural network pruning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Hu, X. Wang, L. Li, and Q. Gu, “Improving one-shot nas with shrinking-and-expanding supernet,” Pattern Recognition
2021
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al
2022
Earlier work this paper cites.
L. Li and Z. Jin, “Shadow knowledge distillation: Bridging offline and online knowledge transfer,” Advances in Neural Information Processing Systems
2022
Cited alongside, same era.
L. Li, L. Shiuan-Ni, Y. Yang, and Z. Jin, “Boosting online feature transfer via separable feature fusion.,” in IJCNN
2022
Cited alongside, same era.
L. Li, “Self-regulated feature learning via teacher-free feature distillation,” in ECCV 2022
2022
Cited alongside, same era.
L. Li, L. Shiuan-Ni, Y. Yang, and Z. Jin, “Teacher-free distillation via regularizing intermediate representation,” in IJCNN
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
K. Chen, L. Yang, Y. Chen, K. Chen, Y. Xu, and L. Li, “Gp-nas-ensemble: a model for the nas performance prediction,” in CVPRW
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
L. Li, P. Dong, Z. Wei, and Y. Yang, “Automated knowledge distillation via monte carlo tree search,” in ICCV
2023
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
E. Frantar and D. Alistarh, “Sparsegpt: Massive language models can be accurately pruned in one-shot,” 2023
2023
Closest in time.
P. Dong, L. Li, and Z. Wei, “Diswot: Student architecture search for distillation without training,” in CVPR
2023
Closest in time.
P. Dong, X. Niu, L. Li, Z. Tian, X. Wang, Z. Wei, H. Pan, and D. Li, “Rd-nas: Enhancing one-shot supernet ranking ability via ranking distillation from zero-cost proxies,” in ICASSP)
2023
Closest in time.
2023
Closest in time.