Fetching the paper…
Reading the bibliography…
Fine-tuning a pre-trained language model (PLM) emerges as the predominant strategy in many natural language processing applications.
H. Robbins and S. Monro, “A stochastic approximation method,” The annals of mathematical statistics , pp. 400–407, 1951
1951
Earlier work this paper cites.
J. M. Hammersley and K. W. Morton, “Poor man’s monte carlo,” Journal of the Royal Statistical Society: Series B (Methodological) , vol. 16, no. 1, pp. 23–38, 1954
1954
Earlier work this paper cites.
K. Fukushima, “Cognitron: A self-organizing multilayered neural network,” Biological cybernetics , vol. 20, no. 3, pp. 121–136, 1975
1975
Earlier work this paper cites.
S. Lloyd, “Least squares quantization in pcm,” IEEE transactions on information theory , vol. 28, no. 2, pp. 129–137, 1982
1982
Earlier work this paper cites.
T. Wei and J. He, “Comprehensive fair meta-learned recommender system,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2022, pp. 1989–1999
1999
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
W. B. Dolan and C. Brockett, “Automatically constructing a corpus of sentential paraphrases,” in Proceedings of the Third International Workshop on Paraphrasing (IWP2005) , 2005. [Online]. Available: https://aclanthology.org/I05-5002
2005
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization , 2005, pp. 65–72
2005
Earlier work this paper cites.
M. Snover, B. Dorr, R. Schwartz, L. Micciulla, and J. Makhoul, “A study of translation edit rate with targeted human annotation,” in Proceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers , 2006, pp. 223–231
2006
Earlier work this paper cites.
N. Ailon and B. Chazelle, “The fast johnson–lindenstrauss transform and approximate nearest neighbors,” SIAM Journal on computing , vol. 39, no. 1, pp. 302–322, 2009
2009
Earlier work this paper cites.
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts, “Recursive deep models for semantic compositionality over a sentiment treebank,” in Proceedings of the 2013 conference on empirical methods in natural language processing , 2013, pp. 1631–1642
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
E. L. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus, “Exploiting linear structure within convolutional networks for efficient evaluation,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
D. P. Woodruff et al. , “Sketching as a tool for numerical linear algebra,” Foundations and Trends® in Theoretical Computer Science , vol. 10, no. 1–2, pp. 1–157, 2014
2014
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , Y. Bengio and Y. LeCun, Eds., 2015
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Z. Liu, J. Li, Z. Shen, G. Huang, S. Yan, and C. Zhang, “Learning efficient convolutional networks through network slimming,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2736–2744
2017
Earlier work this paper cites.
C. Gardent, A. Shimorina, S. Narayan, and L. Perez-Beltrachini, “Creating training corpora for nlg micro-planning,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics (ACL), Aug. 2017, pp. 179–188, 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017 ; Conference date: 30-07-2017 Through 04-08-2017
2017
Earlier work this paper cites.
J. Howard and S. Ruder, “Universal language model fine-tuning for text classification,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 328–339
2018
Earlier work this paper cites.
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in Proceedings of NAACL-HLT , 2018, pp. 2227–2237
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Jacot, F. Gabriel, and C. Hongler, “Neural tangent kernel: Convergence and generalization in neural networks,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
L. Balles and P. Hennig, “Dissecting adam: The sign, magnitude and variance of stochastic gradients,” in International Conference on Machine Learning . PMLR, 2018, pp. 404–413
2018
Earlier work this paper cites.
L. Dong, S. Xu, and B. Xu, “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2018, pp. 5884–5888
2018
Earlier work this paper cites.
N. Lee, T. Ajanthan, and P. Torr, “SNIP: SINGLE-SHOT NETWORK PRUNING BASED ON CONNECTION SENSITIVITY,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
S. Arora, S. S. Du, W. Hu, Z. Li, R. R. Salakhutdinov, and R. Wang, “On exact computation with an infinitely wide neural net,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
G. Peyré, M. Cuturi et al. , “Computational optimal transport: With applications to data science,” Foundations and Trends® in Machine Learning , vol. 11, no. 5-6, pp. 355–607, 2019
2019
Cited alongside, same era.
H. He, J. Wang, Z. Zhang, and F. Wu, “Compressing deep graph neural networks via adversarial knowledge distillation,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , ser. KDD ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 534–544. [Online]. Available: https://doi.org/10.1145/3534678.3539315
2022
Later among the works it cites.
T. Dao, B. Chen, N. S. Sohoni, A. Desai, M. Poli, J. Grogan, A. Liu, A. Rao, A. Rudra, and C. Ré, “Monarch: Expressive structured matrices for efficient and accurate training,” in International Conference on Machine Learning . PMLR, 2022, pp. 4690–4721
2022
Later among the works it cites.
F. Xue, X. He, X. Ren, Y. Lou, and Y. You, “One student knows all experts know: From sparse to dense,” 2022
2022
Later among the works it cites.
C. Liu, C. Lou, R. Wang, A. Y. Xi, L. Shen, and J. Yan, “Deep neural network fusion via graph matching with applications to model ensemble and federated learning,” in Proceedings of the 39th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, Eds., vol. 162. PMLR, 17–23 Jul 2022, pp. 13 857–13 869
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Warstadt, A. Singh, and S. R. Bowman, “Neural network acceptability judgments,” Transactions of the Association for Computational Linguistics , vol. 7, pp. 625–641, 2019. [Online]. Available: https://aclanthology.org/Q19-1040
2019
Cited alongside, same era.
M. Kale and A. Rastogi, “Text-to-text pre-training for data-to-text tasks,” in Proceedings of the 13th International Conference on Natural Language Generation . Dublin, Ireland: Association for Computational Linguistics, Dec. 2020, pp. 97–102
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,” 2020
2020
Cited alongside, same era.
S. Kang, J. Hwang, W. Kweon, and H. Yu, “De-rrd: A knowledge distillation framework for recommender system,” in Proceedings of the 29th ACM International Conference on Information & Knowledge Management , ser. CIKM ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 605–614. [Online]. Available: https://doi.org/10.1145/3340531.3412005
2020
Cited alongside, same era.
2022
Later among the works it cites.
S. Lohit and M. Jones, “Model compression using optimal transport,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2022, pp. 2764–2773
2022
Later among the works it cites.
Y. Chen, Q. Zeng, D. Hakkani-Tur, D. Jin, H. Ji, and Y. Yang, “Sketching as a tool for understanding and accelerating self-attention for long sequences,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . Seattle, United States: Association for Computational Linguistics, Jul. 2022, pp. 5187–5199
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
L. Gu, Y. Du, Y. Zhang, D. Xie, S. Pu, R. C. Qiu, and Z. Liao, “”lossless” compression of deep neural networks: A high-dimensional neural tangent kernel approach,” in Advances in Neural Information Processing Systems , A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, Eds., 2022
2022
Later among the works it cites.
Y. Chen, D. Hazarika, M. Namazifar, Y. Liu, D. Jin, and D. Hakkani-Tur, “Inducer-tuning: Connecting prefix-tuning and adapter-tuning,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2022
2022
Later among the works it cites.
B. Wang, Y. Ren, L. Shang, X. Jiang, and Q. Liu, “Exploring extreme parameter compression for pre-trained language models,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
B. Yuan, C. R. Wolfe, C. Dun, Y. Tang, A. Kyrillidis, and C. Jermaine, “Distributed learning of fully connected neural networks using independent subnet training,” Proceedings of the VLDB Endowment , vol. 15, no. 8, pp. 1581–1590, 2022
2022
Later among the works it cites.
K. Sreenivasan, J. yong Sohn, L. Yang, M. Grinde, A. Nagle, H. Wang, E. Xing, K. Lee, and D. Papailiopoulos, “Rare gems: Finding lottery tickets at initialization,” in Advances in Neural Information Processing Systems , A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, Eds., 2022
2022
Later among the works it cites.
Y. Chen, D. Hazarika, M. Namazifar, Y. Liu, D. Jin, and D. Hakkani-Tur, “Empowering parameter-efficient transfer learning by recognizing the kernel structure in self-attention,” in Findings of the Association for Computational Linguistics: NAACL 2022 . Association for Computational Linguistics, 2022
2022
Later among the works it cites.
S. Geng, S. Liu, Z. Fu, Y. Ge, and Y. Zhang, “Recommendation as language processing (rlp): A unified pretrain, personalized prompt & predict paradigm (p5),” in Proceedings of the 16th ACM Conference on Recommender Systems , 2022, pp. 299–315
2022
Later among the works it cites.
T. Wei, Y. You, T. Chen, Y. Shen, J. He, and Z. Wang, “Augmentations in hypergraph contrastive learning: Fabricated and generative,” in Advances in Neural Information Processing Systems , 2022
2022
Later among the works it cites.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023
2023
Closest in time.
S. He, R.-Z. Fan, L. Ding, L. Shen, T. Zhou, and D. Tao, “Merging experts into one: Improving computational efficiency of mixture of experts,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 14 685–14 691
2023
Closest in time.
2023
Closest in time.
Y. Wang, D. Li, and R. Sun, “NTK-SAP: Improving neural network pruning by aligning training dynamics,” in The Eleventh International Conference on Learning Representations , 2023
2023
Closest in time.
G. He, J. Chen, and J. Zhu, “Preserving pre-trained features helps calibrate fine-tuned language models,” in The Eleventh International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=NI7StoWHJPT
2023
Closest in time.
J. Mukhoti, Y. Gal, P. H. S. Torr, and P. K. Dokania, “Fine-tuning can cripple your foundation model; preserving features may be the solution,” 2023
2023
Closest in time.
2023
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mixtral of experts,” 2024
2024
Closest in time.
2024
Closest in time.
G. Stoica, D. Bolya, J. Bjorner, P. Ramesh, T. Hearn, and J. Hoffman, “Zipit! merging models from different tasks without training,” 2024
2024
Closest in time.
2024
Closest in time.
L. Iurada, M. Ciccone, and T. Tommasi, “Finding lottery tickets in vision models via data-driven spectral foresight pruning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2024, pp. 16 142–16 151
2024
Closest in time.