Fetching the paper…
Reading the bibliography…
Fine-tuning on task-specific datasets is a widely-embraced paradigm of harnessing the powerful capability of pretrained LLMs for various downstream tasks.
J. C. Spall, “Multivariate stochastic approximation using a simultaneous perturbation gradient approximation,” IEEE transactions on automatic control , vol. 37, no. 3, pp. 332–341, 1992
1992
Earlier work this paper cites.
E. M. Voorhees and D. M. Tice, “Building a question answering test collection,” in Proceedings of the 23rd annual international ACM SIGIR conference on Research and development in information retrieval , 2000, pp. 200–207
2000
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out , 2004, pp. 74–81
2004
Earlier work this paper cites.
P. Koehn, “Europarl: A parallel corpus for statistical machine translation,” in Proceedings of machine translation summit x: papers , 2005, pp. 79–86
2005
Earlier work this paper cites.
C. Dwork, “Differential privacy,” in International colloquium on automata, languages, and programming . Springer, 2006, pp. 1–12
2006
Earlier work this paper cites.
C. Dwork, F. McSherry, K. Nissim, and A. Smith, “Calibrating noise to sensitivity in private data analysis,” in Theory of cryptography conference . Springer, 2006, pp. 265–284
2006
Earlier work this paper cites.
S. Ghadimi and G. Lan, “Stochastic first-and zeroth-order methods for nonconvex stochastic programming,” SIAM Journal on Optimization , vol. 23, no. 4, pp. 2341–2368, 2013
2013
Earlier work this paper cites.
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts, “Recursive deep models for semantic compositionality over a sentiment treebank,” in Proceedings of the 2013 conference on empirical methods in natural language processing , 2013, pp. 1631–1642
2013
Earlier work this paper cites.
M. Bun, J. Ullman, and S. Vadhan, “Fingerprinting codes and the price of approximate differential privacy,” in Proceedings of the forty-sixth annual ACM symposium on Theory of computing , 2014, pp. 1–10
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang, “Deep learning with differential privacy,” in Proceedings of the 2016 ACM SIGSAC conference on computer and communications security , 2016, pp. 308–318
2016
Earlier work this paper cites.
H. Karimi, J. Nutini, and M. Schmidt, “Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition,” in Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2016, Riva del Garda, Italy, September 19-23, 2016, Proceedings, Part I 16 . Springer, 2016, pp. 795–811
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y. Nesterov and V. Spokoiny, “Random gradient-free minimization of convex functions,” Foundations of Computational Mathematics , vol. 17, pp. 527–566, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
B. Balle, G. Barthe, and M. Gaboardi, “Privacy amplification by subsampling: Tight analyses via couplings and divergences,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Khashabi, S. Chaturvedi, M. Roth, S. Upadhyay, and D. Roth, “Looking beyond the surface: A challenge set for reading comprehension over multiple sentences,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , 2018, pp. 252–262
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
C. Sun, X. Qiu, Y. Xu, and X. Huang, “How to fine-tune bert for text classification?” in Chinese Computational Linguistics: 18th China National Conference, CCL 2019, Kunming, China, October 18–20, 2019, Proceedings 18 . Springer, 2019, pp. 194–206
2019
Cited alongside, same era.
R. Bassily, C. Guzmán, and M. Menart, “Differentially private stochastic optimization: New results in convex and non-convex settings,” Advances in Neural Information Processing Systems , vol. 34, pp. 9317–9329, 2021
2021
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Carlini, C. Liu, Ú. Erlingsson, J. Kos, and D. Song, “The secret sharer: Evaluating and testing unintended memorization in neural networks,” in 28th USENIX Security Symposium (USENIX Security 19) , 2019, pp. 267–284
2019
Cited alongside, same era.
Z. Yuan, Y. Yan, R. Jin, and T. Yang, “Stagewise training accelerates convergence of testing error over sgd,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “Superglue: A stickier benchmark for general-purpose language understanding systems,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
M.-C. De Marneffe, M. Simons, and J. Tonhauser, “The commitmentbank: Investigating projection in naturally occurring discourse,” in proceedings of Sinn und Bedeutung , vol. 23, no. 2, 2019, pp. 107–124
2019
Cited alongside, same era.
2019
Cited alongside, same era.
M. J. Wainwright, High-dimensional statistics: A non-asymptotic viewpoint . Cambridge university press, 2019, vol. 48
2019
Cited alongside, same era.
2020
Cited alongside, same era.
F. Mireshghallah, A. Backurs, H. A. Inan, L. Wutschitz, and J. Kulkarni, “Differentially private model compression,” Advances in Neural Information Processing Systems , vol. 35, pp. 29 468–29 483, 2022
2022
Later among the works it cites.
W. Shi, H. Gao, and B. Gu, “Gradient-free method for heavily constrained nonconvex optimization,” in International Conference on Machine Learning . PMLR, 2022, pp. 19 935–19 955
2022
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Feng, L. T. Yang, B. Ren, D. Zou, M. Dong, and S. Zhang, “Tensor recurrent neural network with differential privacy,” IEEE Transactions on Computers , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Arora, R. Bassily, T. González, C. A. Guzmán, M. Menart, and E. Ullah, “Faster rates of convergence to stationary points in differentially private optimization,” in International Conference on Machine Learning . PMLR, 2023, pp. 1060–1092
2023
Later among the works it cites.
L. Zhang, B. Li, K. K. Thekumparampil, S. Oh, and N. He, “Dpzero: Private fine-tuning of language models without backpropagation,” in Forty-first International Conference on Machine Learning , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.