Fetching the paper…
Reading the bibliography…
Tool-augmented large language models (LLMs) are often trained on datasets of query-response pairs, which embed the ability to use tools or APIs directly into the parametric knowledge of LLMs.
Likelihood ratio tests for model selection and non-nested hypotheses
Vuong, Q. H · 1989
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Eternal sunshine of the spotless net: Selective forgetting in deep networks
Golatkar, A., Achille, A., and Soatto, S · 2020
Earlier work this paper cites.
Amnesiac machine learning
Graves, L., Nagisetty, V., and Ganesh, V · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Machine unlearning via algorithmic stability
Ullah, E., Mai, T., Rao, A., Rossi, R. A., and Arora, R · 2021
Earlier work this paper cites.
Membership inference attacks from first principles
Carlini, N., Chien, S., Nasr, M., Song, S., Terzis, A., and Tramèr, F · 2022
Earlier work this paper cites.
Gpt3.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Towards a unified view of parameter-efficient transfer learning
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Talm: Tool augmented language models
Parisi, A., Zhao, Y., and Fiedel, N · 2022
Earlier work this paper cites.
Adversarial unlearning: Reducing confidence along adversarial directions
Setlur, A., Eysenbach, B., Smith, V., and Levine, S · 2022
Earlier work this paper cites.
GNNDelete: A general strategy for unlearning in graph neural networks
Cheng, J., Dasoulas, G., He, H., Agarwal, C., and Zitnik, M · 2023
Earlier work this paper cites.
Zero-shot machine unlearning
Chundawat, V. S., Tarun, A. K., Mandal, M., and Kankanhalli, M · 2023
Earlier work this paper cites.
Who’s harry potter? approximate unlearning in llms, 2023
Eldan, R. and Russinovich, M · 2023
Earlier work this paper cites.
Erasing concepts from diffusion models
Gandikota, R., Materzyńska, J., Fiotto-Kaufman, J., and Bau, D · 2023
Earlier work this paper cites.
Pal: program-aided language models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G · 2023
Earlier work this paper cites.
Factual or contextual? disentangling error types in entity description generation
Goyal, N., Nenkova, A., and Daumé III, H · 2023
Earlier work this paper cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Earlier work this paper cites.
Knowledge unlearning for mitigating privacy risks in language models
Jang, J., Yoon, D., Yang, S., Cha, S., Lee, M., Logeswaran, L., and Seo, M · 2023
Earlier work this paper cites.
Model sparsity can simplify machine unlearning
Jia, J., Liu, J., Ram, P., Yao, Y., Liu, G., Liu, Y., Sharma, P., and Liu, S · 2023
Earlier work this paper cites.
Preserving privacy through dememorization: An unlearning technique for mitigating memorization risks in language models
Kassem, A., Mahmoud, O., and Saad, S · 2023
Cited alongside, same era.
Ablating concepts in text-to-image diffusion models
Kumari, N., Zhang, B., Wang, S.-Y., Shechtman, E., Zhang, R., and Zhu, J.-Y · 2023
Cited alongside, same era.
Towards unbounded machine unlearning
Kurmanji, M., Triantafillou, P., Hayes, J., and Triantafillou, E · 2023
Cited alongside, same era.
API-bank: A comprehensive benchmark for tool-augmented LLMs
Li, M., Zhao, Y., Yu, B., Song, F., Li, H., Yu, H., Li, Z., Huang, F., and Li, Y · 2023
Cited alongside, same era.
Muter: Machine unlearning on adversarially trained models
Liu, J., Xue, M., Lou, J., Zhang, X., Xiong, L., and Qin, Z · 2023
Cited alongside, same era.
Evaluating the Ripple Effects of Knowledge Editing in Language Models
Cohen, R., Biran, E., Yoran, O., Globerson, A., and Geva, M · 2024
Later among the works it cites.
Salun: Empowering machine unlearning via gradient-based weight saliency in both image classification and generation
Fan, C., Liu, J., Zhang, Y., Wong, E., Wei, D., and Liu, S · 2024
Later among the works it cites.
Challenging forgets: Unveiling the worst-case forget sets in machine unlearning
Fan, C., Liu, J., Hero, A., and Liu, S · 2024
Later among the works it cites.
Model editing can hurt general abilities of large language models
Gu, J.-C., Xu, H.-X., Ma, J.-Y., Lu, P., Ling, Z.-H., Chang, K.-W., and Peng, N · 2024
Later among the works it cites.
Learn to unlearn for deep neural networks: Minimizing unlearning interference with gradient projection
Hoang, T., Rana, S., Gupta, S., and Venkatesh, S · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Cited alongside, same era.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al · 2023
Cited alongside, same era.
Exploring the impact of model scaling on parameter-efficient tuning
Su, Y., Chan, C.-M., Cheng, J., Qin, Y., Lin, Y., Hu, S., Yang, Z., Ding, N., Sun, X., Xie, G., Liu, Z., and Sun, M · 2023
Cited alongside, same era.
Challenging BIG-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q., Chi, E., Zhou, D., and Wei, J · 2023
Cited alongside, same era.
Toolalpaca: Generalized tool learning for language models with 3000 simulated cases
Tang, Q., Deng, Z., Lin, H., Han, X., Liang, Q., Cao, B., and Sun, L · 2023
Cited alongside, same era.
Reversing the forget-retain objectives: An efficient llm unlearning framework from logit difference
Ji, J., Liu, Y., Zhang, Y., Liu, G., Kompella, R. R., Liu, S., and Chang, S · 2024
Later among the works it cites.
SOUL: Unlocking the power of second-order optimization for LLM unlearning
Jia, J., Zhang, Y., Zhang, Y., Liu, J., Runwal, B., Diffenderfer, J., Kailkhura, B., and Liu, S · 2024
Later among the works it cites.
Rwku: Benchmarking real-world knowledge unlearning for large language models
Jin, Z., Cao, P., Wang, C., He, Z., Yuan, H., Li, J., Chen, Y., Liu, K., and Zhao, J · 2024
Later among the works it cites.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
Liu, H., Li, Z., Hall, D. L. W., Liang, P., and Ma, T · 2024
Later among the works it cites.
Eight methods to evaluate robust unlearning in llms
Lynch, A., Guo, P., Ewart, A., Casper, S., and Hadfield-Menell, D · 2024
Later among the works it cites.
The era of 1-bit llms: All large language models are in 1.58 bits
Ma, S., Wang, H., Ma, L., Wang, L., Wang, W., Huang, S., Dong, L., Wang, R., Xue, J., and Wei, F · 2024
Later among the works it cites.
TOFU: A task of fictitious unlearning for LLMs
Maini, P., Feng, Z., Schwarzschild, A., Lipton, Z. C., and Kolter, J. Z · 2024
Later among the works it cites.
In-context unlearning: Language models as few-shot unlearners
Pawelczyk, M., Neel, S., and Lakkaraju, H · 2024
Later among the works it cites.
ToolLLM: Facilitating large language models to master 16000+ real-world APIs
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., Zhao, S., Hong, L., Tian, R., Xie, R., Zhou, J., Gerstein, M., dahai li, Liu, Z., and Sun, M · 2024
Later among the works it cites.
Knowledge conflicts for llms: A survey
Xu, R., Qi, Z., Guo, Z., Wang, C., Wang, H., Zhang, Y., and Xu, W · 2024
Later among the works it cites.
Machine unlearning of pre-trained large language models
Yao, J., Chien, E., Du, M., Niu, X., Wang, T., Cheng, Z., and Yue, X · 2024
Later among the works it cites.
Evaluating large language models at evaluating instruction following
Zeng, Z., Yu, J., Gao, T., Meng, Y., Goyal, T., and Chen, D · 2024
Later among the works it cites.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2024
Later among the works it cites.
Understanding machine unlearning through the lens of mode connectivity
Cheng, J. and Amiri, H · 2025
Closest in time.
Unlearning or obfuscating? jogging the memory of unlearned LLMs via benign relearning
Hu, S., Fu, Y., Wu, S., and Smith, V · 2025
Closest in time.
An adversarial perspective on machine unlearning for AI safety
Łucki, J., Wei, B., Huang, Y., Henderson, P., Tramèr, F., and Rando, J · 2025
Closest in time.