Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown great potential as general-purpose AI assistants across various domains.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1989
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D. G · 1992
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I. J., and Fergus, R · 2014
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M. and Lee, S · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization
Gliwa, B., Mochol, I., Biesek, M., and Wawer, A · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Accurate post training quantization with small calibration sets
Hubara, I., Nahshan, Y., Hanani, Y., Banner, R., and Soudry, D · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Earlier work this paper cites.
Constitutional AI: harmlessness from AI feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., et al · 2022
Earlier work this paper cites.
Multi-objective deep learning with adaptive reference vectors
Chen, W. and Kwok, J · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Efficient combinatorial optimization for word-level adversarial textual attack
Liu, S., Lu, N., Chen, C., and Tang, K · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Gpt-4 technical report
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Claude, 2023
Anthropic · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Cited alongside, same era.
Beavertails: Towards improved safety alignment of LLM via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., and Yang, Y · 2023
Cited alongside, same era.
Jailbreaking chatgpt via prompt engineering: An empirical study
Liu, Y., Deng, G., Xu, Z., Li, Y., Zheng, Y., Zhang, Y., Zhao, L., Zhang, T., and Liu, Y · 2023
Cited alongside, same era.
Instruction tuning with GPT-4
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Cited alongside, same era.
Safe loRA: The silver lining of reducing safety risks when finetuning large language models
Hsu, C.-Y., Tsai, Y.-L., Lin, C.-H., Chen, P.-Y., Yu, C.-M., and Huang, C.-Y · 2024
Later among the works it cites.
Large language models can be guided to evade ai-generated text detection
Lu, N., Liu, S., He, R., Ong, Y., Wang, Q., and Tang, K · 2024
Later among the works it cites.
Less is more: Understanding word-level textual adversarial attack via n-gram frequency descend
Lu, N., Liu, S., Zhang, Z., Wang, Q., Liu, H., and Tang, K · 2024
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P., Jia, R., Mittal, P., and Henderson, P · 2024
Later among the works it cites.
Modelgrow: Continual text-to-video pre-training with model expansion and language understanding enhancement
Rao, Z., Ji, L., Xing, Y., Liu, R., Liu, Z., Xie, J., Peng, Z., He, Y., and Chen, Q · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
KICGPT: large language model with knowledge in context for knowledge graph completion
Wei, Y., Huang, Q., Zhang, Y., and Kwok, J. T · 2023
Cited alongside, same era.
Shadow alignment: The ease of subverting safely-aligned language models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Cited alongside, same era.
Enhancing meta learning via multi-objective soft improvement functions
Yu, R., Chen, W., Wang, X., and Kwok, J. T · 2023
Cited alongside, same era.
Representation noising: A defence mechanism against harmful finetuning
Rosati, D., Wehner, J., Williams, K., Bartoszcze, L., Gonzales, R., Maple, C., Majumdar, S., Sajjad, H., and Rudzicz, F · 2024
Later among the works it cites.
”do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2024
Later among the works it cites.
Backdooralign: Mitigating fine-tuning based jailbreak attack with backdoor enhanced safety alignment
Wang, J., Li, J., Li, Y., Qi, X., Hu, J., Li, Y., McDaniel, P., Chen, M., Li, B., and Xiao, C · 2024
Later among the works it cites.
GITA: graph to visual and textual integration for vision-language graph reasoning
Wei, Y., Fu, S., Jiang, W., Zhang, Z., Zeng, Z., Wu, Q., Kwok, J. T., and Zhang, Y · 2024
Later among the works it cites.
Backdoor graph condensation
Wu, J., Lu, N., Dai, Z., Fan, W., Liu, S., Li, Q., and Tang, K · 2024
Later among the works it cites.
RLCD: reinforcement learning from contrastive distillation for LM alignment
Yang, K., Klein, D., Celikyilmaz, A., Peng, N., and Tian, Y · 2024
Later among the works it cites.
GPT-4 is too smart to be safe: Stealthy chat with llms via cipher
Yuan, Y., Jiao, W., Wang, W., Huang, J., He, P., Shi, S., and Tu, Z · 2024
Later among the works it cites.
Removing RLHF protections in GPT-4 via fine-tuning
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2024
Later among the works it cites.
Model tailor: Mitigating catastrophic forgetting in multi-modal large language models
Zhu, D., Sun, Z., Li, Z., Shen, T., Yan, K., Ding, S., Wu, C., and Kuang, K · 2024
Later among the works it cites.
Gradient-based multi-objective deep learning: Algorithms, theories, applications, and beyond
Chen, W., Zhang, X., Lin, B., Lin, X., Zhao, H., Zhang, Q., and Kwok, J. T · 2025
Closest in time.
Booster: Tackling harmful fine-tuning for large language models via attenuating harmful perturbation
Huang, T., Hu, S., Ilhan, F., Tekin, S. F., and Liu, L · 2025
Closest in time.
Fine-tuning now available for gpt-4o, 2024
Peng, A., Allard, J., and Heidel, S · 2025
Closest in time.
Open the eyes of mpnn: Vision enhances mpnn in link prediction, 2025
Wei, Y., Wang, X., Zhuang, Z., Chen, Y., Chen, S., Zhang, Y., Zhang, Y., and Kwok, J · 2025
Closest in time.
Tf-dcon: Leveraging large language models (llms) to empower training-free dataset condensation for content-based recommendation, 2025
Wu, J., Liu, Q., Hu, H., Fan, W., Liu, S., Li, Q., Wu, X.-M., and Tang, K · 2025
Closest in time.
NLSR: neuron-level safety realignment of large language models against harmful fine-tuning
Yi, X., Zheng, S., Wang, L., de Melo, G., Wang, X., and He, L · 2025
Closest in time.