Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated powerful capabilities that render them valuable in different applications, including conversational AI products.
Advances in prospect theory: Cumulative representation of uncertainty
Tversky, A.; and Kahneman, D. 1992 · 1992
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022 · 2022
Earlier work this paper cites.
Chain of Thought Prompting Elicits Reasoning in Large Language Models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Chi, E. H.; Le, Q.; and Zhou, D. 2022 · 2022
Earlier work this paper cites.
The Falcon Series of Open Language Models
Almazrouei, E.; Alobeidli, H.; Alshamsi, A.; Cappelli, A.; Cojocaru, R.; Debbah, M.; Étienne Goffinet; Hesslow, D.; Launay, J.; Malartic, Q.; Mazzotta, D.; Noune, B.; Pannier, B.; and Penedo, G. 2023 · 2023
Earlier work this paper cites.
Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)
Biswas, A.; and Talukdar, W. 2023 · 2023
Earlier work this paper cites.
Combating misinformation in the age of llms: Opportunities and challenges
Chen, C.; and Shu, K. 2023 · 2023
Earlier work this paper cites.
RAGAS: Automated Evaluation of Retrieval Augmented Generation
Es, S.; James, J.; Espinosa-Anke, L.; and Schockaert, S. 2023 · 2023
Earlier work this paper cites.
Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects
Hadi, M. U.; Qureshi, R.; Shah, A.; Irfan, M.; Zafar, A.; Shaikh, M. B.; Akhtar, N.; Wu, J.; Mirjalili, S.; et al. 2023 · 2023
Earlier work this paper cites.
Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Huang, Y.; Gupta, S.; Xia, M.; Li, K.; and Chen, D. 2023 · 2023
Earlier work this paper cites.
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Inan, H.; Upasani, K.; Chi, J.; Rungta, R.; Iyer, K.; Mao, Y.; Tontchev, M.; Hu, Q.; Fuller, B.; Testuggine, D.; and Khabsa, M. 2023 · 2023
Earlier work this paper cites.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; de las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; Lavaud, L. R.; Lachaux, M.-A.; Stock, P.; Scao, T. L.; Lavril, T.; Wang, T.; Lacroix, T.; and Sayed, W. E. 2023 · 2023
Earlier work this paper cites.
Automatically Auditing Large Language Models via Discrete Optimization
Jones, E.; Dragan, A.; Raghunathan, A.; and Steinhardt, J. 2023 · 2023
Earlier work this paper cites.
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Li, X.; Zhou, Z.; Zhu, J.; Yao, J.; Liu, T.; and Han, B. 2023 · 2023
Earlier work this paper cites.
Lost in the Middle: How Language Models Use Long Contexts
Liu, N. F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P. 2023 · 2023
Earlier work this paper cites.
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Manakul, P.; Liusie, A.; and Gales, M. J. F. 2023 · 2023
Earlier work this paper cites.
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
Min, S.; Krishna, K.; Lyu, X.; Lewis, M.; tau Yih, W.; Koh, P. W.; Iyyer, M.; Zettlemoyer, L.; and Hajishirzi, H. 2023 · 2023
Earlier work this paper cites.
Fine-Tuned DeBERTa-v3 for Prompt Injection Detection
ProtectAI.com. 2023 · 2023
Earlier work this paper cites.
NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails
Rebedea, T.; Dinu, R.; Sreedhar, M.; Parisien, C.; and Cohen, J. 2023 · 2023
Earlier work this paper cites.
Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Shah, R.; Feuillade-Montixi, Q.; Pour, S.; Tagade, A.; Casper, S.; and Rando, J. 2023 · 2023
Cited alongside, same era.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Ferrer, C. C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goyal, N.; et al. 2023 · 2023
Cited alongside, same era.
”Kelly is a Warm Person, Joseph is a Role Model”: Gender Biases in LLM-Generated Reference Letters
Wan, Y.; Pu, G.; Sun, J.; Garimella, A.; Chang, K.-W.; and Peng, N. 2023 · 2023
Cited alongside, same era.
Adding guardrails to advanced chatbots
Wang, Y.; and Singh, L. 2023 · 2023
Cited alongside, same era.
Generative Judge for Evaluating Alignment
Li, J.; Sun, S.; Yuan, W.; Fan, R.-Z.; hai zhao; and Liu, P. 2024 · 2024
Later among the works it cites.
Automatic and Universal Prompt Injection Attacks against Large Language Models
Liu, X.; Yu, Z.; Zhang, Y.; Zhang, N.; and Xiao, C. 2024 · 2024
Later among the works it cites.
An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
Luo, Y.; Yang, Z.; Meng, F.; Li, Y.; Zhou, J.; and Zhang, Y. 2024 · 2024
Later among the works it cites.
Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning
Mecklenburg, N.; Lin, Y.; Li, X.; Holstein, D.; Nunes, L.; Malvar, S.; Silva, B.; Chandra, R.; Aski, V.; Yannam, P. K. R.; Aktas, T.; and Hendry, T. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei, Z.; Wang, Y.; Li, A.; Mo, Y.; and Wang, Y. 2023 · 2023
Cited alongside, same era.
Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery
Wen, Y.; Jain, N.; Kirchenbauer, J.; Goldblum, M.; Geiping, J.; and Goldstein, T. 2023 · 2023
Cited alongside, same era.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E. P.; Zhang, H.; Gonzalez, J. E.; and Stoica, I. 2023 · 2023
Cited alongside, same era.
AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models
Zhu, S.; Zhang, R.; An, B.; Wu, G.; Barrow, J.; Wang, Z.; Huang, F.; Nenkova, A.; and Sun, T. 2023 · 2023
Cited alongside, same era.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Zou, A.; Wang, Z.; Carlini, N.; Nasr, M.; Kolter, J. Z.; and Fredrikson, M. 2023 · 2023
Cited alongside, same era.
Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices
Abdali, S.; Anarfi, R.; Barberan, C.; and He, J. 2024 · 2024
Cited alongside, same era.
LLM-Based Chatbots for Mining Software Repositories: Challenges and Opportunities
Abedu, S.; Abdellatif, A.; and Shihab, E. 2024 · 2024
Cited alongside, same era.
Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
Adlakha, V.; BehnamGhader, P.; Lu, X. H.; Meade, N.; and Reddy, S. 2024 · 2024
Cited alongside, same era.
Mehrotra, A.; Zampetakis, M.; Kassianik, P.; Nelson, B.; Anderson, H.; Singer, Y.; and Karbasi, A. 2024 · 2024
Later among the works it cites.
Nghiem, H.; Prindle, J.; Zhao, J.; and au2, H. D. I. 2024 · 2024
Later among the works it cites.
OpenAI; Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; Avila, R.; Babuschkin, I.; Balaji, S.; Balcom, V.; et al. 2024 · 2024
Later among the works it cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2024 · 2024
Later among the works it cites.
Analyzing and Reducing Catastrophic Forgetting in Parameter Efficient Tuning
Ren, W.; Li, X.; Wang, L.; Zhao, T.; and Qin, W. 2024 · 2024
Later among the works it cites.
Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
Schimanski, T.; Ni, J.; Kraus, M.; Ash, E.; and Leippold, M. 2024 · 2024
Later among the works it cites.
Shen, X.; Chen, Z.; Backes, M.; Shen, Y.; and Zhang, Y. 2024 · 2024
Later among the works it cites.
Meta Llama Guard 2
Team, L. 2024 · 2024
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Wei, A.; Haghtalab, N.; and Steinhardt, J. 2024 · 2024
Later among the works it cites.
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
Xu, Z.; Liu, Y.; Deng, G.; Li, Y.; and Picek, S. 2024 · 2024
Later among the works it cites.
Large Language Models Meet Text-Centric Multimodal Sentiment Analysis: A Survey
Yang, H.; Zhao, Y.; Wu, Y.; Wang, S.; Zheng, T.; Zhang, H.; Ma, Z.; Che, W.; and Qin, B. 2024 · 2024
Later among the works it cites.
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Yi, S.; Liu, Y.; Sun, Z.; Cong, T.; He, X.; Song, J.; Xu, K.; and Li, Q. 2024 · 2024
Later among the works it cites.
Optimization Techniques for Sentiment Analysis Based on LLM (GPT-3)
Zhan, T.; Shi, C.; Shi, Y.; Li, H.; and Lin, Y. 2024 · 2024
Later among the works it cites.
Defending large language models against jailbreaking attacks through goal prioritization
Zhang, Z.; Yang, J.; Ke, P.; Mi, F.; Wang, H.; and Huang, M. 2024 · 2024
Later among the works it cites.