Fetching the paper…
Reading the bibliography…
Researchers have been studying approaches to steer the behavior of Large Language Models (LLMs) and build personalized LLMs tailored for various applications.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Extracting latent steering vectors from pretrained language models
N. Subramani, N. Suresh, and M. E. Peters · 2022
Earlier work this paper cites.
Calibrating sequence likelihood improves conditional language generation
Y. Zhao, M. Khalman, R. Joshi, S. Narayan, M. Saleh, and P. J. Liu · 2022
Earlier work this paper cites.
Llama2-chinese-7b-chat, 2023
FlagAlpha · 2023
Earlier work this paper cites.
A survey of large language models for healthcare: from data, technology, and applications to accountability and ethics, 2023
K. He, R. Mao, Q. Lin, Y. Ruan, X. Lan, M. Feng, and E. Cambria · 2023
Cited alongside, same era.
A review of opportunities and challenges of chatbots in education
G.-J. Hwang and C.-Y. Chang · 2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Cited alongside, same era.
Large language models in law: A survey, 2023
J. Lai, W. Gan, J. Wu, Z. Qi, and P. S. Yu · 2023
Cited alongside, same era.
Textbooks are all you need ii: phi-1.5 technical report
Y. Li, S. Bubeck, R. Eldan, A. Del Giorno, S. Gunasekar, and Y. T. Lee · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization
A. Turner, L. Thiergart, D. Udell, G. Leech, U. Mini, and M. MacDiarmid · 2023
Later among the works it cites.
H. Wang and K. Shu · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance, 2023
S. Wu, O. Irsoy, S. Lu, V. Dabravolski, M. Dredze, S. Gehrmann, P. Kambadur, D. Rosenberg, and G. Mann · 2023
Later among the works it cites.
Open-source can be dangerous: On the vulnerability of value alignment in open-source llms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Liu, L. Xing, and J. Zou · 2023
Cited alongside, same era.
A brief report on lawgpt 1.0: A virtual legal assistant based on gpt-3
H.-T. Nguyen · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay · 2023
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
E. Perez, S. Ringer, K. Lukosiute, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kadavath, et al · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson · 2023
Cited alongside, same era.
Steering llama 2 via contrastive activation addition
N. Rimsky, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. M. Turner · 2023
Cited alongside, same era.
J. Yi, R. Ye, Q. Chen, B. B. Zhu, S. Chen, D. Lian, G. Sun, X. Xie, and F. Wu · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback
Y. Zhao, R. Joshi, T. Liu, M. Khalman, M. Saleh, and P. J. Liu · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
K. Li, O. Patel, F. Viégas, H. Pfister, and M. Wattenberg · 2024
Closest in time.
Statistical rejection sampling improves preference optimization
T. Liu, Y. Zhao, R. Joshi, M. Khalman, M. Saleh, P. J. Liu, and J. Liu · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Revolutionizing finance with llms: An overview of applications and insights, 2024
H. Zhao, Z. Liu, Z. Wu, Y. Li, T. Yang, P. Shu, S. Xu, H. Dai, L. Zhao, G. Mai, N. Liu, and T. Liu · 2024
Closest in time.