Fetching the paper…
Reading the bibliography…
Safety and trustworthiness are indispensable requirements for real-world applications of AI systems using large language models (LLMs).
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Constrained optimization and Lagrange multiplier methods
D. P. Bertsekas · 2014
Earlier work this paper cites.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
The ethics of artificial intelligence
N. Bostrom and E. Yudkowsky · 2018
Earlier work this paper cites.
Batch policy learning under constraints
H. Le, C. Voloshin, and Y. Yue · 2019
Earlier work this paper cites.
Benchmarking safe exploration in deep reinforcement learning
A. Ray, J. Achiam, and D. Amodei · 2019
Earlier work this paper cites.
Mitigating gender bias in natural language processing: Literature review
T. Sun, A. Gaut, S. Tang, Y. Huang, M. ElSherief, J. Zhao, D. Mirza, E. Belding, K.-W. Chang, and W. Y. Wang · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
TRL: Transformer reinforcement learning
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, and S. Huang · 2020
Earlier work this paper cites.
Constrained Markov decision processes
E. Altman · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Provably efficient safe exploration via primal-dual policy optimization
D. Ding, X. Wei, Z. Yang, Z. Wang, and M. Jovanovic · 2021
Earlier work this paper cites.
The role of permutation invariance in linear mode connectivity of neural networks
R. Entezari, H. Sedghi, O. Saukh, and B. Neyshabur · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2021
Cited alongside, same era.
Mitigating political bias in language models through reinforced calibration
R. Liu, C. Jia, J. Wei, G. Xu, L. Wang, and S. Vosoughi · 2021
Cited alongside, same era.
Git re-basin: Merging models modulo permutation symmetries
S. Ainsworth, J. Hayase, and S. Srinivasa · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
A long way to go: Investigating length correlations in RLHF
P. Singhal, T. Goyal, J. Xu, and G. Durrett · 2023
Later among the works it cites.
Stanford Alpaca: An instruction-following LLaMA model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
Large language models in medicine
A. J. Thirunavukarasu, D. S. J. Ting, K. Elangovan, L. Gutierrez, T. F. Tan, and D. S. W. Ting · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Decodingtrust: A comprehensive assessment of trustworthiness in GPT models
B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Safe policies for reinforcement learning via primal-dual methods
S. Paternain, M. Calvo-Fullana, L. F. Chamon, and A. Ribeiro · 2022
Cited alongside, same era.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, et al · 2022
Cited alongside, same era.
Wordcraft: story writing with large language models
A. Yuan, A. Coenen, E. Reif, and D. Ippolito · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions
F. Bianchi, M. Suzgun, G. Attanasio, P. Rottger, D. Jurafsky, T. Hashimoto, and J. Zou · 2023
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with GPT-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Cited alongside, same era.
Later among the works it cites.
W. Xiong, H. Dong, C. Ye, H. Zhong, N. Jiang, and T. Zhang · 2023
Later among the works it cites.
Prompting large language model for machine translation: A case study
B. Zhang, B. Haddow, and A. Birch · 2023
Later among the works it cites.
Beyond one-preference-for-all: Multi-objective direct preference optimization
Z. Zhou, J. Liu, C. Yang, J. Shao, Y. Liu, X. Yue, W. Ouyang, and Y. Qiao · 2023
Later among the works it cites.
Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs
A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, O. Pietquin, A. Üstün, and S. Hooker · 2024
Closest in time.
A general theoretical paradigm to understand learning from human preferences
M. G. Azar, Z. D. Guo, B. Piot, R. Munos, M. Rowland, M. Valko, and D. Calandriello · 2024
Closest in time.
Safe RLHF: Safe reinforcement learning from human feedback
J. Dai, X. Pan, R. Sun, J. Ji, X. Xu, M. Liu, Y. Wang, and Y. Yang · 2024
Closest in time.
Last-iterate convergent policy gradient primal-dual methods for constrained MDPs
D. Ding, C.-Y. Wei, K. Zhang, and A. Ribeiro · 2024
Closest in time.
KTO: Model alignment as prospect theoretic optimization
K. Ethayarajh, W. Xu, N. Muennighoff, D. Jurafsky, and D. Kiela · 2024
Closest in time.
Arcee’s MergeKit: A toolkit for merging large language models
C. Goddard, S. Siriwardhana, M. Ehghaghi, L. Meyers, V. Karpukhin, B. Benedict, M. McQuade, and J. Solawetz · 2024
Closest in time.
Beavertails: Towards improved safety alignment of LLM via a human-preference dataset
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang · 2024
Closest in time.
Enhancing LLM safety via constrained direct preference optimization
Z. Liu, X. Sun, and Z. Zheng · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Beyond reverse KL: Generalizing direct preference optimization with diverse divergence constraints
C. Wang, Y. Jiang, C. Yang, H. Liu, and Y. Chen · 2024
Closest in time.
Panacea: Pareto alignment via preference adaptation for LLMs
Y. Zhong, C. Ma, X. Zhang, Z. Yang, Q. Zhang, S. Qi, and Y. Yang · 2024
Closest in time.