Fetching the paper…
Reading the bibliography…
Large language models (LLMs) often contain misleading content, emphasizing the need to align them with human values to ensure secure AI systems.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M.; Stiennon, N.; Wu, J.; Brown, T. B.; Radford, A.; Amodei, D.; Christiano, P.; and Irving, G. 2019 · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Bradley, R. A.; and Terry, M. E. 1952 · 1952
Earlier work this paper cites.
The analysis of permutations
Plackett, R. L. 1975 · 1975
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W.-J. 2002 · 2002
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
Luce, R. D. 2012 · 2012
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017 · 2017
Earlier work this paper cites.
Interactive Learning from Policy-Dependent Human Feedback
MacGlashan, J.; Ho, M. K.; Loftin, R.; Peng, B.; Wang, G.; Roberts, D. L.; Taylor, M. E.; and Littman, M. L. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Earlier work this paper cites.
Deep TAMER: Interactive Agent Shaping in High-Dimensional State Spaces
Warnell, G.; Waytowich, N.; Lawhern, V.; and Stone, P. 2018 · 2018
Earlier work this paper cites.
Momentum Contrast for Unsupervised Visual Representation Learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020 · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N.; Ouyang, L.; Wu, J.; Ziegler, D.; Lowe, R.; Voss, C.; Radford, A.; Amodei, D.; and Christiano, P. F. 2020 · 2020
Earlier work this paper cites.
Transformers: State-of-the-Art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Le Scao, T.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. 2020 · 2020
Earlier work this paper cites.
Lee, K.; Smith, L.; and Abbeel, P. 2021 · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R.; Hilton, J.; Balaji, S.; Wu, J.; Ouyang, L.; Kim, C.; Hesse, C.; Jain, S.; Kosaraju, V.; Saunders, W.; et al. 2021 · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2022 · 2022
Cited alongside, same era.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling
Du, Z.; Qian, Y.; Liu, X.; Ding, M.; Qiu, J.; Yang, Z.; and Tang, J. 2022 · 2022
Cited alongside, same era.
Scaling Laws for Reward Model Overoptimization
Gao, L.; Schulman, J.; and Hilton, J. 2022 · 2022
Cited alongside, same era.
Accelerate: Training and inference at scale made simple, efficient and adaptable
OpenAI. 2023 · 2023
Closest in time.
Peng, B.; Li, C.; He, P.; Galley, M.; and Gao, J. 2023 · 2023
Closest in time.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023 · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gugger, S.; Debut, L.; Wolf, T.; Schmid, P.; Mueller, Z.; and Mangrulkar, S. 2022 · 2022
Cited alongside, same era.
Interacting with Non-Cooperative User: A New Paradigm for Proactive Dialogue Policy
Lei, W.; Zhang, Y.; Song, F.; Liang, H.; Mao, J.; Lv, J.; Yang, Z.; and Chua, T.-S. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Cited alongside, same era.
Offline rl for natural language generation with implicit language q learning
Snell, C.; Kostrikov, I.; Su, Y.; Yang, M.; and Levine, S. 2022 · 2022
Cited alongside, same era.
Self-Instruct: Aligning Language Model with Self Generated Instructions
Wang, Y.; Kordi, Y.; Mishra, S.; Liu, A.; Smith, N. A.; Khashabi, D.; and Hajishirzi, H. 2022 · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; et al. 2023 · 2023
Cited alongside, same era.
Raft: Reward ranked finetuning for generative foundation model alignment
Dong, H.; Xiong, W.; Goyal, D.; Pan, R.; Diao, S.; Zhang, J.; Shum, K.; and Zhang, T. 2023 · 2023
Cited alongside, same era.
Api-bank: A benchmark for tool-augmented llms
Li, M.; Song, F.; Yu, B.; Yu, H.; Li, Z.; Huang, F.; and Li, Y. 2023 · 2023
Cited alongside, same era.
Closest in time.
Large language models are not fair evaluators
Wang, P.; Li, L.; Chen, L.; Zhu, D.; Lin, B.; Cao, Y.; Liu, Q.; Liu, T.; and Sui, Z. 2023 · 2023
Closest in time.
Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
Wu, Z.; Hu, Y.; Shi, W.; Dziri, N.; Suhr, A.; Ammanabrolu, P.; Smith, N. A.; Ostendorf, M.; and Hajishirzi, H. 2023 · 2023
Closest in time.
Reinforcement Learning from Diverse Human Preferences
Xue, W.; An, B.; Yan, S.; and Xu, Z. 2023 · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z.; Yuan, H.; Tan, C.; Wang, W.; Huang, S.; and Huang, F. 2023 · 2023
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023 · 2023
Closest in time.
Lima: Less is more for alignment
Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; Yu, L.; et al. 2023 · 2023
Closest in time.
Principled Reinforcement Learning with Human Feedback from Pairwise or K K -wise Comparisons
Zhu, B.; Jiao, J.; and Jordan, M. I. 2023 · 2023
Closest in time.
Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Zhu, B.; Sharma, H.; Frujeri, F. V.; Dong, S.; Zhu, C.; Jordan, M. I.; and Jiao, J. 2023 · 2023
Closest in time.