Fetching the paper…
Reading the bibliography…
Common methods for aligning already-capable models with desired behavior rely on the ability of humans to provide supervision.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019 · 1905
Earlier work this paper cites.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
Huang, L.; Bras, R. L.; Bhagavatula, C.; and Choi, Y. 2019 · 1909
Earlier work this paper cites.
On the computational complexity of combinatorial problems
Karp, R. M. 1975 · 1975
Earlier work this paper cites.
A Framework for Behavioural Cloning
Bain, M.; and Sammut, C. 1995 · 1995
Earlier work this paper cites.
Robot learning from demonstration
Atkeson, C. G.; and Schaal, S. 1997 · 1997
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P.; Schulman, J.; and Mané, D. 2016 · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017 · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B.; Pritzel, A.; and Blundell, C. 2017 · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J.; Liu, N. F.; and Gardner, M. 2017 · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Christiano, P.; Shlegeris, B.; and Amodei, D. 2018 · 2018
Earlier work this paper cites.
Irving, G.; Christiano, P.; and Amodei, D. 2018 · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J.; Krueger, D.; Everitt, T.; Martic, M.; Maini, V.; and Legg, S. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Earlier work this paper cites.
AI safety via market making
Hubinger, E. 2020 · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N.; Ouyang, L.; Wu, J.; Ziegler, D.; Lowe, R.; Voss, C.; Radford, A.; Amodei, D.; and Christiano, P. F. 2020 · 2020
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2021 · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. 2022 · 2022
Cited alongside, same era.
Measuring progress on scalable oversight for large language models
Bowman, S. R.; Hyun, J.; Perez, E.; Chen, E.; Pettit, C.; Heiner, S.; Lukošiūtė, K.; Askell, A.; Jones, A.; Chen, A.; et al. 2022 · 2022
Cited alongside, same era.
Ensemble deep learning: A review
Ai alignment: A comprehensive survey
Ji, J.; Qiu, T.; Chen, B.; Zhang, B.; Lou, H.; Wang, K.; Duan, Y.; He, Z.; Zhou, J.; Zhang, Z.; et al. 2023 · 2023
Later among the works it cites.
Combining weak-to-strong generalization with scalable oversight
Leike, J. 2023 · 2023
Later among the works it cites.
Encouraging divergent thinking in large language models through multi-agent debate
Liang, T.; He, Z.; Jiao, W.; Wang, X.; Wang, Y.; Wang, R.; Yang, Y.; Tu, Z.; and Shi, S. 2023 · 2023
Later among the works it cites.
Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; and Cobbe, K. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ganaie, M. A.; Hu, M.; Malik, A. K.; Tanveer, M.; and Suganthan, P. N. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators
Saunders, W.; Yeh, C.; Wu, J.; Bills, S.; Ouyang, L.; Ward, J.; and Leike, J. 2022 · 2022
Cited alongside, same era.
Introducing claude
Anthropic. 2023 · 2023
Cited alongside, same era.
Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; et al. 2023 · 2023
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C.; Izmailov, P.; Kirchner, J. H.; Baker, B.; Gao, L.; Aschenbrenner, L.; Chen, Y.; Ecoffet, A.; Joglekar, M.; Leike, J.; et al. 2023 · 2023
Cited alongside, same era.
Statement on AI Risk
CAIS. 2023 · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S.; Davies, X.; Shi, C.; Gilbert, T. K.; Scheurer, J.; Rando, J.; Freedman, R.; Korbak, T.; Lindner, D.; Freire, P.; et al. 2023 · 2023
Cited alongside, same era.
Michael, J.; Mahdi, S.; Rein, D.; Petty, J.; Dirani, J.; Padmakumar, V.; and Bowman, S. R. 2023 · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI. 2023 · 2023
Later among the works it cites.
Retrieval meets long context large language models
Xu, P.; Ping, W.; Wu, X.; McAfee, L.; Zhu, C.; Liu, Z.; Subramanian, S.; Bakhturina, E.; Shoeybi, M.; and Catanzaro, B. 2023 · 2023
Later among the works it cites.
Quantifying the Gain in Weak-to-Strong Generalization
Charikar, M.; Pabbaraju, C.; and Shiragur, K. 2024 · 2024
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2024 · 2024
Later among the works it cites.
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Denison, C.; MacDiarmid, M.; Barez, F.; Duvenaud, D.; Kravec, S.; Marks, S.; Schiefer, N.; Soklaski, R.; Tamkin, A.; Kaplan, J.; et al. 2024 · 2024
Later among the works it cites.
On scalable oversight with weak LLMs judging strong LLMs
Kenton, Z.; Siegel, N. Y.; Kramár, J.; Brown-Cohen, J.; Albanie, S.; Bulian, J.; Agarwal, R.; Lindner, D.; Tang, Y.; Goodman, N. D.; et al. 2024 · 2024
Later among the works it cites.
Debating with more persuasive llms leads to more truthful answers
Khan, A.; Hughes, J.; Valentine, D.; Ruis, L.; Sachan, K.; Radhakrishnan, A.; Grefenstette, E.; Bowman, S. R.; Rocktäschel, T.; and Perez, E. 2024 · 2024
Later among the works it cites.
Co-supervised learning: Improving weak-to-strong generalization with hierarchical mixture of experts
Liu, Y.; and Alahi, A. 2024 · 2024
Later among the works it cites.
LLM Critics Help Catch LLM Bugs
McAleese, N.; Pokorny, R. M.; Uribe, J. F. C.; Nitishinskaya, E.; Trebacz, M.; and Leike, J. 2024 · 2024
Later among the works it cites.
Easy-to-hard generalization: Scalable alignment beyond human supervision
Sun, Z.; Yu, L.; Shen, Y.; Liu, W.; Yang, Y.; Welleck, S.; and Gan, C. 2024 · 2024
Later among the works it cites.
Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; Dong, G.; Wei, H.; Lin, H.; Tang, J.; Wang, J.; Yang, J.; Tu, J.; Zhang, J.; Ma, J.; Xu, J.; Zhou, J.; Bai, J.; He, J.; Lin, J.; Dang, K.; Lu, K.; Chen, K.; Yang, K.; Li, M.; Xue, M.; Ni, N.; Zhang, P.; Wang, P.; Peng, R.; Men, R.; Gao, R.; Lin, R.; Wang, S.; Bai, S.; Tan, S.; Zhu, T.; Li, T.; Liu, T.; Ge, W.; Deng, X.; Zhou, X.; Ren, X.; Zhang, X.; Wei, X.; Ren, X.; Fan, Y.; Yao, Y.; Zhang, Y.; Wan, Y.; Chu, Y.; Liu, Y.; Cui, Z.; Zhang, Z.; and Fan, Z. 2024 · 2024
Later among the works it cites.