Fetching the paper…
Reading the bibliography…
Common methods for aligning large language models (LLMs) with desired behaviour heavily rely on human-labelled data.
The Rating of Chessplayers: Past and Present
Elo, A · 1978
Earlier work this paper cites.
Trueskill™: a bayesian skill rating system
Herbrich, R., Minka, T., and Graepel, T · 2006
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Christiano, P., Shlegeris, B., and Amodei, D · 2018
Earlier work this paper cites.
Irving, G., Christiano, P., and Amodei, D · 2018
Earlier work this paper cites.
Finding generalizable evidence by learning to convince Q&A models
Perez, E., Karamcheti, S., Fergus, R., Weston, J., Kiela, D., and Cho, K · 2019
Earlier work this paper cites.
Debate update: Obfuscated arguments problem
Barnes, B · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Does the whole exceed its parts? the effect of ai explanations on complementary team performance
Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., and Weld, D · 2021
Earlier work this paper cites.
I think i get your point, ai! the illusion of explanatory depth in explainable ai
Chromik, M., Eiband, M., Buchner, F., Krüger, A., and Butz, A · 2021
Earlier work this paper cites.
The case for aligning narrowly superhuman models
Cotra, A · 2021
Earlier work this paper cites.
True few-shot learning with language models
Perez, E., Kiela, D., and Cho, K · 2021
Cited alongside, same era.
Constitutional AI: Harmlessness from AI feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Cited alongside, same era.
Measuring progress on scalable oversight for large language models
Bowman, S. R., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner, S., Lukosuite, K., Askell, A., Jones, A., Chen, A., et al · 2022
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Cited alongside, same era.
Teaching language models to support answers with verified quotes
Menick, J., Trebacz, M., Mikulik, V., Aslanides, J., Song, F., Chadwick, M., Glaese, M., Young, S., Campbell-Gillingham, L., Irving, G., et al · 2022
Aqua: A benchmarking tool for label quality assessment
Goswami, M., Sanil, V., Choudhry, A., Srinivasan, A., Udompanyawit, C., and Dubrawski, A · 2023
Later among the works it cites.
Ai control: Improving safety despite intentional subversion
Greenblatt, R., Shlegeris, B., Sachan, K., and Roger, F · 2023
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., and Clark, P · 2023
Later among the works it cites.
Debate helps supervise unreliable experts
Michael, J., Mahdi, S., Rein, D., Petty, J., Dirani, J., Padmakumar, V., and Bowman, S. R · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
QuALITY: Question answering with long input texts, yes!
Pang, R. Y., Parrish, A., Joshi, N., Nangia, N., Phang, J., Chen, A., Padmakumar, V., Ma, J., Thompson, J., He, H., et al · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., Sutskever, I., and Wu, J · 2023
Cited alongside, same era.
Improving factuality and reasoning in language models through multiagent debate
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I · 2023
Cited alongside, same era.
Nikola · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Anthropic fall 2023 debate progress update
Radhakrishnan, A · 2023
Later among the works it cites.
Question decomposition improves the faithfulness of model-generated reasoning
Radhakrishnan, A., Nguyen, K., Chen, A., Chen, C., Denison, C., Hernandez, D., Durmus, E., Hubinger, E., Kernion, J., Lukošiūtė, K., et al · 2023
Later among the works it cites.
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
Sleeper agents: Training deceptive llms that persist through safety training
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., et al · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.