Fetching the paper…
Reading the bibliography…
The diversity of contexts in which large language models (LLMs) are deployed requires the ability to modify or customize default model behaviors to incorporate nuanced requirements and preferences.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
An application of reinforcement learning to dialogue strategy selection in a spoken dialogue system for email
Walker, M. A · 2000
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Peters, J. and Schaal, S · 2007
Earlier work this paper cites.
Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm
Busa-Fekete, R., Szörényi, B., Weng, P., Cheng, W., and Hüllermeier, E · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
Li, J., Monroe, W., Ritter, A., Jurafsky, D., Galley, M., and Gao, J · 2016
Earlier work this paper cites.
Dialogue learning with human-in-the-loop
Li, J., Miller, A. H., Chopra, S., Ranzato, M., and Weston, J · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Better rewards yield better summaries: Learning to summarise without references
Böhm, F., Gao, Y., Meyer, C. M., Shapira, O., Dagan, I., and Gurevych, I · 2019
Earlier work this paper cites.
Learning from dialogue after deployment: Feed yourself, chatbot!
Hancock, B., Bordes, A., Mazare, P.-E., and Weston, J · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Earlier work this paper cites.
Dark experience for general continual learning: a strong, simple baseline
Buzzega, P., Boschini, M., Porrello, A., Abati, D., and Calderara, S · 2020
Earlier work this paper cites.
Editable neural networks
Sinitsin, A., Plokhotnyuk, V., Pyrkin, D., Popov, S., and Babenko, A · 2020
Earlier work this paper cites.
Fine-tuning language models from human preferences, 2020
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment, 2021
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Kernion, J., Ndousse, K., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., and Kaplan, J · 2021
Earlier work this paper cites.
Program synthesis with large language models, 2021
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., and Sutton, C · 2021
Cited alongside, same era.
Editing factual knowledge in language models
Cao, N. D., Aziz, W., and Titov, I · 2021
Cited alongside, same era.
Transformers as soft reasoners over language
Clark, P., Tafjord, O., and Richardson, K · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Fast model editing at scale
Mitchell, E., Lin, C. P., Bosselut, A., Finn, C., and Manning, C. D · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukošiūtė, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T. B., and Kaplan, J · 2022
Supervised fine-tuning and direct preference optimization on intel gaudi2, November 2023
Intel · 2023
Later among the works it cites.
Mistral 7b, 2023
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Later among the works it cites.
The open instruction generalist dataset, 2023
LAION · 2023
Later among the works it cites.
Self: Language-driven self-evolution for large language model
Lu, J., Zhong, W., Huang, W., Wang, Y., Mi, F., Wang, B., Wang, W., Shang, L., and Liu, Q · 2023
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback, 2023
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Welleck, S., Majumder, B. P., Gupta, S., Yazdanbakhsh, A., and Clark, P · 2023
Later among the works it cites.
Editing personality for llms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgements, 2022
Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G · 2022
Cited alongside, same era.
Locating and editing factual knowledge in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Cited alongside, same era.
Memory-based model editing at scale
Mitchell, E., Lin, C. P., Bosselut, A., Manning, C. D., Finn, C., Mitchell, E., Lin, C. P., Bosselut, A., Manning, C. D., and Finn, C · 2022
Cited alongside, same era.
Fixing model bugs with natural language patches
Murty, S., Manning, C. D., Lundberg, S. M., and Ribeiro, M. T · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
When life gives you lemons, make cherryade: Converting feedback from bad responses into good labels, 2022
Shi, W., Dinan, E., Shuster, K., Weston, J., and Xu, J · 2022
Cited alongside, same era.
Mao, S., Zhang, N., Wang, X., Wang, M., Yao, Y., Jiang, Y., Xie, P., Huang, F., and Chen, H · 2023
Later among the works it cites.
Automatically correcting large language models: Surveying the landscape of diverse self-correction strategies
Pan, L., Saxon, M. S., Xu, W., Nathani, D., Wang, X., and Wang, W. Y · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Later among the works it cites.
Dueling rl: Reinforcement learning with trajectory preferences
Saha, A., Pacchiano, A., and Lee, J · 2023
Later among the works it cites.
Training language models with language feedback at scale, 2023
Scheurer, J., Campos, J. A., Korbak, T., Chan, J. S., Chen, A., Cho, K., and Perez, E · 2023
Later among the works it cites.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D. D., Yang, Y., and Gan, C · 2023
Later among the works it cites.
Teaching language models to self-improve through interactive demonstrations, 2023
Yu, X., Peng, B., Galley, M., Gao, J., and Yu, Z · 2023
Later among the works it cites.
Suppressing pink elephants with direct principle feedback, 2024
Castricato, L., Lile, N., Anand, S., Schoelkopf, H., Verma, S., and Biderman, S · 2024
Closest in time.
Model editing with canonical examples, 2024
Hewitt, J., Chen, S., Xie, L. L., Adams, E., Liang, P., and Manning, C. D · 2024
Closest in time.
RLCD: Reinforcement learning from contrastive distillation for LM alignment
Yang, K., Klein, D., Celikyilmaz, A., Peng, N., and Tian, Y · 2024
Closest in time.
Self-rewarding language models, 2024
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Closest in time.