Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but gathering high-quality preference labels is expensive.
The Problem of m m Rankings
Kendall, M. G. and Smith, B. B · 1939
Earlier work this paper cites.
Dynamic programming and markov processes
Howard, R. A · 1960
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Avoiding wireheading with value reinforcement learning
Everitt, T. and Hutter, M · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
Jaques, N., Gu, S., Bahdanau, D., Hernández-Lobato, J. M., Turner, R. E., and Eck, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Earlier work this paper cites.
A study of reinforcement learning for neural machine translation
Wu, L., Tian, F., Qin, T., Lai, J., and Liu, T.-Y · 2018
Earlier work this paper cites.
Learning to extract coherent summary via deep reinforcement learning
Wu, Y. and Hu, B · 2018
Earlier work this paper cites.
Reward learning for efficient reinforcement learning in extractive document summarisation
Gao, Y., Meyer, C. M., Mesgar, M., and Gurevych, I · 2019
Earlier work this paper cites.
A theory of regularized markov decision processes
Geist, M., Scherrer, B., and Pietquin, O · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Cited alongside, same era.
A survey of data augmentation approaches for NLP
Feng, S. Y., Gangal, V., Wei, J., Chandar, S., Vosoughi, S., Mitamura, T., and Hovy, E · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Cited alongside, same era.
Is GPT-3 a good data annotator?
Ding, B., Qin, C., Liu, L., Chia, Y. K., Li, B., Joty, S., and Bing, L · 2023
Closest in time.
RAFT: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., SHUM, K., and Zhang, T · 2023
Closest in time.
Chatgpt outperforms crowd-workers for text-annotation tasks
Gilardi, F., Alizadeh, M., and Kubli, M · 2023
Closest in time.
Ai platform data labeling service pricing
Google · 2023
Closest in time.
Palm 2 technical report, 2023
Google, R. A., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., Chu, E., Clark, J. H., Shafey, L. E., Huang, Y., Meier-Hellstern, K., Mishra, G., Moreira, E., Omernick, M., Robinson, K., Ruder, S., Tay, Y., Xiao, K., Xu, Y., Zhang, Y., Abrego, G. H., Ahn, J., Austin, J., Barham, P., Botha, J., Bradbury, J., Brahma, S., Brooks, K., Catasta, M., Cheng, Y., Cherry, C., Choquette-Choo, C. A., Chowdhery, A., Crepy, C., Dave, S., Dehghani, M., Dev, S., Devlin, J., Díaz, M., Du, N., Dyer, E., Feinberg, V., Feng, F., Fienber, V., Freitag, M., Garcia, X., Gehrmann, S., Gonzalez, L., Gur-Ari, G., Hand, S., Hashemi, H., Hou, L., Howland, J., Hu, A., Hui, J., Hurwitz, J., Isard, M., Ittycheriah, A., Jagielski, M., Jia, W., Kenealy, K., Krikun, M., Kudugunta, S., Lan, C., Lee, K., Lee, B., Li, E., Li, M., Li, W., Li, Y., Li, J., Lim, H., Lin, H., Liu, Z., Liu, F., Maggioni, M., Mahendru, A., Maynez, J., Misra, V., Moussalem, M., Nado, Z., Nham, J., Ni, E., Nystrom, A., Parrish, A., Pellat, M., Polacek, M., Polozov, A., Pope, R., Qiao, S., Reif, E., Richter, B., Riley, P., Ros, A. C., Roy, A., Saeta, B., Samuel, R., Shelby, R., Slone, A., Smilkov, D., So, D. R., Sohn, D., Tokumine, S., Valter, D., Vasudevan, V., Vodrahalli, K., Wang, X., Wang, P., Wang, Z., Wang, T., Wieting, J., Wu, Y., Xu, K., Xu, Y., Xue, L., Yin, P., Yu, J., Zhang, Q., Zheng, S., Zheng, C., Zhou, W., Zhou, D., Petrov, S., and Wu, Y · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Want to reduce labeling cost? gpt-3 can help
Wang, S., Liu, Y., Xu, Y., Zhu, C., and Zeng, M · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Cited alongside, same era.
Understanding dataset difficulty with 𝒱 \mathcal{V} -usable information
Ethayarajh, K., Choi, Y., and Swayamdipta, S · 2022
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al · 2022
Cited alongside, same era.
Large language models can self-improve
Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2022
Cited alongside, same era.
Reward design with language models
Kwon, M., Xie, S. M., Bullard, K., and Sadigh, D · 2022
Cited alongside, same era.
Closest in time.
Lai, V. D., Van Nguyen, C., Ngo, N. T., Nguyen, T., Dernoncourt, F., Rossi, R. A., and Nguyen, T. H · 2023
Closest in time.
Summary of chatgpt/gpt-4 research and perspective towards the future of large language models
Liu, Y., Han, T., Ma, S., Zhang, J., Yang, Y., Tian, J., He, H., Li, A., He, M., Liu, Z., et al · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2023
Closest in time.
An overview of bard: an early experiment with generative ai
Manyika, J · 2023
Closest in time.
Tuning language models as training data generators for augmentation-enhanced few-shot learning
Meng, Y., Michalski, M., Huang, J., Zhang, Y., Abdelzaher, T., and Han, J · 2023
Closest in time.
Openai pricing
OpenAI · 2023
Closest in time.
Large language models sensitivity to the order of options in multiple-choice questions
Pezeshkpour, P. and Hruschka, E · 2023
Closest in time.
Factually consistent summarization via reinforcement learning with textual entailment feedback
Roit, P., Ferret, J., Shani, L., Aharoni, R., Cideron, G., Dadashi, R., Geist, M., Girgin, S., Hussenot, L., Keller, O., et al · 2023
Closest in time.
Large language models are not fair evaluators
Wang, P., Li, L., Chen, L., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., and Sui, Z · 2023
Closest in time.
Rlcd: Reinforcement learning from contrast distillation for language model alignment, 2023
Yang, K., Klein, D., Celikyilmaz, A., Peng, N., and Tian, Y · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.