Fetching the paper…
Reading the bibliography…
Reward-guided text generation (RGTG) has emerged as a viable alternative to offline reinforcement learning from human feedback (RLHF).
Rank analysis of incomplete block designs: The method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Adversarial learning for neural dialogue generation
Li, J., Monroe, W., Shi, T., Jean, S., Ritter, A., and Jurafsky, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Tl; dr: Mining reddit to learn automatic summarization
Völske, M., Potthast, M., Syed, S., and Stein, B · 2017
Earlier work this paper cites.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Earlier work this paper cites.
Proximal policy optimization and its dynamic version for sequence generation
Tuan, Y.-L., Zhang, J., Li, Y., and Lee, H.-y · 2018
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R · 2019
Earlier work this paper cites.
Improving conditional sequence generative adversarial networks by stepwise evaluation
Tuan, Y.-L. and Lee, H.-Y · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2020
Earlier work this paper cites.
GeDi: Generative discriminator guided sequence generation
Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F · 2021
Earlier work this paper cites.
PEBBLE: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
Lee, K., Smith, L. M., and Abbeel, P · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Cited alongside, same era.
Fudge: Controlled text generation with future discriminators
Yang, K. and Klein, D · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Cited alongside, same era.
PPL-MCTS: Constrained textual generation through discriminator-guided MCTS decoding
Chaffin, A., Claveau, V., and Kijak, E · 2022
Cited alongside, same era.
Rankgen: Improving text generation with large ranking models
Krishna, K., Chang, Y., Wieting, J., and Iyyer, M · 2022
Cited alongside, same era.
Let’s verify step by step
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Later among the works it cites.
Fine-grained human feedback gives better rewards for language model training
Wu, Z., Hu, Y., Shi, W., Dziri, N., Suhr, A., Ammanabrolu, P., Smith, N. A., Ostendorf, M., and Hajishirzi, H · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Offline RL for natural language generation with implicit language Q learning
Snell, C., Kostrikov, I., Su, Y., Yang, M., and Levine, S · 2022
Cited alongside, same era.
Solving math word problems with process-and outcome-based feedback
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Cited alongside, same era.
Naturalprover: Grounded mathematical proof generation with language models
Welleck, S., Liu, J., Lu, X., Hajishirzi, H., and Choi, Y · 2022
Cited alongside, same era.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Reward-augmented decoding: Efficient controlled text generation with a unidirectional reward model
Deng, H. and Raffel, C · 2023
Cited alongside, same era.
RAFT: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., KaShun, S., and Zhang, T · 2023
Cited alongside, same era.
Enhancing reinforcement learning with dense rewards from language model critic
Cao, M., Shu, L., Yu, L., Zhu, Y., Wichers, N., Liu, Y., and Meng, L · 2024
Later among the works it cites.
Ultrafeedback: Boosting language models with scaled ai feedback
Ganqu Cui, L. Y., Ding, N., Yao, G., He, B., Zhu, W., Ni, Y., Xie, G., Xie, R., Lin, Y., Liu, Z., and Sun, M · 2024
Later among the works it cites.
Value augmented sampling for language model alignment and personalization
Han, S., Shenfeld, I., Srivastava, A., Kim, Y., and Agrawal, P · 2024
Later among the works it cites.
Alignment as reward-guided search
Khanov, M., Burapacheep, J., and Li, Y · 2024
Later among the works it cites.
Cascade reward sampling for efficient decoding-time alignment
Li, B., Wang, Y., Grama, A., and Zhang, R · 2024
Later among the works it cites.
Controlled decoding from language models
Mudgal, S., Lee, J., Ganapathy, H., Li, Y., Wang, T., Huang, Y., Chen, Z., Cheng, H.-T., Collins, M., Strohman, T., et al · 2024
Later among the works it cites.
From r r to Q ∗ Q^{*} : Your language model is secretly a Q-function
Rafailov, R., Hejna, J., Park, R., and Finn, C · 2024
Later among the works it cites.
A critical look at tokenwise reward-guided text generation
Rashid, A., Wu, R., Grosse, J., Kristiadi, A., and Poupart, P · 2024
Later among the works it cites.
Probabilistic inference in language models via twisted sequential Monte Carlo
Zhao, S., Brekelmans, R., Makhzani, A., and Grosse, R · 2024
Later among the works it cites.