Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) has emerged as a powerful technique to make large language models (LLMs) easier to use and more effective.
Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A.I., Babaei, H., LeJeune, D.Baraniuk, R.G · 2023
Earlier work this paper cites.
Gudibande, A., Wallace, E., Snell, C., Geng, X., Liu, H., Abbeel, P.Song, D · 2023
Cited alongside, same era.
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N. & Anderson, R · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…