Fetching the paper…
Reading the bibliography…
We present evidence of substantial benefit from efficient exploration in gathering human feedback to improve large language models.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Information-based objective functions for active data selection
D. J. MacKay · 1992
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Planning to be surprised: Optimal Bayesian exploration in dynamic environments
Y. Sun, F. Gomez, and J. Schmidhuber · 2011
Earlier work this paper cites.
Entropy search for information-efficient global optimization
P. Hennig and C. J. Schuler · 2012
Earlier work this paper cites.
The knowledge gradient algorithm for a general class of online learning problems
I. O. Ryzhov, W. B. Powell, and P. I. Frazier · 2012
Earlier work this paper cites.
The K K -armed dueling bandits problem
Y. Yue, J. Broder, R. Kleinberg, and T. Joachims · 2012
Earlier work this paper cites.
Learning to optimize via information-directed sampling
D. Russo and B. Van Roy · 2014
Earlier work this paper cites.
Contextual dueling bandits
M. Dudík, K. Hofmann, R. E. Schapire, A. Slivkins, and M. Zoghi · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
R. Houthooft, X. Chen, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped DQN
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Earlier work this paper cites.
Double Thompson sampling for dueling bandits
H. Wu and X. Liu · 2016
Cited alongside, same era.
Ensemble Sampling
X. Lu and B. Van Roy · 2017
Cited alongside, same era.
Count-based exploration with neural density models
G. Ostrovski, M. G. Bellemare, A. Oord, and R. Munos · 2017
Cited alongside, same era.
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2018
Cited alongside, same era.
C. Riquelme, G. Tucker, and J. Snoek · 2018
Cited alongside, same era.
A Tutorial on Thompson Sampling
D. J. Russo, B. Van Roy, A. Kazerouni, I. Osband, and Z. Wen · 2018
Neural contextual bandits with ucb-based exploration
D. Zhou, L. Li, and Q. Gu · 2020
Later among the works it cites.
Deciding what to learn: A rate-distortion approach
D. Arumugam and B. Van Roy · 2021
Later among the works it cites.
Optimal algorithms for stochastic contextual preference bandits
A. Saha · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
A. Glaese, N. McAleese, M. Trebacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Planning for cars that coordinate with people: Leveraging effects on human actions for planning and active information gathering over human internal state
D. Sadigh, N. Landolfi, S. S. Sastry, S. A. Seshia, and A. D. Dragan · 2018
Cited alongside, same era.
Deep exploration via randomized value functions
I. Osband, B. Van Roy, D. J. Russo, and Z. Wen · 2019
Cited alongside, same era.
Never give up: Learning directed exploration strategies
A. P. Badia, P. Sprechmann, A. Vitvitskyi, D. Guo, B. Piot, S. Kapturowski, O. Tieleman, M. Arjovsky, A. Pritzel, A. Bolt, et al · 2020
Cited alongside, same era.
Hypermodels for exploration
V. Dwaracherla, X. Lu, M. Ibrahimi, I. Osband, Z. Wen, and B. Van Roy · 2020
Cited alongside, same era.
Bandit Algorithms
T. Lattimore and C. Szepesvári · 2020
Cited alongside, same era.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
An empirical analysis of compute-optimal large language model training
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, O. Vinyals, J. Rae, and L. Sifre · 2022
Later among the works it cites.
ChatGPT: Optimizing Language Models for Dialogue, 2022
OpenAI · 2022
Later among the works it cites.
Evaluating high-order predictive distributions in deep learning
I. Osband, Z. Wen, S. M. Asghari, V. Dwaracherla, X. Lu, and B. Van Roy · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe · 2022
Later among the works it cites.
Palm 2 technical report, 2023
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier-Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G. H. Abrego, J. Ahn, J. Austin, P. Barham, J. Botha, J. Bradbury, S. Brahma, K. Brooks, M. Catasta, Y. Cheng, C. Cherry, C. A. Choquette-Choo, A. Chowdhery, C. Crepy, S. Dave, M. Dehghani, S. Dev, J. Devlin, M. Díaz, N. Du, E. Dyer, V. Feinberg, F. Feng, V. Fienber, M. Freitag, X. Garcia, S. Gehrmann, L. Gonzalez, G. Gur-Ari, S. Hand, H. Hashemi, L. Hou, J. Howland, A. Hu, J. Hui, J. Hurwitz, M. Isard, A. Ittycheriah, M. Jagielski, W. Jia, K. Kenealy, M. Krikun, S. Kudugunta, C. Lan, K. Lee, B. Lee, E. Li, M. Li, W. Li, Y. Li, J. Li, H. Lim, H. Lin, Z. Liu, F. Liu, M. Maggioni, A. Mahendru, J. Maynez, V. Misra, M. Moussalem, Z. Nado, J. Nham, E. Ni, A. Nystrom, A. Parrish, M. Pellat, M. Polacek, A. Polozov, R. Pope, S. Qiao, E. Reif, B. Richter, P. Riley, A. C. Ros, A. Roy, B. Saeta, R. Samuel, R. Shelby, A. Slone, D. Smilkov, D. R. So, D. Sohn, S. Tokumine, D. Valter, V. Vasudevan, K. Vodrahalli, X. Wang, P. Wang, Z. Wang, T. Wang, J. Wieting, Y. Wu, K. Xu, Y. Xu, L. Xue, P. Yin, J. Yu, Q. Zhang, S. Zheng, C. Zheng, W. Zhou, D. Zhou, S. Petrov, and Y. Wu · 2023
Later among the works it cites.
GPT-4 Technical Report
OpenAI · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models, 2023
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.