Fetching the paper…
Reading the bibliography…
Language models often pre-train on large unsupervised text corpora, then fine-tune on additional task-specific data.
Query by committee
Seung, H. S., Opper, M., and Sompolinsky, H · 1992
Earlier work this paper cites.
A sequential algorithm for training text classifiers: Corrigendum and additional data
Lewis, D. D · 1995
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E · 2001
Earlier work this paper cites.
Margin-based active learning for structured output spaces
Roth, D. and Small, K · 2006
Earlier work this paper cites.
An analysis of active learning strategies for sequence labeling tasks
Settles, B. and Craven, M · 2008
Earlier work this paper cites.
Active learning literature survey
Settles, B · 2009
Earlier work this paper cites.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. and Ba, J · 2015
Earlier work this paper cites.
Bootstrapped Thompson sampling and deep exploration
Osband, I. and Van Roy, B · 2015
Earlier work this paper cites.
How much does your data exploration overfit? controlling bias via information usage
Russo, D. and Zou, J · 2015
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Risk versus uncertainty in deep learning: Bayes, bootstrap and the dangers of dropout
Osband, I · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Deep Bayesian active learning with image data
Gal, Y., Islam, R., and Ghahramani, Z · 2017
Earlier work this paper cites.
Variational Gaussian dropout is not Bayesian
Hron, J., Matthews, A. G. d. G., and Ghahramani, Z · 2017
Earlier work this paper cites.
What uncertainties do we need in Bayesian deep learning for computer vision?
Kendall, A. and Gal, Y · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Sadigh, D., Dragan, A. D., Sastry, S., and Seshia, S. A · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. R · 2017
Cited alongside, same era.
The power of ensembles for active learning in image classification
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Later among the works it cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beluch, W. H., Genewein, T., Nürnberger, A., and Köhler, J. M · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A · 2018
Cited alongside, same era.
Do cifar-10 classifiers generalize to cifar-10?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning
Kirsch, A., Van Amersfoort, J., and Gal, Y · 2019
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021
Fedus, W., Zoph, B., and Shazeer, N · 2021
Later among the works it cites.
Reinforcement learning, bit by bit
Lu, X., Van Roy, B., Dwaracherla, V., Ibrahimi, M., Osband, I., and Wen, Z · 2021
Later among the works it cites.
Osband, I., Wen, Z., Asghari, M., Ibrahimi, M., Lu, X., and Van Roy, B · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Closest in time.
Uncertainty estimation for language reward models
Gleave, A. and Irving, G · 2022
Closest in time.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Closest in time.
Language models (mostly) know what they know
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Dodds, Z. H., DasSarma, N., Tran-Johnson, E., et al · 2022
Closest in time.
On the importance of effectively adapting pretrained language models for active learning
Margatina, K., Barrault, L., and Aletras, N · 2022
Closest in time.
The neural testbed: Evaluating joint predictions
Osband, I., Wen, Z., Asghari, S. M., Dwaracherla, V., Hao, B., Ibrahimi, M., Lawson, D., Lu, X., O’Donoghue, B., and Van Roy, B · 2022
Closest in time.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Closest in time.
From predictions to decisions: The importance of joint predictive distributions, 2022
Wen, Z., Osband, I., Qin, C., Lu, X., Ibrahimi, M., Dwaracherla, V., Asghari, M., and Van Roy, B · 2022
Closest in time.