On the partition function and random maximum a-posteriori perturbations,
T. Hazan, T. S. Jaakkola, · 2012
Earlier work this paper cites.
Semeval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning,
A. Gordon, Z. Kozareva, M. Roemmele, · 2012
Earlier work this paper cites.
Neural variational inference and learning in belief networks,
A. Mnih, K. Gregor, · 2014
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables,
C. J. Maddison, A. Mnih, Y. W. Teh, · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks,
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al., · 2017
Earlier work this paper cites.
Universal language model fine-tuning for text classification,
J. Howard, S. Ruder, · 2018
Earlier work this paper cites.
Xnli: Evaluating cross-lingual sentence representations,
A. Conneau, R. Rinott, G. Lample, A. Williams, S. Bowman, H. Schwenk, V. Stoyanov, · 2018
Earlier work this paper cites.
Xcopa: A multilingual dataset for causal commonsense reasoning,
E. M. Ponti, G. Glavaš, O. Majewska, Q. Liu, I. Vulić, A. Korhonen, · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment,
Original
A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, et al, · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models,
E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al., · 2021
Earlier work this paper cites.
A survey on multi-task learning,
Y. Zhang, Q. Yang, · 2021
Earlier work this paper cites.
Efficiently identifying task groupings for multi-task learning,
C. Fifty, E. Amid, Z. Zhao, T. Yu, R. Anil, C. Finn, · 2021
Earlier work this paper cites.
Evaluating large language models trained on code,
Original
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al., · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems,
Original
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al., · 2021
Earlier work this paper cites.
Glm-130b: An open bilingual pre-trained model,
A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y. Xu, W. Zheng, X. Xia, et al., · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback,
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, e. a. Zhang, · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback,
Original
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, DasSarma, et al., · 2022
Earlier work this paper cites.
Dataless knowledge fusion by merging weights of language models,
X. Jin, X. Ren, D. Preotiuc-Pietro, P. Cheng, · 2022
Earlier work this paper cites.