Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Convex optimization
Boyd, S. P. and Vandenberghe, L · 2004
Earlier work this paper cites.
Learning all optimal policies with multiple criteria
Barrett, L. and Narayanan, S · 2008
Earlier work this paper cites.
Ava: A large-scale database for aesthetic visual analysis
Murray, N., Marchesotti, L., and Perronnin, F · 2012
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Roijers, D. M., Vamplew, P., Whiteson, S., and Dazeley, R · 2013
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of pareto dominating policies
Van Moffaert, K. and Nowé, A · 2014
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Original
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Original
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
Human-aligned artificial intelligence is a multiobjective problem
Vamplew, P., Dazeley, R., Foale, C., Firmin, S., and Mummery, J · 2018
Earlier work this paper cites.
Curriculum-guided hindsight experience replay
Fang, M., Zhou, T., Du, Y., Han, L., and Zhang, Z · 2019
Earlier work this paper cites.
Learning to reach goals via iterated supervised learning
Original
Ghosh, D., Gupta, A., Reddy, A., Fu, J., Devin, C., Eysenbach, B., and Levine, S · 2019
Earlier work this paper cites.
Reward-conditioned policies
Original
Kumar, A., Peng, X. B., and Levine, S · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Policy continuation with hindsight inverse dynamics
Sun, H., Li, Z., Liu, X., Zhou, B., and Lin, D · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Original
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Recall and learn: Fine-tuning deep pretrained language models with less forgetting
Original
Chen, S., Hou, Y., Cui, Y., Che, W., Liu, T., and Yu, X · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Original
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.