Training a helpful and harmless assistant with reinforcement learning from human feedback
Original
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Later among the works it cites.
Predictability and surprise in large generative models
Ganguli, D., Hernandez, D., Lovitt, L., Askell, A., Bai, Y., Chen, A., Conerly, T., Dassarma, N., Drain, D., Elhage, N., et al · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Original
Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al · 2022
Later among the works it cites.
URL https://twitter.com/tomgoldsteincs/status/1600196990905614336
Goldstein, T., 2022 · 2022
Later among the works it cites.
A survey on automated fact-checking
Guo, Z., Schlichtkrull, M., and Vlachos, A · 2022
Later among the works it cites.
Opt-iml: Scaling language model instruction meta learning through the lens of generalization
Original
Iyer, S., Lin, X. V., Pasunuru, R., Mihaylov, T., Simig, D., Yu, P., Shuster, K., Wang, T., Liu, Q., Koura, P. S., et al · 2022
Later among the works it cites.
Holistic evaluation of language models
Original
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al · 2022
Later among the works it cites.
A holistic approach to undesired content detection in the real world
Original
Markov, T., Zhang, C., Agarwal, S., Eloundou, T., Lee, T., Adler, S., Jiang, A., and Weng, L · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Original
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Ignore previous prompt: Attack techniques for language models
Original
Perez, F. and Ribeiro, I · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Original
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Later among the works it cites.
Veil 3.1.x, 2022
Truncer, C · 2022
Later among the works it cites.
Benchmarking generalization via in-context instructions on 1,600+ language tasks
Original
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A. S., Naik, A., Stap, D., et al · 2022
Later among the works it cites.
Taxonomy of risks posed by language models
Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., et al · 2022
Later among the works it cites.
Templm: Distilling language models into template-based generators
Original
Zhang, T., Lee, M., Li, L., Shen, E., and Hashimoto, T. B · 2022
Later among the works it cites.