Fetching the paper…
Reading the bibliography…
Assessing response quality to instructions in language models is vital but challenging due to the complexity of human language across different contexts.
A mathematical theory of communication
Shannon, C. E · 1948
Earlier work this paper cites.
Philosophical Investigations
Wittgenstein, L · 1953
Earlier work this paper cites.
On a measure of the information provided by an experiment
Lindley, D. V · 1956
Earlier work this paper cites.
Aleatory or epistemic? does it matter?
Der Kiureghian, A. and Ditlevsen, O · 2009
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning, 2016
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Deep bayesian active learning with image data
Gal, Y., Islam, R., and Ghahramani, Z · 2017
Earlier work this paper cites.
What uncertainties do we need in bayesian deep learning for computer vision?
Kendall, A. and Gal, Y · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Sadigh, D., Dragan, A. D., Sastry, S., and Seshia, S. A · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Test-time data augmentation for estimation of heteroscedastic aleatoric uncertainty in deep neural networks
Ayhan, M. S. and Berens, P · 2018
Earlier work this paper cites.
Batch active preference-based learning of reward functions
Biyik, E. and Sadigh, D · 2018
Earlier work this paper cites.
Decomposition of uncertainty in bayesian deep learning for efficient and risk-sensitive learning
Depeweg, S., Hernandez-Lobato, J.-M., Doshi-Velez, F., and Udluft, S · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction, 2018
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Earlier work this paper cites.
Curriculum learning for natural answer generation
Liu, C., He, S., Liu, K., Zhao, J., et al · 2018
Earlier work this paper cites.
Unsupervised visual domain adaptation: A deep max-margin gaussian process approach
Kim, M., Sahu, P., Gholami, B., and Pavlovic, V · 2019
Earlier work this paper cites.
Direct uncertainty prediction for medical second opinions
Raghu, M., Blumer, K., Sayres, R., Obermeyer, Z., Kleinberg, B., Mullainathan, S., and Kleinberg, J · 2019
Earlier work this paper cites.
Active preference-based gaussian process regression for reward learning
Biyik, E., Huynh, N., Kochenderfer, M. J., and Sadigh, D · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
DIME: An information-theoretic difficulty measure for AI datasets
Zhang, P., Wang, H., Naik, N., Xiong, C., and richard socher · 2020
Earlier work this paper cites.
Fine-tuning language models from human preferences, 2020
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2020
Earlier work this paper cites.
Curriculum learning for language modeling
Campos, D · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
Lee, K., Smith, L., and Abbeel, P · 2021
Cited alongside, same era.
Li, C., Zhang, M., and He, Y · 2021
Cited alongside, same era.
Pre-training a bert with curriculum learning by increasing block-size of input text
Nagatsuka, K., Broni-Bediako, C., and Atsumi, M · 2021
Cited alongside, same era.
Review and arrange: Curriculum learning for natural language understanding
Zhang, L., Mao, Z., Xu, B., Wang, Q., and Zhang, Y · 2021
Cited alongside, same era.
Scaling laws for reward model overoptimization, 2022
Gao, L., Schulman, J., and Hilton, J · 2022
Cited alongside, same era.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset, 2023
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Zhang, C., Sun, R., Wang, Y., and Yang, Y · 2023
Later among the works it cites.
Unsupervised domain adaptation based on the predictive uncertainty of models
Lee, J. and Lee, G · 2023
Later among the works it cites.
Unsupervised accuracy estimation of deep visual models using domain-adaptive adversarial perturbation without source samples
Lee, J., Woo, J. O., Moon, H., and Lee, K · 2023
Later among the works it cites.
Policy optimization in rlhf: The impact of out-of-preference data
Li, Z., Xu, T., and Yu, Y · 2023
Later among the works it cites.
Openorca: An open dataset of gpt augmented flan reasoning traces
Lian, W., Goodson, B., Pentland, E., Cook, A., Vong, C., and "Teknium" · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving alignment of dialogue agents via targeted human judgements, 2022
Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G · 2022
Cited alongside, same era.
Efficient pre-training of masked language model via concept-based curriculum masking
Lee, M., Park, J.-H., Kim, J., Kim, K.-M., and Lee, S · 2022
Cited alongside, same era.
Reward uncertainty for exploration in preference-based reinforcement learning
Liang, X., Shu, K., Lee, K., and Abbeel, P · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y · 2022
Cited alongside, same era.
Curriculum learning: A survey
Soviany, P., Ionescu, R. T., Rota, P., and Sebe, N · 2022
Cited alongside, same era.
Learning to summarize from human feedback, 2022
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2022
Cited alongside, same era.
Orca: Progressive learning from complex explanation traces of gpt-4, 2023
Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., and Awadallah, A · 2023
Later among the works it cites.
GPT-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Are all languages equal? curriculum learning over different languages
Pucci, G., Ranaldi, L., Zanzotto, F. M., et al · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model, 2023
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Modeling easiness for training transformers with curriculum learning
Ranaldi, L., Pucci, G., and Zanzotto, F. M · 2023
Later among the works it cites.
Sharegpt dataset, 2023
RyokoAI · 2023
Later among the works it cites.
Preference ranking optimization for human alignment, 2023
Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model, 2023
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Later among the works it cites.
Openchat: Advancing open-source language models with mixed-quality data, 2023
Wang, G., Cheng, S., Zhan, X., Li, X., Song, S., and Liu, Y · 2023
Later among the works it cites.
Curriculum learning with adam: The devil is in the wrong details
Weber, L., Jumelet, J., Michel, P., Bruni, E., and Hupkes, D · 2023
Later among the works it cites.
Active learning in bayesian neural networks with balanced entropy learning principle
Woo, J. O · 2023
Later among the works it cites.
Contrastive post-training large language models on data curriculum
Xu, C., Rosset, C., Del Corro, L., Mahajan, S., McAuley, J., Neville, J., Awadallah, A. H., and Rao, N · 2023
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Later among the works it cites.
Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles, 2023
Zhai, Y., Zhang, H., Lei, Y., Yu, Y., Xu, K., Feng, D., Ding, B., and Wang, H · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
Self-play fine-tuning converts weak language models to strong language models
Chen, Z., Deng, Y., Yuan, H., Ji, K., and Gu, Q · 2024
Closest in time.
Reward model ensembles help mitigate overoptimization
Coste, T., Anwar, U., Kirk, R., and Krueger, D · 2024
Closest in time.