Fetching the paper…
Reading the bibliography…
Aligning large language models (LLMs) with human values is a vital task for LLM practitioners.
Language models are few-shot learners
Brown, T · 1901
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D · 2022
Earlier work this paper cites.
Large language models can self-improve
Huang, J · 2022
Earlier work this paper cites.
Unsupervised cross-task generalization via retrieval augmentation
Lin, B. Y · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L · 2022
Cited alongside, same era.
Self-instruct: Aligning language model with self generated instructions
Wang, Y · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J · 2022
Cited alongside, same era.
Wu, Y · 2022
Cited alongside, same era.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Ji, J · 2023
Later among the works it cites.
Longform: Optimizing instruction tuning for long text generation with corpus extraction
Köksal, A · 2023
Later among the works it cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Lee, H · 2023
Later among the works it cites.
The unlocking spell on base llms: Rethinking alignment via in-context learning
Lin, B. Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, S · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S · 2023
Cited alongside, same era.
Reinforced self-training (rest) for language modeling
Gulcehre, C · 2023
Cited alongside, same era.
In-context alignment: Chat with vanilla language models before fine-tuning
Han, X · 2023
Cited alongside, same era.
Self-alignment with instruction backtranslation
Li, X
Cited in the paper.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X
Cited in the paper.
Rain: Your language models can align themselves without finetuning
Li, Y
Cited in the paper.
Liu, Z · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Later among the works it cites.
Baize: An open-source chat model with parameter-efficient tuning on self-chat data
Xu, C · 2023
Later among the works it cites.
Self-qa: Unsupervised knowledge guided language model alignment
Zhang, X · 2023
Later among the works it cites.
Lima: Less is more for alignment
Zhou, C · 2023
Later among the works it cites.