Fetching the paper…
Reading the bibliography…
Since the release of T\"ULU [Wang et al., 2023b], open resources for instruction tuning have developed quickly, from better base models to new finetuning techniques.
Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction
Y. Luan, L. He, M. Ostendorf, and H. Hajishirzi · 2018
Earlier work this paper cites.
Inferring which medical treatments work from reports of clinical trials
E. Lehman, J. DeYoung, R. Barzilay, and B. C. Wallace · 2019
Earlier work this paper cites.
TLDR: Extreme summarization of scientific documents
I. Cachola, K. Lo, A. Cohan, and D. Weld · 2020
Earlier work this paper cites.
Fact or fiction: Verifying scientific claims
D. Wadden, S. Lin, K. Lo, L. L. Wang, M. van Zuylen, A. Cohan, and H. Hajishirzi · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers
P. Dasigi, K. Lo, I. Beltagy, A. Cohan, N. A. Smith, and M. Gardner · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
B. Lester, R. Al-Rfou, and N. Constant · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2022
Earlier work this paper cites.
TOXIGEN: Controlling Language Models to Generate Implied and Adversarial Toxicity
T. Hartvigsen, S. Gabriel, H. Palangi, M. Sap, D. Ray, and E. Kamar · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Earlier work this paper cites.
Training Language Models to Follow Instructions with Human Feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Offline rl for natural language generation with implicit language q learning
C. Snell, I. Kostrikov, Y. Su, M. Yang, and S. Levine · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Cited alongside, same era.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al · 2023
Cited alongside, same era.
Code alpaca: An instruction-following llama model for code generation
S. Chaudhary · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Cited alongside, same era.
Ultrafeedback: Boosting language models with high-quality feedback
G. Cui, L. Yuan, N. Ding, G. Yao, W. Zhu, Y. Ni, G. Xie, Z. Liu, and M. Sun · 2023
Openorca: An open dataset of gpt augmented flan reasoning traces
W. Lian, B. Goodson, E. Pentland, A. Cook, C. Vong, and "Teknium" · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, commercially usable llms
MosaicML · 2023
Closest in time.
Orca: Progressive learning from complex explanation traces of gpt-4, 2023
S. Mukherjee, A. Mitra, G. Jawahar, S. Agarwal, H. Palangi, and A. Awadallah · 2023
Closest in time.
B. Peng, C. Li, P. He, M. Galley, and J. Gao · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Free dolly: Introducing the world’s first truly open instruction-tuned llm
Databricks · 2023
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer · 2023
Cited alongside, same era.
Ultrachat: A large-scale auto-generated multi-round dialogue data
N. Ding, Y. Chen, B. Xu, S. Hu, Y. Qin, Z. Liu, M. Sun, and B. Zhou · 2023
Cited alongside, same era.
Alpacafarm: A simulation framework for methods that learn from human feedback
Y. Dubois, X. Li, R. Taori, T. Zhang, I. Gulrajani, J. Ba, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Cited alongside, same era.
Easylm: A simple and scalable training framework for large language models, 2023
X. Geng · 2023
Cited alongside, same era.
Reinforced self-training (rest) for language modeling
C. Gulcehre, T. L. Paine, S. Srinivasan, K. Konyushkova, L. Weerts, A. Sharma, A. Siddhant, A. Ahern, M. Wang, C. Gu, et al · 2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Cited alongside, same era.
B. Roziere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, T. Remez, J. Rapin, et al · 2023
Closest in time.
Efficient rlhf: Reducing the memory usage of ppo, 2023
M. Santacroce, Y. Lu, H. Yu, Y. Li, and Y. Shen · 2023
Closest in time.
A long way to go: Investigating length correlations in rlhf
P. Singhal, T. Goyal, J. Xu, and G. Durrett · 2023
Closest in time.
Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of rlhf, 2023
S. Sun, D. Gupta, and M. Iyyer · 2023
Closest in time.
Zephyr: Direct distillation of lm alignment
L. Tunstall, E. Beeching, N. Lambert, N. Rajani, K. Rasul, Y. Belkada, S. Huang, L. von Werra, C. Fourrier, N. Habib, et al · 2023
Closest in time.
Wizardlm: Empowering large language models to follow complex instructions
C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, and D. Jiang · 2023
Closest in time.
Xwin-lm, 2023
Xwin-LM Team · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Closest in time.
Lima: Less is more for alignment
C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y. Mao, X. Ma, A. Efrat, P. Yu, L. Yu, et al · 2023
Closest in time.