Fetching the paper…
Reading the bibliography…
Given the increasing use of synthetic data in language model (LM) post-training, an LM's ability to generate high-quality data has become nearly as crucial as its ability to solve problems directly.
Instruction induction: From few examples to natural language task descriptions
Honovich, O., Shaham, U., Bowman, S., and Levy, O · 1952
Earlier work this paper cites.
Learning to mine aligned code and natural language pairs from stack overflow
Yin, P., Deng, B., Chen, E., Vasilescu, B., and Neubig, G · 2018
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Earlier work this paper cites.
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H · 2022
Earlier work this paper cites.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Naik, A., Ashok, A., Dhanasekaran, A. S., Arunkumar, A., Stap, D., et al · 2022
Earlier work this paper cites.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2022
Earlier work this paper cites.
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Del Giorno, A., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., et al · 2023
Earlier work this paper cites.
Evaluating large language models: A comprehensive survey
Guo, Z., Jin, R., Liu, C., Huang, Y., Shi, D., Yu, L., Liu, Y., Li, J., Xiong, B., Xiong, D., et al · 2023
Earlier work this paper cites.
The cot collection: Improving zero-shot and few-shot learning of language models via chain-of-thought fine-tuning
Kim, S., Joo, S., Kim, D., Jang, J., Ye, S., Shin, J., and Seo, M · 2023
Earlier work this paper cites.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Cited alongside, same era.
Orca: Progressive learning from complex explanation traces of gpt-4
Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., and Awadallah, A · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Prompt2model: Generating deployable models from natural language instructions
Viswanathan, V., Zhao, C., Bertsch, A., Wu, T., and Neubig, G · 2023
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2023
Cited alongside, same era.
Rewardbench: Evaluating reward models for language modeling
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., et al · 2024
Closest in time.
From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline
Li, T., Chiang, W.-L., Frick, E., Dunlap, L., Wu, T., Zhu, B., Gonzalez, J. E., and Stoica, I · 2024
Closest in time.
Universal and context-independent triggers for precise control of llm outputs
Liang, J., Li, G., and Yu, Y · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
MetaAI · 2024
Closest in time.
Mixeval: Deriving wisdom of the crowd from LLM benchmark mixtures
Ni, J., Xue, F., Yue, X., Deng, Y., Shah, M., Jain, K., Neubig, G., and You, Y · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mammoth: Building math generalist models through hybrid instruction tuning
Yue, X., Qu, X., Zhang, G., Fu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2023
Cited alongside, same era.
Claude 3.5 sonnet model card addendum
Anthropic, A · 2024
Cited alongside, same era.
A survey on evaluation of large language models
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., et al · 2024
Cited alongside, same era.
How abilities in large language models are affected by supervised fine-tuning data composition
Dong, G., Yuan, H., Lu, K., Li, C., Xue, M., Liu, D., Wang, W., Yuan, Z., Zhou, C., and Zhou, J · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Length-controlled alpacaeval: A simple debiasing of automatic evaluators
Dubois, Y., Liang, P., and Hashimoto, T · 2024
Cited alongside, same era.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Cited alongside, same era.
Closest in time.
Leverage the Latest Open Models for Synthetic Data Generation with NVIDIA Nemotron-4-340B
Nvidia · 2024
Closest in time.
Observational scaling laws and the predictability of langauge model performance
Ruan, Y., Maddison, C. J., and Hashimoto, T · 2024
Closest in time.
Structuredrag: Json response formatting with large language models
Shorten, C., Pierse, C., Smith, T. B., Cardenas, E., Sharma, A., Trengrove, J., and van Luijt, B · 2024
Closest in time.
Tam, Z. R., Wu, C.-K., Tsai, Y.-L., Lin, C.-Y., Lee, H.-y., and Chen, Y.-N · 2024
Closest in time.
Qwen2.5: A party of foundation models, September 2024
Team, Q · 2024
Closest in time.
MAmmoTH2: Scaling instructions from the web
Yue, X., Zheng, T., Zhang, G., and Chen, W · 2024
Closest in time.
Unveiling the impact of coding data instruction fine-tuning on large language models reasoning
Zhang, X., Chen, Z. Z., Ye, X., Yang, X., Chen, L., Wang, W. Y., and Petzold, L. R · 2025
Closest in time.
Mteb: Massive text embedding benchmark
Muennighoff, N., Tazi, N., Magne, L., and Reimers, N · 2037
Closest in time.