Fetching the paper…
Reading the bibliography…
Chain-of-thought distillation is a powerful technique for transferring reasoning abilities from large language models (LLMs) to smaller student models.
Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters
Bridle, J · 1989
Earlier work this paper cites.
Learning and development in neural networks: The importance of starting small
Elman, J. L · 1993
Earlier work this paper cites.
The generalized sigmoid activation function: Competitive supervised learning
Narayan, S · 1997
Earlier work this paper cites.
Modulation of the palmar grasp behavior in neonates according to texture property
Molina, M. and Jouen, F · 1998
Earlier work this paper cites.
Constrained k-means clustering
Bradley, P. S., Bennett, K. P., and Demiriz, A · 2000
Earlier work this paper cites.
A day of great illumination: Bf skinner’s discovery of shaping
Peterson, G. B · 2004
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Flexible shaping: How learning in small steps helps
Krueger, K. A. and Dayan, P · 2009
Earlier work this paper cites.
Young children’s mapping between arrays, number words, and digits
Benoit, L., Lehalle, H., Molina, M., Tijus, C., and Jouen, F · 2013
Earlier work this paper cites.
Self-paced learning with diversity
Jiang, L., Meng, D., Yu, S.-I., Lan, Z., Shan, S., and Hauptmann, A · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Ling, W., Yogatama, D., Dyer, C., and Blunsom, P · 2017
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Talmor, A., Herzig, J., Lourie, N., and Berant, J · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Self-paced learning for neural machine translation
Wan, Y., Yang, B., Wong, D. F., Zhou, Y., Chao, L. S., Zhang, H., and Chen, B · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Adaptive curriculum learning
Kong, Y., Liu, L., Wang, J., and Tao, D · 2021
Cited alongside, same era.
Token-wise curriculum learning for neural machine translation
Liang, C., Jiang, H., Liu, X., He, P., Chen, W., Gao, J., and Zhao, T · 2021
Cited alongside, same era.
MCC-KD: Multi-CoT consistent knowledge distillation
Chen, H., Wu, S., Quan, X., Wang, R., Yan, M., and Zhang, J · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Later among the works it cites.
Active prompting with chain-of-thought for large language models
Diao, S., Wang, P., Lin, Y., and Zhang, T · 2023
Later among the works it cites.
Specializing smaller language models towards multi-step reasoning
Fu, Y., Peng, H., Ou, L., Sabharwal, A., and Khot, T · 2023
Later among the works it cites.
Large language models are reasoning teachers
Ho, N., Schmid, L., and Yun, S.-Y · 2023
Later among the works it cites.
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Miao, S.-Y., Liang, C.-C., and Su, K.-Y · 2021
Cited alongside, same era.
Are nlp models really able to solve simple math word problems?
Patel, A., Bhattamishra, S., and Goyal, N · 2021
Cited alongside, same era.
Curgraph: Curriculum learning for graph classification
Wang, Y., Wang, W., Liang, Y., Cai, Y., and Hooi, B · 2021
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Cited alongside, same era.
Freekd: Free-direction knowledge distillation for graph neural networks
Feng, K., Li, C., Yuan, Y., and Wang, G · 2022
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Cited alongside, same era.
Hsieh, C.-Y., Li, C.-L., Yeh, C.-k., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C.-Y., and Pfister, T · 2023
Later among the works it cites.
In-sample curriculum learning by sequence completion for natural language generation
Jia, Q., Liu, Y., Tang, H., and Zhu, K · 2023
Later among the works it cites.
Symbolic chain-of-thought distillation: Small models can also “think” step-by-step
Li, L. H., Hessel, J., Yu, Y., Ren, X., Chang, K.-W., and Choi, Y · 2023
Later among the works it cites.
Teaching small language models to reason
Magister, L. C., Mallinson, J., Adamek, J., Malmi, E., and Severyn, A · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Cue-cot: Chain-of-thought prompting for responding to in-depth dialogue questions with llms
Wang, H., Wang, R., Mi, F., Deng, Y., Wang, Z., Liang, B., Xu, R., and Wong, K.-F · 2023
Later among the works it cites.
Large language models as optimizers
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X · 2023
Later among the works it cites.
A comprehensive capability analysis of gpt-3 and gpt-3.5 series models
Ye, J., Chen, X., Xu, N., Zu, C., Shao, Z., Liu, S., Cui, Y., Zhou, Z., Gong, C., Shen, Y., et al · 2023
Later among the works it cites.
Multimodal chain-of-thought reasoning in language models
Zhang, Z., Zhang, A., Li, M., Zhao, H., Karypis, G., and Smola, A · 2023
Later among the works it cites.
On the road to portability: Compressing end-to-end motion planner for autonomous driving
Feng, K., Li, C., Ren, D., Yuan, Y., and Wang, G · 2024
Closest in time.