Fetching the paper…
Reading the bibliography…
Chain-of-thought (CoT) decoding enables language models to improve reasoning performance at the cost of high generation latency in decoding.
Burtsev, M. S., Kuratov, Y., Peganov, A., and Sapunov, G. V · 2006
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Implicit chain of thought reasoning via knowledge distillation, 2023
Deng, Y., Prasad, K., Fernandez, R., Smolensky, P., Chaudhary, V., and Shieber, S · 2023
Earlier work this paper cites.
Llmlingua: Compressing prompts for accelerated inference of large language models, 2023
Jiang, H., Wu, Q., Lin, C.-Y., Yang, Y., and Qiu, L · 2023
Earlier work this paper cites.
Large language models are zero-shot reasoners, 2023
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2023
Earlier work this paper cites.
Measuring faithfulness in chain-of-thought reasoning, 2023
Lanham, T., Chen, A., Radhakrishnan, A., Steiner, B., Denison, C., Hernandez, D., Li, D., Durmus, E., Hubinger, E., Kernion, J., Lukošiūtė, K., Nguyen, K., Cheng, N., Joseph, N., Schiefer, N., Rausch, O., Larson, R., McCandlish, S., Kundu, S., Kadavath, S., Yang, S., Henighan, T., Maxwell, T., Telleen-Lawton, T., Hume, T., Hatfield-Dodds, Z., Kaplan, J., Brauner, J., Bowman, S. R., and Perez, E · 2023
Earlier work this paper cites.
Chain of hindsight aligns language models with feedback, 2023
Liu, H., Sferrazza, C., and Abbeel, P · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Cited alongside, same era.
Attention is all you need, 2023
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2023
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2023
Cited alongside, same era.
Thinking tokens for language modeling, 2024
Herel, D. and Mikolov, T · 2024
Closest in time.
Cllms: Consistency large language models, 2024
Kou, S., Hu, L., He, Z., Deng, Z., and Zhang, H · 2024
Closest in time.
Cachegen: Kv cache compression and streaming for fast large language model serving
Liu, Y., Li, H., Cheng, Y., Ray, S., Huang, Y., Zhang, Q., Du, K., Yao, J., Lu, S., Ananthanarayanan, G., Maire, M., Hoffmann, H., Holtzman, A., and Jiang, J · 2024
Closest in time.
Skeleton-of-thought: Prompting llms for efficient parallel generation, 2024
Ning, X., Lin, Z., Zhou, Z., Wang, Z., Yang, H., and Wang, Y · 2024
Closest in time.
Let’s think dot by dot: Hidden computation in transformer language models, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Cited alongside, same era.
From explicit cot to implicit cot: Learning to internalize cot step by step, 2024
Deng, Y., Choi, Y., and Shieber, S · 2024
Cited alongside, same era.
In-context autoencoder for context compression in a large language model, 2024
Ge, T., Hu, J., Wang, L., Wang, X., Chen, S.-Q., and Wei, F · 2024
Cited alongside, same era.
Think before you speak: Training language models with pause tokens, 2024
Goyal, S., Ji, Z., Rawat, A. S., Menon, A. K., Kumar, S., and Nagarajan, V · 2024
Cited alongside, same era.
Training large language models to reason in a continuous latent space, 2024
Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., and Tian, Y · 2024
Cited alongside, same era.
Pfau, J., Merrill, W., and Bowman, S. R · 2024
Closest in time.
Puerto, H., Chubakov, T., Zhu, X., Madabushi, H. T., and Gurevych, I · 2024
Closest in time.
Dodo: Dynamic contextual compression for decoder-only lms, 2024
Qin, G., Rosset, C., Chau, E. C., Rao, N., and Durme, B. V · 2024
Closest in time.
Fast chain-of-thought: A glance of future from parallel decoding leads to answers faster, 2024
Zhang, H., Liu, Z., Zhao, Y., Zheng, J., Zhuang, C., Gu, J., and Chen, G · 2024
Closest in time.