Fetching the paper…
Reading the bibliography…
Techniques enabling large language models (LLMs) to "think more" by generating and attending to intermediate reasoning steps have shown promise in solving complex problems.
Learning feed-forward one-shot learners
L. Bertinetto, J. F. Henriques, J. Valmadre, P. H. S. Torr, and A. Vedaldi · 2016
Earlier work this paper cites.
Modulating early visual processing by language
H. de Vries, F. Strub, J. Mary, H. Larochelle, O. Pietquin, and A. C. Courville · 2017
Earlier work this paper cites.
Hypernetworks
D. Ha, A. M. Dai, and Q. V. Le · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
S. Rebuffi, H. Bilen, and A. Vedaldi · 2017
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. C. Courville · 2018
Earlier work this paper cites.
Contextual parameter generation for universal neural machine translation
E. A. Platanios, M. Sachan, G. Neubig, and T. Mitchell · 2018
Earlier work this paper cites.
Efficient parametrization of multi-domain deep neural networks
S. Rebuffi, H. Bilen, and A. Vedaldi · 2018
Earlier work this paper cites.
On self modulation for generative adversarial networks
T. Chen, M. Lučić, N. Houlsby, and S. Gelly · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Earlier work this paper cites.
UDapter: Language adaptation for truly Universal Dependency parsing
A. Üstün, A. Bisazza, G. Bouma, and G. van Noord · 2020
Earlier work this paper cites.
MAD-G: Multilingual adapter generation for efficient cross-lingual transfer
A. Ansell, E. M. Ponti, J. Pfeiffer, S. Ruder, G. Glavaš, I. Vulić, and A. Korhonen · 2021
Earlier work this paper cites.
Recurrent independent mechanisms
A. Goyal, A. Lamb, J. Hoffmann, S. Sodhani, S. Levine, Y. Bengio, and B. Schölkopf · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
B. Lester, R. Al-Rfou, and N. Constant · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
X. L. Li and P. Liang · 2021
Earlier work this paper cites.
Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
R. K. Mahabadi, S. Ruder, M. Dehghani, and J. Henderson · 2021
Cited alongside, same era.
Parameter space factorization for zero-shot learning across tasks and languages
E. M. Ponti, I. Vulić, R. Cotterell, M. Parovic, R. Reichart, and A. Korhonen · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Cited alongside, same era.
Pali: A jointly-scaled multilingual language-image model
X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, et al · 2022
Cited alongside, same era.
Hyperprompt: Prompt-based task-conditioning of transformers
Y. He, H. S. Zheng, Y. Tay, J. P. Gupta, Y. Du, V. Aribandi, Z. Zhao, Y. Li, Z. Chen, D. Metzler, H. Cheng, and E. H. Chi · 2022
Cited alongside, same era.
Least-to-most prompting enables complex reasoning in large language models
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. V. Le, and E. H. Chi · 2023
Later among the works it cites.
Llm augmented llms: Expanding capabilities through composition
R. Bansal, B. Samanta, S. Dalmia, N. Gupta, S. Vashishth, S. Ganapathy, A. Bapna, P. Jain, and P. Talukdar · 2024
Closest in time.
Hopping too late: Exploring the limitations of large language models on multi-hop queries
E. Biran, D. Gottesman, S. Yang, M. Geva, and A. Globerson · 2024
Closest in time.
From explicit cot to implicit cot: Learning to internalize cot step by step
Y. Deng, Y. Choi, and S. Shieber · 2024
Closest in time.
In-context autoencoder for context compression in a large language model
T. Ge, H. Jing, L. Wang, X. Wang, S.-Q. Chen, and F. Wei · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Cited alongside, same era.
Confident adaptive language modeling
T. Schuster, A. Fisch, J. Gupta, M. Dehghani, D. Bahri, V. Tran, Y. Tay, and D. Metzler · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Coca: Contrastive captioners are image-text foundation models. arxiv 2022
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. Goodman · 2022
Cited alongside, same era.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Cited alongside, same era.
Closest in time.
Training large language models to reason in a continuous latent space, 2024
S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y. Tian · 2024
Closest in time.
Training chain-of-thought via latent-variable inference
M. D. Hoffman, D. Phan, D. Dohan, S. Douglas, T. A. Le, A. Parisi, P. Sountsov, C. Sutton, S. Vikram, and R. A Saurous · 2024
Closest in time.
Learning to compress prompts with gist tokens
J. Mu, X. Li, and N. Goodman · 2024
Closest in time.
Let’s think dot by dot: Hidden computation in transformer language models
J. Pfau, W. Merrill, and S. R. Bowman · 2024
Closest in time.
Distributional reasoning in llms: Parallel reasoning processes in multi-hop reasoning
Y. Shalev, A. Feder, and A. Goldstein · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Team-Gemma, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, et al · 2024
Closest in time.
Chain-of-thought reasoning without prompting
X. Wang and D. Zhou · 2024
Closest in time.
Y. Wu, Z. Sun, S. Li, S. Welleck, and Y. Yang · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan · 2024
Closest in time.
Quiet-star: Language models can teach themselves to think before speaking
E. Zelikman, G. Harik, Y. Shao, V. Jayasiri, N. Haber, and N. D. Goodman · 2024
Closest in time.