Fetching the paper…
Reading the bibliography…
Large language models (LLMs) can be seen as atomic units of computation mapping sequences to a distribution over sequences.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Distilling task-specific knowledge from BERT into simple neural networks
Tang, R., Lu, Y., Liu, L., Mou, L., Vechtomova, O., and Lin, J. (2019) · 1903
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T. (2019) · 1910
Earlier work this paper cites.
Language as a latent variable: Discrete generative models for sentence compression
Miao, Y. and Blunsom, P. (2016a) · 2016
Earlier work this paper cites.
Language as a latent variable: Discrete generative models for sentence compression
Miao, Y. and Blunsom, P. (2016b) · 2016
Earlier work this paper cites.
Variational inference: A review for statisticians
Blei, D. M., Kucukelbir, A., and McAuliffe, J. D. (2017) · 2017
Earlier work this paper cites.
Learning to few-shot learn across diverse natural language classification tasks
Bansal, T., Jha, R., and McCallum, A. (2020) · 2020
Earlier work this paper cites.
Xtremedistil: Multi-stage distillation for massive multilingual models
Mukherjee, S. and Awadallah, A. H. (2020) · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S. (2020) · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Mitchell, M. (2021) · 2021
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Gao, T., Fisch, A., and Chen, D. (2021) · 2021
Earlier work this paper cites.
Understanding by understanding not: Modeling negation in language models
Hosseini, A., Reddy, S., Bahdanau, D., Hjelm, R. D., Sordoni, A., and Courville, A. C. (2021) · 2021
Earlier work this paper cites.
What makes good in-context examples for gpt- 3 3 ?
Liu, J., Shen, D., Zhang, Y., Dolan, B., Carin, L., and Chen, W. (2021) · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., et al. (2021) · 2021
Earlier work this paper cites.
Responsible language technologies: Foreseeing and mitigating harms
Blodgett, S. L., Liao, Q. V., Olteanu, A., Mihalcea, R., Muller, M., Scheuerman, M. K., Tan, C., and Yang, Q. (2022) · 2022
Earlier work this paper cites.
Selection-inference: Exploiting large language models for interpretable logical reasoning
Creswell, A., Shanahan, M., and Higgins, I. (2022) · 2022
Cited alongside, same era.
Rlprompt: Optimizing discrete text prompts with reinforcement learning
Deng, M., Wang, J., Hsieh, C., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E. P., and Hu, Z. (2022) · 2022
Cited alongside, same era.
Dohan, D., Xu, W., Lewkowycz, A., Austin, J., Bieber, D., Lopes, R. G., Wu, Y., Michalewski, H., Saurous, R. A., Sohl-Dickstein, J., et al. (2022) · 2022
Cited alongside, same era.
Instruction induction: From few examples to natural language task descriptions
Honovich, O., Shaham, U., Bowman, S. R., and Levy, O. (2022) · 2022
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
Recitation-augmented language models
Sun, Z., Wang, X., Tay, Y., Yang, Y., and Zhou, D. (2022) · 2022
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., , and Wei, J. (2022) · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D. (2022) · 2022
Later among the works it cites.
Taxonomy of risks posed by language models
Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., et al. (2022) · 2022
Later among the works it cites.
Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts
Wu, T., Terry, M., and Cai, C. J. (2022) · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., et al. (2022) · 2022
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M. (2022) · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. (2022) · 2022
Cited alongside, same era.
Internet-augmented language models through few-shot prompting for open-domain question answering
Lazaridou, A., Gribovskaya, E., Stokowiec, W., and Grigorev, N. (2022) · 2022
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P. (2022) · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Cited alongside, same era.
Grips: Gradient-free, edit-based instruction search for prompting large language models
Prasad, A., Hase, P., Zhou, X., and Bansal, M. (2022) · 2022
Cited alongside, same era.
Later among the works it cites.
Improving factuality and reasoning in language models through multiagent debate
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I. (2023) · 2023
Closest in time.
Reasoning with language model is planning with world model
Hao, S., Gu, Y., Ma, H., Hong, J. J., Wang, Z., Wang, D. Z., and Hu, Z. (2023) · 2023
Closest in time.
Decomposed prompting: A modular approach for solving complex tasks
Khot, T., Trivedi, H., Finlayson, M., Fu, Y., Richardson, K., Clark, P., and Sabharwal, A. (2023) · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G. (2023) · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Welleck, S., Majumder, B. P., Gupta, S., Yazdanbakhsh, A., and Clark, P. (2023) · 2023
Closest in time.
Augmented language models: a survey
Mialon, G., Dessì, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Rozière, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., Grave, E., LeCun, Y., and Scialom, T. (2023) · 2023
Closest in time.
Automatic prompt optimization with" gradient descent" and beam search
Pryzant, R., Iter, D., Li, J., Lee, Y. T., Zhu, C., and Zeng, M. (2023) · 2023
Closest in time.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Shinn, N., Labash, B., and Gopinath, A. (2023) · 2023
Closest in time.
Selective annotation makes language models better few-shot learners
Su, H., Kasai, J., Wu, C. H., Shi, W., Wang, T., Xin, J., Zhang, R., Ostendorf, M., Zettlemoyer, L., Smith, N. A., et al. (2023) · 2023
Closest in time.
Automatic chain of thought prompting in large language models
Zhang, Z., Zhang, A., Li, M., and Smola, A. (2023) · 2023
Closest in time.