Fetching the paper…
Reading the bibliography…
Autoregressive Large Language Models have transformed the landscape of Natural Language Processing.
2013
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in
2014
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,”
2017
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever
2018
Earlier work this paper cites.
F. Petroni, T. Rocktäschel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu, and A. Miller, “Language models as knowledge bases?” in
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh, “Universal adversarial triggers for attacking and analyzing NLP,” in
2019
Earlier work this paper cites.
J. Davison, J. Feldman, and A. Rush, “Commonsense knowledge mining from pretrained models,” in
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,”
2020
Earlier work this paper cites.
Y. Liu, J. Gu, N. Goyal, X. Li, S. Edunov, M. Ghazvininejad, M. Lewis, and L. Zettlemoyer, “Multilingual denoising pre-training for neural machine translation,”
2020
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,”
2020
Earlier work this paper cites.
Z. Jiang, F. F. Xu, J. Araki, and G. Neubig, “How can we know what language models know?”
2020
Earlier work this paper cites.
J. Sarzynska-Wawer, A. Wawer, A. Pawlak, J. Szymanowska, I. Stefaniak, M. Jarkiewicz, and L. Okruszek, “Detecting formal thought disorder by deep contextualized word representations,”
2021
Cited alongside, same era.
X. Liu, F. Zhang, Z. Hou, L. Mian, Z. Wang, J. Zhang, and J. Tang, “Self-supervised learning: Generative or contrastive,”
2021
Cited alongside, same era.
2021
Cited alongside, same era.
W. Yuan, G. Neubig, and P. Liu, “Bartscore: Evaluating generated text as text generation,”
2021
Cited alongside, same era.
A. Haviv, J. Berant, and A. Globerson, “BERTese: Learning to speak to BERT,” in
2021
Cited alongside, same era.
2022
Later among the works it cites.
Y. Lu, M. Bartolo, A. Moore, S. Riedel, and P. Stenetorp, “Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity,” in
2022
Later among the works it cites.
M. Deng, J. Wang, C.-P. Hsieh, Y. Wang, H. Guo, T. Shu, M. Song, E. Xing, and Z. Hu, “RLPrompt: Optimizing discrete text prompts with reinforcement learning,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Chen, N. Zhang, X. Xie, S. Deng, Y. Yao, C. Tan, F. Huang, L. Si, and H. Chen, “Knowprompt: Knowledge-aware prompt-tuning with synergistic optimization for relation extraction,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” in
2021
Cited alongside, same era.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,” in
2021
Cited alongside, same era.
2022
Cited alongside, same era.
H. Wang, J. Li, H. Wu, E. Hovy, and Y. Sun, “Pre-trained language models and their applications,”
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou
2022
Cited alongside, same era.
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,”
2022
Cited alongside, same era.
2022
Later among the works it cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray
2022
Later among the works it cites.
S. Armstrong and R. Gorman, “Using gpt-eliezer against chatgpt jailbreaking,”
2022
Later among the works it cites.
2023
Closest in time.
D. Hulbert, “Tree of knowledge: Tok aka tree of knowledge dataset for large language models llm,” https://github.com/dave1010/tree-of-thought-prompting, 2023
2023
Closest in time.
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig, “Pal: Program-aided language models,” in
2023
Closest in time.
J. Long, “Large language model guided tree-of-thought,”
2023
Closest in time.
2023
Closest in time.
X. Liu, Y. Zheng, Z. Du, M. Ding, Y. Qian, Z. Yang, and J. Tang, “Gpt understands, too,”
2023
Closest in time.
2023
Closest in time.
Z. Liu, X. Yu, Y. Fang, and X. Zhang, “Graphprompt: Unifying pre-training and downstream tasks for graph neural networks,” in
2023
Closest in time.
C. Nardo, “The waluigi effect (mega-post),”
2023
Closest in time.