Fetching the paper…
Reading the bibliography…
Prompt-based learning has shown considerable promise in reformulating various downstream tasks as cloze problems by combining original input with a predetermined template.
Wordnet: a lexical database for english
Miller, G. A · 1995
Earlier work this paper cites.
Vicinal risk minimization
Chapelle, O., Weston, J., Bottou, L., and Vapnik, V · 2000
Earlier work this paper cites.
Transformation invariance in pattern recognition: Tangent distance and propagation
Simard, P. Y., LeCun, Y., Denker, J. S., and Victorri, B · 2000
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Bar-Haim, R., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., and Magnini, B · 2006
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Roemmele, M., Bejan, C. A., and Gordon, A. S · 2011
Earlier work this paper cites.
The winograd schema challenge
Levesque, H. J., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S. E., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2014
Earlier work this paper cites.
That’s so annoying!!!: A lexical and frame-semantic embedding based data augmentation approach to automatic categorization of annoying behaviors using #petpeeve tweets
Wang, W. Y. and Yang, D · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J. J., and LeCun, Y · 2015
Earlier work this paper cites.
Pubmed 200k rct: a dataset for sequential sentence classification in medical abstracts
Dernoncourt, F. and Lee, J. Y · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Khashabi, D., Chaturvedi, S., Roth, M., Upadhyay, S., and Roth, D · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H., Cissé, M., Dauphin, Y. N., and Lopez-Paz, D · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Structural scaffolds for citation intent classification in scientific publications
Cohan, A., Ammar, W., van Zuylen, M., and Cady, F · 2019
Earlier work this paper cites.
Commonsense knowledge mining from pretrained models
Davison, J., Feldman, J., and Rush, A · 2019
Earlier work this paper cites.
The commitmentbank: Investigating projection in naturally occurring discourse
de Marneffe, M.-C., Simons, M., and Tonhauser, J · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Augmenting data with mixup for sentence classification: An empirical study
Guo, H., Mao, Y., and Zhang, R · 2019
Earlier work this paper cites.
Learning data manipulation for augmentation and weighting
Hu, Z., Tan, B., Salakhutdinov, R., Mitchell, T. M., and Xing, E. P · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Riedel, S., Lewis, P., Bakhtin, A., Wu, Y., and Miller, A · 2019
Earlier work this paper cites.
WiC: the word-in-context dataset for evaluating context-sensitive meaning representations
Pilehvar, M. T. and Camacho-Collados, J · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N. M., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing NLP
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
EDA: Easy data augmentation techniques for boosting performance on text classification tasks
Wei, J. and Zou, K · 2019
Cited alongside, same era.
Learning how to ask: Querying LMs with mixtures of soft prompts
Qin, G. and Eisner, J · 2021
Later among the works it cites.
Text AutoAugment: Learning compositional augmentation policy for text classification
Ren, S., Zhang, J., Li, L., Sun, X., and Zhou, J · 2021
Later among the works it cites.
Open aspect target sentiment classification with natural language prompts
Seoh, R., Birle, I., Tak, M., Chang, H.-S., Pinette, B., and Hough, A · 2021
Later among the works it cites.
Multimodal few-shot learning with frozen language models
Tsimpoukelli, M., Menick, J., Cabi, S., Eslami, S. M. A., Vinyals, O., and Hill, F · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation, 2021
Yuan, W., Neubig, G., and Liu, P · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Do not have enough data? deep learning to the rescue!
Anaby-Tavor, A., Carmeli, B., Goldbraich, E., Kantor, A., Kour, G., Shlomov, S., Tepper, N., and Zwerdling, N · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
AdvAug: Robust adversarial augmentation for neural machine translation
Cheng, Y., Jiang, L., Macherey, W., and Eisenstein, J · 2020
Cited alongside, same era.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Dodge, J., Ilharco, G., Schwartz, R., Farhadi, A., Hajishirzi, H., and Smith, N. A · 2020
Cited alongside, same era.
Augmix: A simple data processing method to improve robustness and uncertainty
Hendrycks, D., Mu, N., Cubuk, E. D., Zoph, B., Gilmer, J., and Lakshminarayanan, B · 2020
Cited alongside, same era.
How can we know what language models know?
Jiang, Z., Xu, F. F., Araki, J., and Neubig, G · 2020
Cited alongside, same era.
TinyBERT: Distilling BERT for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 2020
Cited alongside, same era.
Adapting language models for zero-shot learning by meta-tuning on dataset and prompt collections
Zhong, R., Lee, K., Zhang, Z., and Klein, D · 2021
Later among the works it cites.
Flipda: Effective and robust data augmentation for few-shot learning, 2021
Zhou, J., Zheng, Y., Tang, J., Li, J., and Yang, Z · 2021
Later among the works it cites.
PADA: Example-based Prompt Learning for on-the-fly Adaptation to Unseen Domains
Ben-David, E., Oved, N., and Reichart, R · 2022
Later among the works it cites.
Aug-fedprompt: Practical few-shot federated nlp with data-augmented prompts
Cai, D., Wu, Y., Yuan, H., Wang, S., Lin, F. X., and Xu, M · 2022
Later among the works it cites.
Can prompt probe pretrained language models? understanding the invisible risks from a causal view
Cao, B., Lin, H., Han, X., Liu, F., and Sun, L · 2022
Later among the works it cites.
Promptda: Label-guided data augmentation for prompt-based few shot learners
Chen, C. and Shu, K · 2022
Later among the works it cites.
Novelty controlled paraphrase generation with retrieval augmented conditional prompt tuning
Chowdhury, J. R., Zhuang, Y., and Wang, S · 2022
Later among the works it cites.
Data augmentation approaches in natural language processing: A survey
Li, B., Hou, Y., and Che, W · 2022
Later among the works it cites.
Improving few-shot performance of language models via nearest neighbor calibration
Nie, F., Chen, M., Zhang, Z., and Cheng, X · 2022
Later among the works it cites.
Enhancing cross-lingual natural language inference by prompt-learning from cross-lingual templates
Qi, K., Wan, H., Du, J., and Chen, H · 2022
Later among the works it cites.
Don’t prompt, search! mining-based zero-shot learning with language models
van de Kar, M., Xia, M., Chen, D., and Artetxe, M · 2022
Later among the works it cites.
Promda: Prompt-based data augmentation for low-resource nlu tasks
Wang, Y., Xu, C., Sun, Q., Hu, H., Tao, C., Geng, X., and Jiang, D · 2022
Later among the works it cites.
Palm 2 technical report
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A. T., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., Chu, E., Clark, J., Shafey, L. E., Huang, Y., Meier-Hellstern, K. S., Mishra, G., Moreira, E., Omernick, M., Robinson, K., Ruder, S., Tay, Y., Xiao, K., Xu, Y., Zhang, Y., Abrego, G. H., Ahn, J., Austin, J., Barham, P., Botha, J. A., Bradbury, J., Brahma, S., Brooks, K. M., Catasta, M., Cheng, Y., Cherry, C., Choquette-Choo, C. A., Chowdhery, A., Crépy, C., Dave, S., Dehghani, M., Dev, S., Devlin, J., D’iaz, M. C., Du, N., Dyer, E., Feinberg, V., Feng, F., Fienber, V., Freitag, M., García, X., Gehrmann, S., González, L., Gur-Ari, G., Hand, S., Hashemi, H., Hou, L., Howland, J., Hu, A. R., Hui, J., Hurwitz, J., Isard, M., Ittycheriah, A., Jagielski, M., Jia, W. H., Kenealy, K., Krikun, M., Kudugunta, S., Lan, C., Lee, K., Lee, B., Li, E., Li, M.-L., Li, W., Li, Y., Li, J. Y., Lim, H., Lin, H., Liu, Z.-Z., Liu, F., Maggioni, M., Mahendru, A., Maynez, J., Misra, V., Moussalem, M., Nado, Z., Nham, J., Ni, E., Nystrom, A., Parrish, A., Pellat, M., Polacek, M., Polozov, A., Pope, R., Qiao, S., Reif, E., Richter, B., Riley, P., Ros, A., Roy, A., Saeta, B., Samuel, R., Shelby, R. M., Slone, A., Smilkov, D., So, D. R., Sohn, D., Tokumine, S., Valter, D., Vasudevan, V., Vodrahalli, K., Wang, X., Wang, P., Wang, Z., Wang, T., Wieting, J., Wu, Y., Xu, K., Xu, Y., Xue, L. W., Yin, P., Yu, J., Zhang, Q., Zheng, S., Zheng, C., Zhou, W., Zhou, D., Petrov, S., and Wu, Y · 2023
Closest in time.
Think outside the code: Brainstorming boosts large language models in code generation
Li, X., Xue, J.-T., Xie, Z., and Li, M · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K. R., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D. M., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A. S., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I. M., Korenev, A. V., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Closest in time.
Cognitive distortion based explainable depression detection and analysis technologies for the adolescent internet users on social media
Wang, B., Zhao, Y., Lu, X., and Qin, B · 2023
Closest in time.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T · 2023
Closest in time.
Baichuan 2: Open large-scale language models
Yang, A. M., Xiao, B., Wang, B., Zhang, B., Bian, C., Yin, C., Lv, C., Pan, D., Wang, D., Yan, D., Yang, F., Deng, F., Wang, F., Liu, F., Ai, G., Dong, G., Zhao, H., Xu, H., Sun, H., Zhang, H., Liu, H., Ji, J., Xie, J., Dai, J., Fang, K., Su, L., Song, L., Liu, L., Ru, L., Ma, L., Wang, M., Liu, M., Lin, M., Nie, N., Guo, P., Sun, R., Zhang, T., Li, T., Li, T., Cheng, W., Chen, W., Zeng, X., Wang, X., Chen, X., Men, X., Yu, X., Pan, X., Shen, Y.-B., Wang, Y., Li, Y., Jiang, Y., Gao, Y., Zhang, Y., Zhou, Z., and Wu, Z · 2023
Closest in time.
Distilling script knowledge from large language models for constrained language planning
Yuan, S., Chen, J., Fu, Z., Ge, X., Shah, S., Jankowski, C. R., Yang, D., and Xiao, Y · 2023
Closest in time.