Fetching the paper…
Reading the bibliography…
Recently, Language Models (LMs) instruction-tuned on multiple tasks, also known as multitask-prompted fine-tuning (MT), have shown the capability to generalize to unseen tasks.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 1908
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M. and Cohen, N. J · 1989
Earlier work this paper cites.
Learning question classifiers
Li, X. and Roth, D · 2002
Earlier work this paper cites.
English gigaword
Graff, D., Kong, J., Chen, K., and Maeda, K · 2003
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2005
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Pang, B. and Lee, L · 2005
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Roemmele, M., Bejan, C. A., and Gordon, A. S · 2011
Earlier work this paper cites.
The winograd schema challenge
Levesque, H., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
Hidden factors and hidden topics: understanding rating dimensions with review text
McAuley, J. J. and Leskovec, J · 2013
Earlier work this paper cites.
Dbpedia - a large-scale, multilingual knowledge base extracted from wikipedia
Lehmann, J., Isele, R., Jakob, M., Jentzsch, A., Kontokostas, D., Mendes, P. N., Hellmann, S., Morsey, M., van Kleef, P., Auer, S., and Bizer, C · 2015
Earlier work this paper cites.
WikiQA: A challenge dataset for open-domain question answering
Yang, Y., Yih, W.-t., and Meek, C · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J. J., and LeCun, Y · 2015
Earlier work this paper cites.
Neural text generation from structured data with application to the biography domain
Lebret, R., Grangier, D., and Auli, M · 2016
Earlier work this paper cites.
A corpus and cloze evaluation for deeper understanding of commonsense stories
Mostafazadeh, N., Chambers, N., He, X., Parikh, D., Batra, D., Vanderwende, L., Kohli, P., and Allen, J · 2016
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
See, A., Liu, P. J., and Manning, C. D · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J., Liu, N. F., and Gardner, M · 2017
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Camburu, O.-M., Rocktäschel, T., Lukasiewicz, T., and Blunsom, P · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
DuoRC: Towards complex language understanding with paraphrased reading comprehension
Saha, A., Aralikatte, R., Khapra, M. M., and Sankaranarayanan, K · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Earlier work this paper cites.
Constructing datasets for multi-hop reading comprehension across documents
Welbl, J., Stenetorp, P., and Riedel, S · 2018
Earlier work this paper cites.
The commitmentbank: Investigating projection in naturally occurring discourse
De Marneffe, M.-C., Simons, M., and Tonhauser, J · 2019
Earlier work this paper cites.
Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model
Fabbri, A., Li, I., She, T., Li, S., and Radev, D · 2019
Earlier work this paper cites.
ELI5: Long form question answering
Fan, A., Jernite, Y., Perez, E., Grangier, D., Weston, J., and Auli, M · 2019
Earlier work this paper cites.
SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization
Gliwa, B., Mochol, I., Biesek, M., and Wawer, A · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
Huang, L., Le Bras, R., Bhagavatula, C., and Choi, Y · 2019
Earlier work this paper cites.
Reasoning over paragraph effects in situations
Lin, K., Tafjord, O., Clark, P., and Gardner, M · 2019
Earlier work this paper cites.
WiC: the word-in-context dataset for evaluating context-sensitive meaning representations
Pilehvar, M. T. and Camacho-Collados, J · 2019
Cited alongside, same era.
Explain yourself! leveraging language models for commonsense reasoning
Rajani, N. F., McCann, B., Xiong, C., and Socher, R · 2019
Cited alongside, same era.
Towards empathetic open-domain conversation models: A new benchmark and dataset
Rashkin, H., Smith, E. M., Li, M., and Boureau, Y.-L · 2019
Cited alongside, same era.
Social IQa: Commonsense reasoning about social interactions
Sap, M., Rashkin, H., Chen, D., Le Bras, R., and Choi, Y · 2019
Cited alongside, same era.
DREAM: A challenge data set and models for dialogue-based reading comprehension
Sun, K., Yu, D., Chen, J., Yu, D., Choi, Y., and Cardie, C · 2019
Cited alongside, same era.
Quarel: A dataset and models for answering questions about qualitative relationships
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al · 2021
Later among the works it cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2021
Later among the works it cites.
Transformer-based lexically constrained headline generation
Yamada, K., Hitomi, Y., Tamori, H., Sasano, R., Okazaki, N., Inui, K., and Takeda, K · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tafjord, O., Clark, P., Gardner, M., Yih, W.-t., and Sabharwal, A · 2019
Cited alongside, same era.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Talmor, A., Herzig, J., Lourie, N., and Berant, J · 2019
Cited alongside, same era.
WIQA: A dataset for “what if…” reasoning over procedural text
Tandon, N., Dalvi, B., Sakaguchi, K., Clark, P., and Bosselut, A · 2019
Cited alongside, same era.
HellaSwag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
PAWS: Paraphrase adversaries from word scrambling
Zhang, Y., Baldridge, J., and He, L · 2019
Cited alongside, same era.
Beat the AI: Investigating adversarial human annotation for reading comprehension
Bartolo, M., Roberts, A., Welbl, J., Riedel, S., and Stenetorp, P · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S. K., Hayase, J., and Srinivasa, S · 2022
Later among the works it cites.
Attempt: Parameter-efficient multi-task tuning via attentional mixtures of soft prompts
Asai, A., Salehi, M., Peters, M. E., and Hajishirzi, H · 2022
Later among the works it cites.
PromptSource: An integrated development environment and repository for natural language prompts
Bach, S., Sanh, V., Yong, Z. X., Webson, A., Raffel, C., Nayak, N. V., Sharma, A., Kim, T., Bari, M. S., Fevry, T., Alyafeai, Z., Dey, M., Santilli, A., Sun, Z., Ben-david, S., Xu, C., Chhablani, G., Wang, H., Fries, J., Al-shaibani, M., Sharma, S., Thakker, U., Almubarak, K., Tang, X., Radev, D., Jiang, M. T.-j., and Rush, A · 2022
Later among the works it cites.
Spt: Semi-parametric prompt tuning for multitask prompted learning
Bari, M. S., Zhang, A., Zheng, S., Shi, X., Zhu, Y., Joty, S., and Li, M · 2022
Later among the works it cites.
Petals: Collaborative inference and fine-tuning of large models
Borzunov, A., Baranchuk, D., Dettmers, T., Ryabinin, M., Belkada, Y., Chumachenko, A., Samygin, P., and Raffel, C · 2022
Later among the works it cites.
Fine-tuned language models can be continual learners
Chakrabarty, T., Scialom, T., and Muresan, S · 2022
Later among the works it cites.
Few-shot adaptation works with unpredictable data
Chan, J. S., Pieler, M., Jao, J., Scheurer, J., and Perez, E · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Later among the works it cites.
Diffcse: Difference-based contrastive learning for sentence embeddings
Chuang, Y.-S., Dangovski, R., Luo, H., Zhang, Y., Chang, S., Soljačić, M., Li, S.-W., Yih, S., Kim, Y., and Glass, J · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Later among the works it cites.
Cold fusion: Collaborative descent for distributed multitask finetuning
Don-Yehiya, S., Venezian, E., Raffel, C., Slonim, N., Katz, Y., and Choshen, L · 2022
Later among the works it cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2022
Later among the works it cites.
Decomposed prompting: A modular approach for solving complex tasks
Khot, T., Trivedi, H., Finlayson, M., Fu, Y., Richardson, K., Clark, P., and Sabharwal, A · 2022
Later among the works it cites.
Standing on the shoulders of giant frozen language models
Levine, Y., Dalmedigos, I., Ram, O., Zeldes, Y., Jannai, D., Muhlgay, D., Osin, Y., Lieber, O., Lenz, B., Shalev-Shwartz, S., et al · 2022
Later among the works it cites.
Branch-train-merge: Embarrassingly parallel training of expert language models
Li, M., Gururangan, S., Dettmers, T., Lewis, M., Althoff, T., Smith, N. A., and Zettlemoyer, L · 2022
Later among the works it cites.
Unsupervised cross-task generalization via retrieval augmentation
Lin, B. Y., Tan, K., Miller, C., Tian, B., and Ren, X · 2022
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Muennighoff, N., Wang, T., Sutawika, L., Roberts, A., Biderman, S., Scao, T. L., Bari, M. S., Shen, S., Yong, Z.-X., Schoelkopf, H., et al · 2022
Later among the works it cites.
Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models
Ni, J., Abrego, G. H., Constant, N., Ma, J., Hall, K., Cer, D., and Yang, Y · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Lifting the curse of multilinguality by pre-training modular transformers
Pfeiffer, J., Goyal, N., Lin, X. V., Li, X., Cross, J., Riedel, S., and Artetxe, M · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Later among the works it cites.
One embedder, any task: Instruction-finetuned text embeddings
Su, H., Kasai, J., Wang, Y., Hu, Y., Ostendorf, M., Yih, W.-t., Smith, N. A., Zettlemoyer, L., Yu, T., et al · 2022
Later among the works it cites.
SPoT: Better frozen model adaptation through soft prompt transfer
Vu, T., Lester, B., Constant, N., Al-Rfou’, R., and Cer, D · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D · 2022
Later among the works it cites.
A survey on negative transfer
Zhang, W., Deng, L., Zhang, L., and Wu, D · 2022
Later among the works it cites.