Fetching the paper…
Reading the bibliography…
Large pretrained language models (LMs) have shown impressive In-Context Learning (ICL) ability, where the model learns to do an unseen task via a prompt consisting of input-output examples as the demonstration, without any parameter updates.
An exact algorithm for maximum entropy sampling
Ko, C.-W., Lee, J., and Queyranne, M · 1995
Earlier work this paper cites.
Learning to parse database queries using inductive logic programming
Zelle, J. M. and Mooney, R. J · 1996
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Schölkopf, B., Smola, A. J., Bach, F., et al · 2002
Earlier work this paper cites.
Unsupervised construction of large paraphrase corpora: Exploiting massively parallel news sources
Dolan, W. B., Quirk, C., and Brockett, C · 2004
Earlier work this paper cites.
Learning to rank for information retrieval
Liu, T.-Y. et al · 2009
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Robertson, S. and Zaragoza, H · 2009
Earlier work this paper cites.
k-dpps: Fixed-size determinantal point processes
Kulesza, A. and Taskar, B · 2011
Earlier work this paper cites.
Near-optimal map inference for determinantal point processes
Gillenwater, J., Kulesza, A., and Taskar, B · 2012
Earlier work this paper cites.
Determinantal point processes for machine learning
Kulesza, A., Taskar, B., et al · 2012
Earlier work this paper cites.
Semantic parsing on Freebase from question-answer pairs
Berant, J., Chou, A., Frostig, R., and Liang, P · 2013
Earlier work this paper cites.
Metric learning: A survey
Kulis, B. et al · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Diverse sequential subset selection for supervised video summarization
Gong, B., Chao, W.-L., Grauman, K., and Sha, F · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Heilbron, F. C., Escorcia, V., Ghanem, B., and Niebles, J. C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Learning detection with diverse proposals
Azadi, S., Feng, J., and Darrell, T · 2017
Earlier work this paper cites.
Faster greedy map inference for determinantal point processes
Han, I., Kambadur, P., Park, K., and Shin, J · 2017
Earlier work this paper cites.
Movie description
Rohrbach, A., Torabi, A., Rohrbach, M., Tandon, N., Pal, C., Larochelle, H., Courville, A., and Schiele, B · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Deep determinantal point process for large-scale multi-label classification
Xie, P., Salakhutdinov, R., Mou, L., and Xing, E. P · 2017
Earlier work this paper cites.
Fast greedy map inference for determinantal point process to improve recommendation diversity
Chen, L., Zhang, G., and Zhou, E · 2018
Earlier work this paper cites.
Improving text-to-SQL evaluation methodology
Finegan-Dollak, C., Kummerfeld, J. K., Zhang, L., Ramanathan, K., Sadasivam, S., Zhang, R., and Radev, D · 2018
Cited alongside, same era.
Nl2bash: A corpus and semantic parser for natural language interface to the linux operating system
Lin, X. V., Wang, C., Zettlemoyer, L., and Ernst, M. D · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Cited alongside, same era.
Mtop: A comprehensive multilingual task-oriented semantic parsing benchmark
Li, H., Arora, A., Chen, S., Gupta, A., Gupta, S., and Mehdad, Y · 2021
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Later among the works it cites.
Compositional generalization and natural language variation: Can a semantic parsing approach handle both?
Shaw, P., Chang, M.-W., Pasupat, P., and Toutanova, K · 2021
Later among the works it cites.
Compositional generalization for neural semantic parsing via span-level supervised attention
Yin, P., Fang, H., Neubig, G., Pauls, A., Platanios, E. A., Su, Y., Thomson, S., and Andreas, J · 2021
Later among the works it cites.
In-context examples selection for machine translation
Agrawal, S., Zhou, C., Lewis, M., Zettlemoyer, L., and Ghazvininejad, M · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Talmor, A., Herzig, J., Lourie, N., and Berant, J · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
Task-oriented dialogue as dataflow synthesis
Andreas, J., Bufe, J., Burkett, D., Chen Jr, C., Clausman, J., Crawford, J., Crim, K., DeLoach, J., Dorner, L., Eisner, J., et al · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Later among the works it cites.
Cont: Contrastive neural text generation
An, C., Feng, J., Lv, K., Kong, L., Qiu, X., and Huang, X · 2022
Later among the works it cites.
Human-level play in the game of diplomacy by combining language models with strategic reasoning
FAIR, Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., et al · 2022
Later among the works it cites.
Diverse demonstrations improve in-context compositional generalization
Levy, I., Bogin, B., and Berant, J · 2022
Later among the works it cites.
On the advance of making language models better reasoners
Li, Y., Lin, Z., Zhang, S., Fu, Q., Chen, B., Lou, J.-G., and Chen, W · 2022
Later among the works it cites.
What makes good in-context examples for gpt-3?
Liu, J., Shen, D., Zhang, Y., Dolan, W. B., Carin, L., and Chen, W · 2022
Later among the works it cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P · 2022
Later among the works it cites.
Chatgpt: Optimizing language models for dialogue
OpenAI, T · 2022
Later among the works it cites.
Improving compositional generalization with latent structure and data augmentation
Qiu, L., Shaw, P., Pasupat, P., Nowak, P., Linzen, T., Sha, F., and Toutanova, K · 2022
Later among the works it cites.
Learning to retrieve prompts for in-context learning
Rubin, O., Herzig, J., and Berant, J · 2022
Later among the works it cites.
Natural language to code translation with execution
Shi, F., Fried, D., Ghazvininejad, M., Zettlemoyer, L., and Wang, S. I · 2022
Later among the works it cites.
Selective annotation makes language models better few-shot learners
Su, H., Kasai, J., Wu, C. H., Shi, W., Wang, T., Xin, J., Zhang, R., Ostendorf, M., Zettlemoyer, L., Smith, N. A., et al · 2022
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., and Zhou, D · 2022
Later among the works it cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Later among the works it cites.
Self-adaptive in-context learning
Wu, Z., Wang, Y., Ye, J., and Kong, L · 2022
Later among the works it cites.
ProGen: Progressive zero-shot dataset generation via in-context feedback
Ye, J., Gao, J., Wu, Z., Feng, J., Yu, T., and Kong, L · 2022
Later among the works it cites.
Generating data for symbolic language with large language models
Ye, J., Li, C., Kong, L., and Yu, T · 2023
Closest in time.