Fetching the paper…
Reading the bibliography…
Pretrained language models (LMs) perform well on many tasks even when learning from a few examples, but prior work uses many held-out examples to tune various aspects of learning, such as hyperparameters, training objectives, and natural language templates ("prompts").
Some comments on cp
C. L. Mallows · 1973
Earlier work this paper cites.
The relationship between variable selection and data agumentation and a method for prediction
David M. Allen · 1974
Earlier work this paper cites.
Cross-validatory choice and assessment of statistical predictions
M. Stone · 1974
Earlier work this paper cites.
A new look at the statistical model identification
H. Akaike · 1974
Earlier work this paper cites.
The predictive sample reuse method with applications
Seymour Geisser · 1975
Earlier work this paper cites.
Modeling by shortest data description
J. Rissanen · 1978
Earlier work this paper cites.
Universal coding, information, prediction, and estimation
J. Rissanen · 1984
Earlier work this paper cites.
Present position and potential developments: Some personal views: Statistical theory: The prequential approach
A. P. Dawid · 1984
Earlier work this paper cites.
Occam’s razor
Alselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K. Warmuth · 1987
Earlier work this paper cites.
Learning many related tasks at the same time with backpropagation
Rich Caruana · 1995
Earlier work this paper cites.
Design and regularization of neural networks: the optimal use of a validation set
J. Larsen, L.K. Hansen, C. Svarer, and M. Ohlsson · 1996
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
The feret evaluation methodology for face-recognition algorithms
P. J. Phillips, Hyeonjoon Moon, S. A. Rizvi, and P. J. Rauss · 2000
Earlier work this paper cites.
Gradient-Based Optimization of Hyperparameters
Yoshua Bengio · 2000
Earlier work this paper cites.
The Elements of Statistical Learning
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2001
Earlier work this paper cites.
A tutorial introduction to the minimum description length principle
Peter Grünwald · 2004
Earlier work this paper cites.
Choosing multiple parameters for support vector machines
O. Chapelle, V. Vapnik, O. Bousquet, and S. Mukherjee · 2004
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2006
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Roy Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Luisa Bentivogli, Ido Dagan, Hoa Trang Dang, Danilo Giampiccolo, and Bernardo Magnini · 2009
Earlier work this paper cites.
Asymptotic equivalence of bayes cross validation and widely applicable information criterion in singular learning theory
Sumio Watanabe · 2010
Earlier work this paper cites.
Sequential model-based optimization for general algorithm configuration
Frank Hutter, Holger H. Hoos, and Kevin Leyton-Brown · 2011
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
The winograd schema challenge
Hector J. Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier García, Fern, and o Fernández · 2015
Earlier work this paper cites.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, koray kavukcuoglu, and Daan Wierstra · 2016
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul F. Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel · 2017
Earlier work this paper cites.
Optimization as a model for few-shot learning
S. Ravi and H. Larochelle · 2017
Earlier work this paper cites.
Ke Li and Jitendra Malik · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V. Le · 2017
Cited alongside, same era.
Practical bayesian model evaluation using leave-one-out cross-validation and waic
Aki Vehtari, Andrew Gelman, and Jonah Gabry · 2017
Cited alongside, same era.
The description length of deep learning models
Léonard Blier and Yann Ollivier · 2018
Cited alongside, same era.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
Exploiting cloze questions for few-shot text classification and natural language inference
Timo Schick and Hinrich Schütze · 2020
Later among the works it cites.
How can we know what language models know?
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig · 2020
Later among the works it cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen · 2020
Later among the works it cites.
It’s not just size that matters: Small language models are also few-shot learners
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ethical challenges in data-driven dialogue systems
Peter Henderson, Koustuv Sinha, Nicolas Angelard-Gontier, Nan Rosemary Ke, Genevieve Fried, Ryan Lowe, and Joelle Pineau · 2018
Cited alongside, same era.
Advancing the state of the art in open domain dialog systems through the alexa prize
Chandra Khatri, Behnam Hedayatnia, Anu Venkatesh, Jeff Nunn, Yi Pan, Qing Liu, Han Song, Anna Gottardi, Sanjeev Kwatra, Sanju Pancholi, Ming Cheng, Qinglang Chen, Lauren Stubel, Karthik Gopalakrishnan, Kate Bland, Raefer Gabriel, Arindam Mandal, Dilek Hakkani-Tür, Gene Hwang, Nate Michel, Eric King, and Rohit Prasad · 2018
Cited alongside, same era.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R. Bowman · 2018
Cited alongside, same era.
Unsupervised neural machine translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho · 2018
Cited alongside, same era.
Unsupervised machine translation using monolingual corpora only
Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato · 2018
Cited alongside, same era.
Realistic evaluation of deep semi-supervised learning algorithms
Avital Oliver, Augustus Odena, Colin Raffel, Ekin D. Cubuk, and Ian J. Goodfellow · 2018
Cited alongside, same era.
Timo Schick and Hinrich Schütze · 2020
Later among the works it cites.
Few-shot text generation with pattern-exploiting training
Timo Schick and H. Schutze · 2020
Later among the works it cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh · 2020
Later among the works it cites.
E-BERT: Efficient-yet-effective entity embeddings for BERT
Nina Poerner, Ulli Waltinger, and Hinrich Schütze · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Later among the works it cites.
Information-theoretic probing with minimum description length
Elena Voita and Ivan Titov · 2020
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Later among the works it cites.
Meta-dataset: A dataset of datasets for learning to learn from few examples
Eleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin, Utku Evci, Kelvin Xu, Ross Goroshin, Carles Gelada, Kevin Swersky, Pierre-Antoine Manzagol, and Hugo Larochelle · 2020
Later among the works it cites.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le · 2020
Later among the works it cites.
MixText: Linguistically-informed interpolation of hidden space for semi-supervised text classification
Jiaao Chen, Zichao Yang, and Diyi Yang · 2020
Later among the works it cites.
Generative data augmentation for commonsense reasoning
Yiben Yang, Chaitanya Malaviya, Jared Fernandez, Swabha Swayamdipta, Ronan Le Bras, Ji-Ping Wang, Chandra Bhagavatula, Yejin Choi, and Doug Downey · 2020
Later among the works it cites.
Unsupervised question decomposition for question answering
Ethan Perez, Patrick Lewis, Wen-tau Yih, Kyunghyun Cho, and Douwe Kiela · 2020
Later among the works it cites.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger · 2020
Later among the works it cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah A. Smith · 2020
Later among the works it cites.
Calibrate before use: Improving few-shot performance of language models, 2021
Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Closest in time.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity, 2021
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2021
Closest in time.
What makes good in-context examples for gpt-3?
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, L. Carin, and W. Chen · 2021
Closest in time.
Rissanen data analysis: Examining dataset characteristics via description length
Ethan Perez, Douwe Kiela, and Kyunghyun Cho · 2021
Closest in time.
Improving and simplifying pattern exploiting training
Derek Tam, Rakesh R Menon, Mohit Bansal, Shashank Srivastava, and Colin Raffel · 2021
Closest in time.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Closest in time.
Entailment as few-shot learner, 2021
Sinong Wang, Han Fang, Madian Khabsa, Hanzi Mao, and Hao Ma · 2021
Closest in time.
How many data points is a prompt worth?, 2021
Teven Le Scao and Alexander M. Rush · 2021
Closest in time.
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela · 2021
Closest in time.
Gpt understands, too, 2021
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang · 2021
Closest in time.
Factual probing is [MASK]: learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen · 2021
Closest in time.
Crossfit: A few-shot learning challenge for cross-task generalization in NLP
Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren · 2021
Closest in time.