Fetching the paper…
Reading the bibliography…
Evaluating the in-context learning classification performance of language models poses challenges due to small dataset sizes, extensive prompt-selection using the validation set, and intentionally difficult tasks that lead to near-random performance.
Machine Learning
Tom Mitchell · 1997
Earlier work this paper cites.
Permutation tests for classification
Polina Golland, Feng Liang, Sayan Mukherjee, and Dmitry Panchenko · 2000
Earlier work this paper cites.
Statistical Inference
George Casella and Roger L. Berger · 2002
Earlier work this paper cites.
Permutation tests for classification: towards statistical significance in image-based studies
Polina Golland and Bruce Fischl · 2003
Earlier work this paper cites.
Permutation tests for studying classifier performance
Markus Ojala and Gemma C Garriga · 2010
Earlier work this paper cites.
On computing the distribution function for the Poisson binomial distribution
Yili Hong · 2013
Earlier work this paper cites.
Generalization in adaptive data analysis and holdout reuse
Cynthia Dwork, Vitaly Feldman, Moritz Hardt, Toni Pitassi, Omer Reingold, and Aaron Roth · 2015
Earlier work this paper cites.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart · 2018
Earlier work this paper cites.
Show your work: Improved reporting of experimental results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A Smith · 2019
Earlier work this paper cites.
Deep dominance-how to properly compare deep neural models
Rotem Dror, Segev Shlomov, and Roi Reichart · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
With little power comes great responsibility
Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky · 2020
Earlier work this paper cites.
Array programming with NumPy
Charles R. Harris, K. Jarrod Millman, Stéfan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fernández del Río, Mark Wiebe, Pearu Peterson, Pierre Gérard-Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, Christoph Gohlke, and Travis E. Oliphant · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2020
Earlier work this paper cites.
RAFT: A real-world few-shot text classification benchmark
Neel Alex, Eli Lifland, Lewis Tunstall, Abhishek Thakur, Pegah Maham, C Jess Riedel, Emmie Hine, Carolyn Ashurst, Paul Sedille, Alexis Carlier, et al · 2021
Cited alongside, same era.
FLEX: Unifying evaluation for few-shot NLP
Jonathan Bragg, Arman Cohan, Kyle Lo, and Iz Beltagy · 2021
Cited alongside, same era.
Expected validation performance and estimation of a random variable’s maximum
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A Smith · 2021
Cited alongside, same era.
Surface form competition: Why the highest probability answer isn’t always right
Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer · 2021
Cited alongside, same era.
Reduced, reused and recycled: The life of a dataset in machine learning research
Bernard Koch, Emily Denton, Alex Hanna, and Jacob G Foster · 2021
Cited alongside, same era.
True few-shot learning with language models
QLoRA: Efficient finetuning of quantized LLMs
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Later among the works it cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2023
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2023
Later among the works it cites.
Can generalist foundation models outcompete special-purpose tuning? Case study in medicine
Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, et al · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ethan Perez, Douwe Kiela, and Kyunghyun Cho · 2021
Cited alongside, same era.
Better than average: Paired evaluation of nlp systems
Maxime Peyrard, Wei Zhao, Steffen Eger, and Robert West · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2022
Cited alongside, same era.
Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, et al · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2022
Cited alongside, same era.
Exact paired-permutation testing for structured test statistics
Ran Zmigrod, Tim Vieira, and Ryan Cotterell · 2022
Cited alongside, same era.
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, and Jason Wei · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
OLMo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al · 2024
Closest in time.
State of what art? a call for multi-prompt LLM evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky · 2024
Closest in time.
tinyBenchmarks: evaluating LLMs with fewer examples
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin · 2024
Closest in time.
Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr · 2024
Closest in time.
Mind your format: Towards consistent evaluation of in-context learning improvements
Anton Voronov, Lena Wolf, and Max Ryabinin · 2024
Closest in time.