Fetching the paper…
Reading the bibliography…
Large language models (LLMs) exhibit in-context learning abilities which enable the same model to perform several tasks without any task-specific training.
“Language models are few-shot learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 1901
Earlier work this paper cites.
“Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning”
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal and Colin Raffel · 1965
Earlier work this paper cites.
“Thumbs Up? Sentiment Classification Using Machine Learning Techniques”
Bo Pang, Lillian Lee and Shivakumar Vaithyanathan · 2002
Earlier work this paper cites.
Cambridge university press, 2003
“Probability theory: The logic of science” · 2003
Earlier work this paper cites.
“Learning multiple layers of features from tiny images”, 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
“MNIST handwritten digit database”
Yann LeCun, Corinna Cortes and CJ Burges · 2010
Earlier work this paper cites.
“Contributions to the Study of SMS Spam Filtering: New Collection and Results”
Tiago. Almeida, Jose Hidalgo and Akebo Yamakami · 2011
Earlier work this paper cites.
“Learning Word Vectors for Sentiment Analysis”
Andrew. Maas, Raymond. Daly, Peter. Pham, Dan Huang, Andrew. Ng and Christopher Potts · 2011
Earlier work this paper cites.
“Recursive deep models for semantic compositionality over a sentiment treebank”
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher Manning, Andrew Ng and Christopher Potts · 2013
Earlier work this paper cites.
“Understanding machine learning: From theory to algorithms”
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
“Character-level convolutional networks for text classification”
Xiang Zhang, Junbo Zhao and Yann LeCun · 2015
Earlier work this paper cites.
“Improving language understanding by generative pre-training”
Alec Radford, Karthik Narasimhan, Tim Salimans and Ilya Sutskever · 2018
Earlier work this paper cites.
“Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition”
P. Warden · 2018
Earlier work this paper cites.
“Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification”
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain and Lucy Vasserman · 2019
Earlier work this paper cites.
“Parameter-efficient transfer learning for NLP”
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De, Andrea Gesmundo, Mona Attariyan and Sylvain Gelly · 2019
Earlier work this paper cites.
“Visual Transformers: Token-based Image Representation and Processing for Computer Vision”, 2020
Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Zhicheng Yan, Masayoshi Tomizuka, Joseph Gonzalez, Kurt Keutzer and Peter Vajda · 2020
Earlier work this paper cites.
“RAFT: A Real-World Few-Shot Text Classification Benchmark”
Neel Alex, Eli Lifland, Lewis Tunstall, Abhishek Thakur, Pegah Maham, C Riedel, Emmie Hine, Carolyn Ashurst, Paul Sedille and Alexis Carlier · 2021
Earlier work this paper cites.
“GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow”
Sid Black, Leo Gao, Phil Wang, Connor Leahy and Stella Biderman · 2021
Cited alongside, same era.
“On the opportunities and risks of foundation models”
Rishi Bommasani, Drew Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael Bernstein, Jeannette Bohg, Antoine Bosselut and Emma Brunskill · 2021
Cited alongside, same era.
“The Power of Scale for Parameter-Efficient Prompt Tuning”
Brian Lester, Rami Al-Rfou and Noah Constant · 2021
Cited alongside, same era.
“Prefix-Tuning: Optimizing Continuous Prompts for Generation”
Xiang Li and Percy Liang · 2021
Cited alongside, same era.
“GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model”, https://github.com/kingoflolz/mesh-transformer-jax , 2021
Ben Wang and Aran Komatsuzaki · 2021
Cited alongside, same era.
“Transformers learn in-context by gradient descent”
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov and Max Vladymyrov · 2022
Later among the works it cites.
“Robust Speech Recognition via Large-Scale Weak Supervision”
Alec Radford, Jong Kim, Tao Xu, Greg Brockman, Christine McLeavey and Ilya Sutskever · 2022
Later among the works it cites.
“Bloom: A 176b-parameter open-access multilingual language model”
Teven Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Luccioni, François Yvon and Matthias Gallé · 2022
Later among the works it cites.
“Rationale-augmented ensembles in language models”
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi and Denny Zhou · 2022
Later among the works it cites.
“Self-consistency improves chain of thought reasoning in language models”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sang Xie, Aditi Raghunathan, Percy Liang and Tengyu Ma · 2021
Cited alongside, same era.
“WRENCH: A Comprehensive Benchmark for Weak Supervision”
Jieyu Zhang, Yue Yu, Yinghao Li, Yujing Wang, Yaming Yang, Mao Yang and Alexander Ratner · 2021
Cited alongside, same era.
“Large language models are zero-shot clinical information extractors”
Monica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim and David Sontag · 2022
Cited alongside, same era.
“What can transformers learn in-context? a case study of simple function classes”
Shivam Garg, Dimitris Tsipras, Percy Liang and Gregory Valiant · 2022
Cited alongside, same era.
“LoRA: Low-Rank Adaptation of Large Language Models”
Edward Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang and Weizhu Chen · 2022
Cited alongside, same era.
“Large Language Models are Zero-Shot Reasoners”
Takeshi Kojima, Shixiang Gu, Machel Reid, Yutaka Matsuo and Yusuke Iwasawa · 2022
Cited alongside, same era.
“Fine-tuning can distort pretrained features and underperform out-of-distribution”
Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma and Percy Liang · 2022
Cited alongside, same era.
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi and Denny Zhou · 2022
Later among the works it cites.
“Emergent Abilities of Large Language Models” Survey Certification
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean and William Fedus · 2022
Later among the works it cites.
“Chain of thought prompting elicits reasoning in large language models”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le and Denny Zhou · 2022
Later among the works it cites.
“STaR: Bootstrapping Reasoning With Reasoning”
Eric Zelikman, Yuhuai Wu, Jesse Mu and Noah Goodman · 2022
Later among the works it cites.
“Ask Me Anything: A simple strategy for prompting language models”
Simran Arora, Avanika Narayan, Mayee Chen, Laurel Orr, Neel Guha, Kush Bhatia, Ines Chami, Frederic Sala and Christopher Ré · 2023
Closest in time.
“Pythia: A suite for analyzing large language models across training and scaling”
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Khan, Shivanshu Purohit, USVSN Prashanth and Edward Raff · 2023
Closest in time.
“Active Prompting with Chain-of-Thought for Large Language Models”
Shizhe Diao, Pengcheng Wang, Yong Lin and Tong Zhang · 2023
Closest in time.
“Prompting vs. Finetuning vs. Alternatives”, 2023
Chip Huyen · 2023
Closest in time.
“Chatgpt: Jack of all trades, master of none”
Jan Kocoń, Igor Cichecki, Oliwier Kaszyca, Mateusz Kochanek, Dominika Szydło, Joanna Baran, Julita Bielaniewicz, Marcin Gruza, Arkadiusz Janz and Kamil Kanclerz · 2023
Closest in time.
“Hyena hierarchy: Towards larger convolutional language models”
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon and Christopher Ré · 2023
Closest in time.
“Larger language models do in-context learning differently”
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang and Denny Zhou · 2023
Closest in time.
“LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention”
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao and Yu Qiao · 2023
Closest in time.