Fetching the paper…
Reading the bibliography…
Learning with limited labelled data, such as prompting, in-context learning, fine-tuning, meta-learning or few-shot learning, aims to effectively train a model using only a small amount of labelled samples.
Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems , Vol. 33. Curran Associates, Inc., 1877–1901
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2002
Earlier work this paper cites.
Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement
David Moher, Alessandro Liberati, Jennifer Tetzlaff, Douglas G Altman, and PRISMA Group*. 2009 · 2009
Earlier work this paper cites.
Is Support Set Diversity Necessary for Meta-Learning?
Amrith Setlur, Oscar Li, and Virginia Smith. 2021a · 2011
Earlier work this paper cites.
Matching Networks for One Shot Learning. In Advances in Neural Information Processing Systems , Vol. 29. Curran Associates, Inc
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, koray kavukcuoglu, and Daan Wierstra. 2016 · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning . PMLR, 1126–1135
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Earlier work this paper cites.
Reporting Score Distributions Makes a Difference: Performance Study of LSTM-networks for Sequence Tagging. In Proc. of the 2017 Conference on Empirical Methods in Natural Language Processing . ACL, Copenhagen, Denmark, 338–348
Nils Reimers and Iryna Gurevych. 2017 · 2017
Earlier work this paper cites.
Prototypical Networks for Few-shot Learning. In Advances in Neural Information Processing Systems , Vol. 30. Curran Associates, Inc
Jake Snell, Kevin Swersky, and Richard Zemel. 2017 · 2017
Earlier work this paper cites.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman. 2018 · 2018
Earlier work this paper cites.
mixup: Beyond Empirical Risk Minimization. In International Conference on Learning Representations
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. 2018 · 2018
Earlier work this paper cites.
Reproducibility and Stability Analysis in Metric-Based Few-Shot Learning. In RML@ ICLR , Vol. 3
Thomas Boquet, Laure Delisle, Denis Kochetkov, Nathan Schucher, Boris N Oreshkin, and Julien Cornebise. 2019 · 2019
Earlier work this paper cites.
Unreproducible Research is Reproducible. In Proc. of the 36th International Conf. on Machine Learning . PMLR, 725–734
Xavier Bouthillier, César Laurent, and Pascal Vincent. 2019 · 2019
Earlier work this paper cites.
Momentum-Based Variance Reduction in Non-Convex SGD. In Advances in Neural Information Processing Systems , Vol. 32. Curran Associates, Inc
Ashok Cutkosky and Francesco Orabona. 2019 · 2019
Earlier work this paper cites.
MetaInit: Initializing learning by learning to initialize. In Advances in Neural Information Processing Systems , Vol. 32. Curran Associates, Inc
Yann N Dauphin and Samuel Schoenholz. 2019 · 2019
Earlier work this paper cites.
Deep Dominance - How to Properly Compare Deep Neural Models. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . ACL, Florence, Italy, 2773–2785
Rotem Dror, Segev Shlomov, and Roi Reichart. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP. In International conference on machine learning . PMLR, 2790–2799
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT . 4171–4186
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In International Conference on Learning Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-data Tasks
Jason Phang, Thibault Févry, and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML. In International Conference on Learning Representations
Aniruddh Raghu, Maithra Raghu, Samy Bengio, and Oriol Vinyals. 2019 · 2019
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Earlier work this paper cites.
Multimodal Model-Agnostic Meta-Learning via Task-Aware Modulation. In Advances in Neural Information Processing Systems , Vol. 32. Curran Associates, Inc
Risto Vuorio, Shao-Hua Sun, Hexiang Hu, and Joseph J Lim. 2019 · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Earlier work this paper cites.
Learning to Few-Shot Learn Across Diverse Natural Language Classification Tasks. In Proceedings of the 28th International Conference on Computational Linguistics . 5108–5123
Trapit Bansal, Rishikesh Jha, and Andrew McCallum. 2020 · 2020
Earlier work this paper cites.
BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance. In Proc. of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP . ACL, Online, 217–227
R. Thomas McCoy, Junghyun Min, and Tal Linzen. 2020 · 2020
Earlier work this paper cites.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Earlier work this paper cites.
Which *BERT? A Survey Organizing Contextualized Encoders. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Association for Computational Linguistics, Online, 7516–7533
Patrick Xia, Shijie Wu, and Benjamin Van Durme. 2020 · 2020
Earlier work this paper cites.
The Curse of Performance Instability in Analysis Datasets: Consequences, Source, and Suggestions. In Proc. of the 2020 Conf. on Empirical Methods in Natural Language Processing (EMNLP) . ACL, Online, 8215–8228
Xiang Zhou, Yixin Nie, Hao Tan, and Mohit Bansal. 2020 · 2020
Earlier work this paper cites.
On sensitivity of meta-learning to support data. In Advances in Neural Information Processing Systems , Vol. 34. Curran Associates, Inc., 20447–20460
Mayank Agarwal, Mikhail Yurochkin, and Yuekai Sun. 2021 · 2021
Earlier work this paper cites.
Composable sparse fine-tuning for cross-lingual transfer
Alan Ansell, Edoardo Maria Ponti, Anna Korhonen, and Ivan Vulić. 2021 · 2021
Earlier work this paper cites.
Accounting for Variance in Machine Learning Benchmarks. In Proceedings of Machine Learning and Systems , Vol. 3. 747–769
Xavier Bouthillier, Pierre Delaunay, Mirko Bronzi, Assya Trofimov, Brennan Nichyporuk, Justin Szeto, Naz Sepah, Edward Raff, Kanika Madan, Vikram Voleti, Samira Ebrahimi Kahou, Vincent Michalski, Dmitriy Serdyuk, Tal Arbel, Chris Pal, Gaël Varoquaux, and Pascal Vincent. 2021 · 2021
Earlier work this paper cites.
FLEX: Unifying Evaluation for Few-Shot NLP. In Advances in Neural Information Processing Systems , Vol. 34. Curran Associates, Inc., 15787–15800
Jonathan Bragg, Arman Cohan, Kyle Lo, and Iz Beltagy. 2021 · 2021
Earlier work this paper cites.
On Training Instance Selection for Few-Shot Neural Text Generation. In Proc. of the 59th Annual Meeting of the ACL and the 11th International Joint Conference on Natural Language Processing . ACL, Online, 8–13
Ernie Chang, Xiaoyu Shen, Hui-Syuan Yeh, and Vera Demberg. 2021 · 2021
Earlier work this paper cites.
The Benchmark Lottery
Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. 2021 · 2021
Earlier work this paper cites.
Making Pre-trained Language Models Better Few-shot Learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing . 3816–3830
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021 · 2021
Earlier work this paper cites.
Meta-learning in neural networks: A survey
Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey. 2021 · 2021
Earlier work this paper cites.
Noise Stability Regularization for Improving BERT Fine-tuning. In Proc. of the 2021 Conference of the NAACL: Human Language Technologies . ACL, Online, 3229–3241
Hang Hua, Xingjian Li, Dejing Dou, Chengzhong Xu, and Jiebo Luo. 2021 · 2021
Earlier work this paper cites.
How Emotionally Stable is ALBERT? Testing Robustness with Stochastic Weight Averaging on a Sentiment Analysis Task. In Proc. of the 2nd Workshop on Evaluation and Comparison of NLP Systems . ACL, Punta Cana, Dominican Republic, 16–31
Urja Khurana, Eric Nalisnick, and Antske Fokkens. 2021 · 2021
Earlier work this paper cites.
Reordering Examples Helps during Priming-based Few-Shot Learning. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 . ACL, Online, 4507–4518
Sawan Kumar and Partha Talukdar. 2021 · 2021
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 3045–3059
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Earlier work this paper cites.
STORM+: Fully Adaptive SGD with Recursive Momentum for Nonconvex Optimization. In Advances in Neural Information Processing Systems , Vol. 34. Curran Associates, Inc., 20571–20582
Kfir Levy, Ali Kavis, and Volkan Cevher. 2021 · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing . Association for Computational Linguistics, Online, 4582–4597
Xiang Lisa Li and Percy Liang. 2021 · 2021
Earlier work this paper cites.
On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong Baselines
Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow. 2021 · 2021
Earlier work this paper cites.
CLUES: Few-Shot Learning Evaluation in Natural Language Understanding
Subhabrata Mukherjee, Xiaodong Liu, Guoqing Zheng, Saghar Hosseini, Hao Cheng, Greg Yang, Christopher Meek, Ahmed Hassan Awadallah, and Jianfeng Gao. 2021 · 2021
Earlier work this paper cites.
The PRISMA 2020 statement: an updated guideline for reporting systematic reviews
Matthew J Page, Joanne E McKenzie, Patrick M Bossuyt, Isabelle Boutron, Tammy C Hoffmann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Brennan, et al · 2021
Earlier work this paper cites.
True Few-Shot Learning with Language Models. In Advances in Neural Information Processing Systems , Vol. 34. Curran Associates, Inc., 11054–11070
Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021 · 2021
Earlier work this paper cites.
AdapterFusion: Non-Destructive Task Composition for Transfer Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume . Association for Computational Linguistics, Online, 487–503
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021 · 2021
Earlier work this paper cites.
Problems and opportunities in training deep learning software systems: an analysis of variance. In Proceedings of the 35th IEEE/ACM International Conference on Automated Software Engineering (ASE ’20) . Association for Computing Machinery, New York, NY, USA, 771–783
Hung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier, Jonathan Rosenthal, Lin Tan, Yaoliang Yu, and Nachiappan Nagappan. 2021 · 2021
Earlier work this paper cites.
It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . ACL, Online, 2339–2352
Timo Schick and Hinrich Schütze. 2021 · 2021
Earlier work this paper cites.
Nondeterminism and Instability in Neural Network Optimization. In Proceedings of the 38th International Conference on Machine Learning . PMLR, 9913–9922
Cecilia Summers and Michael J. Dinneen. 2021 · 2021
Earlier work this paper cites.
Variance-reduced First-order Meta-learning for Natural Language Processing Tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . ACL, Online, 2609–2615
Lingxiao Wang, Kevin Huang, Tengyu Ma, Quanquan Gu, and Jing Huang. 2021 · 2021
Earlier work this paper cites.
Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . ACL, Online and Punta Cana, Dominican Republic, 9514–9528
Runxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan, Baobao Chang, Songfang Huang, and Fei Huang. 2021 · 2021
Earlier work this paper cites.
CrossFit: A Few-shot Learning Challenge for Cross-task Generalization in NLP. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . ACL, Online and Punta Cana, Dominican Republic, 7163–7189
Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren. 2021 · 2021
Earlier work this paper cites.
Revisiting Few-sample BERT Fine-tuning. arXiv
Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger, and Yoav Artzi. 2021 · 2021
Earlier work this paper cites.
A Closer Look at Few-Shot Crosslingual Transfer: The Choice of Shots Matters. In Proc. of the 59th Annual Meeting of the ACL and the 11th International Joint Conference on Natural Language Processing . ACL, Online, 5751–5767
Mengjie Zhao, Yi Zhu, Ehsan Shareghi, Ivan Vulić, Roi Reichart, Anna Korhonen, and Hinrich Schütze. 2021b · 2021
Earlier work this paper cites.
Are Larger Pretrained Language Models Uniformly Better? Comparing Performance at the Instance Level. In Findings of the Assoc. for Comp. Linguistics: ACL-IJCNLP 2021 . ACL, Online, 3813–3827
Ruiqi Zhong, Dhruba Ghosh, Dan Klein, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Revisiting Parameter-Efficient Tuning: Are We Really There Yet?. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . ACL, Abu Dhabi, United Arab Emirates, 2612–2626
Guanzheng Chen, Fangyu Liu, Zaiqiao Meng, and Shangsong Liang. 2022a · 2022
Earlier work this paper cites.
How to Distribute Data across Tasks for Meta-Learning?. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 6394–6401
Alexandru Cioba, Michael Bromberg, Qian Wang, Ritwik Niyogi, Georgios Batzolis, Jezabel Garcia, Da-shan Shiu, and Alberto Bernacchia. 2022 · 2022
Earlier work this paper cites.
Underspecification Presents Challenges for Credibility in Modern Machine Learning
Alexander D’Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai, and D. Sculley. 2022 · 2022
Earlier work this paper cites.
A Survey for In-context Learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022 · 2022
Earlier work this paper cites.
LMentry: A Language Model Benchmark of Elementary Language Tasks
Avia Efrat, Or Honovich, and Omer Levy. 2022 · 2022
Cited alongside, same era.
Worst Case Matters for Few-Shot Recognition. In European Conference on Computer Vision . Springer, 99–115
Minghao Fu, Yun-Hao Cao, and Jianxin Wu. 2022 · 2022
Cited alongside, same era.
Sources of Irreproducibility in Machine Learning: A Review
Odd Erik Gundersen, Kevin Coakley, and Christine Kirkpatrick. 2022 · 2022
Cited alongside, same era.
Reducing Model Churn: Stable Re-training of Conversational Agents. In Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue . ACL, Edinburgh, UK, 14–25
Christopher Hidey, Fei Liu, and Rahul Goel. 2022 · 2022
Cited alongside, same era.
In-context Example Selection with Influences
Tai Nguyen and Eric Wong. 2023 · 2023
Closest in time.
Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization. In Findings of the ACL: EMNLP 2023 . ACL, Singapore, 1059–1077
Kaihang Pan, Juncheng Li, Hongye Song, Jun Lin, Xiaozhong Liu, and Siliang Tang. 2023 · 2023
Closest in time.
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka. 2023 · 2023
Closest in time.
Automatic Prompt Optimization with “Gradient Descent” and Beam Search. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . ACL, Singapore, 7957–7968
Reid Pryzant, Dan Iter, Jerry Li, Yin Lee, Chenguang Zhu, and Michael Zeng. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MetaPrompting: Learning to Learn Better Prompts. In Proc. of the 29th International Conference on Computational Linguistics . International Committee on Computational Linguistics, Gyeongju, Republic of Korea, 3251–3262
Yutai Hou, Hongyuan Dong, Xinghao Wang, Bohan Li, and Wanxiang Che. 2022 · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
How to Translate Your Samples and Choose Your Shots? Analyzing Translate-train & Few-shot Cross-lingual Transfer. In Findings of the Assoc. for Comp. Linguistics: NAACL 2022 . ACL, Seattle, United States, 129–150
Iman Jundi and Gabriella Lapesa. 2022 · 2022
Cited alongside, same era.
Meta Learning for Natural Language Processing: A Survey. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . Association for Computational Linguistics, Seattle, United States, 666–684
Hung-yi Lee, Shang-Wen Li, and Thang Vu. 2022b · 2022
Cited alongside, same era.
Surgical Fine-Tuning Improves Adaptation to Distribution Shifts. In NeurIPS 2022 Workshop on Distribution Shifts: Connecting Methods and Applications
Yoonho Lee, Annie S Chen, Fahim Tajwar, Ananya Kumar, Huaxiu Yao, Percy Liang, and Chelsea Finn. 2022a · 2022
Cited alongside, same era.
CAMERO: Consistency Regularized Ensemble of Perturbed Language Models with Weight Sharing. In Proc. of the 60th Annual Meeting of the Association for Computational Linguistics . ACL, Dublin, Ireland, 7162–7175
Chen Liang, Pengcheng He, Yelong Shen, Weizhu Chen, and Tuo Zhao. 2022 · 2022
Cited alongside, same era.
What Makes Good In-Context Examples for GPT-3?. In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures . ACL, Dublin, Ireland and Online, 100–114
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
Cutting Down on Prompts and Parameters: Simple Few-Shot Learning with Language Models. In Findings of the Association for Computational Linguistics: ACL 2022 . Association for Computational Linguistics, Dublin, Ireland, 2824–2835
Robert Logan IV, Ivana Balazevic, Eric Wallace, Fabio Petroni, Sameer Singh, and Sebastian Riedel. 2022 · 2022
Cited alongside, same era.
Learning to Initialize: Can Meta Learning Improve Cross-task Generalization in Prompt Tuning?. In Proc. of the 61st Annual Meeting of the ACL . ACL, Toronto, Canada, 11802–11832
Chengwei Qin, Shafiq Joty, Qian Li, and Ruochen Zhao. 2023a · 2023
Closest in time.
In-context learning with iterative demonstration selection
Chengwei Qin, Aston Zhang, Anirudh Dagar, and Wenming Ye. 2023b · 2023
Closest in time.
Residual Prompt Tuning: improving prompt tuning with residual reparameterization. In Findings of the ACL: ACL 2023 . ACL, Toronto, Canada, 6740–6757
Anastasiia Razdaibiedina, Yuning Mao, Madian Khabsa, Mike Lewis, Rui Hou, Jimmy Ba, and Amjad Almahairi. 2023 · 2023
Closest in time.
Whose Opinions Do Language Models Reflect?. In Proceedings of the 40th International Conference on Machine Learning . PMLR, 29971–30004
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.
Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2023 · 2023
Closest in time.
Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data. In Findings of the Association for Computational Linguistics: EMNLP 2023 . ACL, Singapore, 12113–12139
Kashun Shum, Shizhe Diao, and Tong Zhang. 2023 · 2023
Closest in time.
A Comprehensive Survey of Few-Shot Learning: Evolution, Applications, Challenges, and Opportunities
Yisheng Song, Ting Wang, Puyu Cai, Subrota K Mondal, and Jyoti Prakash Sahoo. 2023 · 2023
Closest in time.
Resources and Few-shot Learners for In-context Learning in Slavic Languages. In Proc. of the 9th Workshop on Slavic Natural Language Processing 2023 (SlavicNLP 2023) . ACL, Dubrovnik, Croatia, 94–105
Michal Štefánik, Marek Kadlčík, Piotr Gramacki, and Petr Sojka. 2023 · 2023
Closest in time.
Evaluating the Zero-shot Robustness of Instruction-tuned Language Models
Jiuding Sun, Chantal Shaib, and Byron C. Wallace. 2023 · 2023
Closest in time.
Two-Stage Fine-Tuning for Improved Bias and Variance for Large Pretrained Language Models. In Proc. of the 61st Annual Meeting of the ACL . ACL, Toronto, Canada, 15746–15761
Lijing Wang, Yingya Li, Timothy Miller, Steven Bethard, and Guergana Savova. 2023 · 2023
Closest in time.
Mind the instructions: a holistic evaluation of consistency and interactions in prompt-based learning. In Proc. of the 27th Conference on Computational Natural Language Learning (CoNLL) . ACL, Singapore, 294–313
Lucas Weber, Elia Bruni, and Dieuwke Hupkes. 2023 · 2023
Closest in time.
Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics . ACL, Toronto, Canada, 1423–1436
Zhiyong Wu, Yaoxiang Wang, Jiacheng Ye, and Lingpeng Kong. 2023b · 2023
Closest in time.
MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics . ACL, Toronto, Canada, 11445–11465
Zhiyang Xu, Ying Shen, and Lifu Huang. 2023 · 2023
Closest in time.
Representative Demonstration Selection for In-Context Learning with Two-Stage Determinantal Point Process. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . ACL, Singapore, 5443–5456
Zhao Yang, Yuanzhe Zhang, Dianbo Sui, Cao Liu, Jun Zhao, and Kang Liu. 2023 · 2023
Closest in time.
Compositional exemplars for in-context learning. In International Conference on Machine Learning . PMLR, 39818–39833
Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong. 2023 · 2023
Closest in time.
Cold-Start Data Selection for Better Few-shot Language Model Fine-tuning: A Prompt-based Uncertainty Propagation Approach. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics . ACL, Toronto, Canada, 2499–2521
Yue Yu, Rongzhi Zhang, Ran Xu, Jieyu Zhang, Jiaming Shen, and Chao Zhang. 2023 · 2023
Closest in time.
Instruction tuning for large language models: A survey
Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, et al · 2023
Closest in time.
Auto-Instruct: Automatic Instruction Generation and Ranking for Black-Box Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2023 . ACL, Singapore, 9850–9867
Zhihan Zhang, Shuohang Wang, Wenhao Yu, Yichong Xu, Dan Iter, Qingkai Zeng, Yang Liu, Chenguang Zhu, and Meng Jiang. 2023c · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Closest in time.
Large Language Models Are Not Robust Multiple Choice Selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2023 · 2023
Closest in time.
Can chatgpt understand too? a comparative study on chatgpt and fine-tuned bert
Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, and Dacheng Tao. 2023 · 2023
Closest in time.
Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering
Han Zhou, Xingchen Wan, Lev Proleev, Diana Mincu, Jilin Chen, Katherine A. Heller, and Subhrajit Roy. 2023 · 2023
Closest in time.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, et al · 2023
Closest in time.
Fool your (vision and) language model with embarrassingly simple permutations
Yongshuo Zong, Tingyang Yu, Bingchen Zhao, Ruchika Chavhan, and Timothy Hospedales. 2023 · 2023
Closest in time.
Designing Informative Metrics for Few-Shot Example Selection
Rishabh Adiga, Lakshminarayanan Subramanian, and Varun Chandrasekaran. 2024 · 2024
Closest in time.
When benchmarks are targets: Revealing the sensitivity of large language model leaderboards
Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Alsubaie, Yusef Almushaykeh, Faisal Mirza, Nouf Alotaibi, Nora Altwairesh, Areeb Alowisheq, et al · 2024
Closest in time.
Prompt Design Matters for Computational Social Science Tasks but in Unpredictable Ways
Shubham Atreja, Joshua Ashkinaze, Lingyao Li, Julia Mendelsohn, and Libby Hemphill. 2024 · 2024
Closest in time.
Measuring model variability using robust non-parametric testing
Sinjini Banerjee, Tim Marrinan, Reilly Cannon, Tony Chiang, and Anand D. Sarwate. 2024 · 2024
Closest in time.
In-context learning with long-context models: An in-depth exploration
Amanda Bertsch, Maor Ivgi, Uri Alon, Jonathan Berant, Matthew R Gormley, and Graham Neubig. 2024 · 2024
Closest in time.
Lessons from the Trenches on Reproducible Evaluation of Language Models
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, et al · 2024
Closest in time.
Changing answer order can decrease mmlu accuracy
Vipul Gupta, David Pantoja, Candace Ross, Adina Williams, and Megan Ung. 2024 · 2024
Closest in time.
Parameter-efficient fine-tuning for large models: A comprehensive survey
Zeyu Han, Chao Gao, Jinyang Liu, Sai Qian Zhang, et al · 2024
Closest in time.
Submodular-based In-context Example Selection for LLMs-based Machine Translation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) . ELRA and ICCL, Torino, Italia, 15398–15409
Baijun Ji, Xiangyu Duan, Zhenyu Qiu, Tong Zhang, Junhui Li, Hao Yang, and Min Zhang. 2024 · 2024
Closest in time.
Vision-Language Instruction Tuning: A Review and Analysis
Chen Li, Yixiao Ge, Dian Li, and Ying Shan. 2024 · 2024
Closest in time.
StablePT: Towards Stable Prompting for Few-shot Learning via Input Separation
Xiaoming Liu, Chen Liu, Zhaohan Zhang, Chengzhengxu Li, Longtian Wang, Yu Lan, and Chao Shen. 2024b · 2024
Closest in time.
Let’s Learn Step by Step: Enhancing In-Context Learning Ability with Curriculum Learning
Yinpeng Liu, Jiawei Liu, Xiang Shi, Qikai Cheng, and Wei Lu. 2024a · 2024
Closest in time.
Quantifying Variance in Evaluation Benchmarks
Lovish Madaan, Aaditya K Singh, Rylan Schaeffer, Andrew Poulton, Sanmi Koyejo, Pontus Stenetorp, Sharan Narang, and Dieuwke Hupkes. 2024 · 2024
Closest in time.
Branislav Pecher, Jan Cegin, Robert Belanec, Jakub Simko, Ivan Srba, and Maria Bielikova. 2024a · 2024
Closest in time.
Branislav Pecher, Ivan Srba, and Maria Bielikova. 2024b · 2024
Closest in time.
Branislav Pecher, Ivan Srba, and Maria Bielikova. 2024c · 2024
Closest in time.
Automatic Combination of Sample Selection Strategies for Few-Shot Learning
Branislav Pecher, Ivan Srba, Maria Bielikova, and Joaquin Vanschoren. 2024d · 2024
Closest in time.
Revisiting Demonstration Selection Strategies in In-Context Learning
Keqin Peng, Liang Ding, Yancheng Yuan, Xuebo Liu, Min Zhang, Yuanxin Ouyang, and Dacheng Tao. 2024 · 2024
Closest in time.
Efficient multi-prompt evaluation of LLMs
Felipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. 2024 · 2024
Closest in time.
Sub-SA: Strengthen In-context Learning via Submodular Selective Annotation
Jian Qian, Miao Sun, Sifan Zhou, Ziyu Zhao, Ruizhi Hun, and Patrick Chiang. 2024 · 2024
Closest in time.
Beyond Performance: Quantifying and Mitigating Label Bias in LLMs. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . ACL, Mexico City, Mexico, 6784–6798
Yuval Reif and Roy Schwartz. 2024 · 2024
Closest in time.
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts. In Proceedings of the 2024 Conference of the NAACL: Human Language Technologies . ACL, Mexico City, Mexico, 4936–4953
Sai Ashish Somayajula, Youwei Liang, Li Zhang, Abhishek Singh, and Pengtao Xie. 2024 · 2024
Closest in time.
Mind your format: Towards consistent evaluation of in-context learning improvements
Anton Voronov, Lena Wolf, and Max Ryabinin. 2024 · 2024
Closest in time.
Unveiling Selection Biases: Exploring Order and Token Sensitivity in Large Language Models
Sheng-Lun Wei, Cheng-Kuang Wu, Hen-Hsen Huang, and Hsin-Hsi Chen. 2024 · 2024
Closest in time.
Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars. In ICML 2024 Workshop on In-Context Learning
Zhaoxuan Wu, Xiaoqiang Lin, Zhongxiang Dai, Wenyang Hu, Yao Shu, See-Kiong Ng, Patrick Jaillet, and Bryan Kian Hsiang Low. 2024 · 2024
Closest in time.
Misconfidence-based demonstration selection for llm in-context learning
Shangqing Xu and Chao Zhang. 2024 · 2024
Closest in time.
In-context learning with retrieved demonstrations for language models: A survey
Xin Xu, Yue Liu, Panupong Pasupat, Mehran Kazemi, et al · 2024
Closest in time.
Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement
Pengwei Zhan, Zhen Xu, Qian Tan, Jie Song, and Ru Xie. 2024 · 2024
Closest in time.
Batch-ICL: Effective, Efficient, and Order-Agnostic In-Context Learning
Kaiyi Zhang, Ang Lv, Yuhan Chen, Hansen Ha, Tao Xu, and Rui Yan. 2024b · 2024
Closest in time.
The Impact of Demonstrations on Multilingual In-Context Learning: A Multidimensional Analysis
Miaoran Zhang, Vagrant Gautam, Mingyang Wang, Jesujoba O Alabi, Xiaoyu Shen, Dietrich Klakow, and Marius Mosbach. 2024a · 2024
Closest in time.
Correcting Language Model Bias for Text Classification in True Zero-Shot Learning. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) . ELRA and ICCL, Torino, Italia, 4036–4046
Feng Zhao, Wan Xianlin, Cheng Yan, and Chu Kiong Loo. 2024b · 2024
Closest in time.
NoisyICL: A Little Noise in Model Parameters Calibrates In-context Learning
Yufeng Zhao, Yoshihiro Sakai, and Naoya Inoue. 2024a · 2024
Closest in time.
Towards Robust In-Context Learning for Machine Translation with Large Language Models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) . ELRA and ICCL, Torino, Italia, 16619–16629
Shaolin Zhu, Menglong Cui, and Deyi Xiong. 2024 · 2024
Closest in time.