Fetching the paper…
Reading the bibliography…
Finetuning large language models inflates the costs of NLU applications and remains the bottleneck of development cycles.
Bert for joint intent classification and slot filling
Qian Chen, Zhu Zhuo, and Wen Wang. 2019 · 1902
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Evaluation of spoken language systems: The atis domain
Patti Price. 1990 · 1990
Earlier work this paper cites.
Smaller coresets for k-median and k-means clustering
Sariel Har-Peled and Akash Kushal. 2005 · 2005
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Self-paced learning for latent variable models
M Kumar, Benjamin Packer, and Daphne Koller. 2010 · 2010
Earlier work this paper cites.
What is left to be understood in atis?
Gokhan Tur, Dilek Hakkani-Tür, and Larry Heck. 2010 · 2010
Earlier work this paper cites.
Variance reduction in sgd by distributed importance sampling
Guillaume Alain, Alex Lamb, Chinnadhurai Sankar, Aaron Courville, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015 · 2015
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017 · 2017
Earlier work this paper cites.
Bayesian coreset construction via greedy iterative geodesic ascent
Trevor Campbell and Tamara Broderick. 2018 · 2018
Earlier work this paper cites.
Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, et al. 2018 · 2018
Earlier work this paper cites.
Not all samples are created equal: Deep learning with importance sampling
Angelos Katharopoulos and François Fleuret. 2018 · 2018
Earlier work this paper cites.
Mixed precision training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. 2018 · 2018
Earlier work this paper cites.
On coresets for logistic regression
Alexander Munteanu, Chris Schwiegelshohn, Christian Sohler, and David Woodruff. 2018 · 2018
Earlier work this paper cites.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon. 2018 · 2018
Cited alongside, same era.
Deep batch active learning by diverse, uncertain gradient lower bounds
Jordan T Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Revealing the dark secrets of bert
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Cited alongside, same era.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020 · 2020
Later among the works it cites.
Accelerating training of transformer-based language models with progressive layer dropping
Minjia Zhang and Yuxiong He. 2020 · 2020
Later among the works it cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021 · 2021
Later among the works it cites.
Mtop: A comprehensive multilingual task-oriented semantic parsing benchmark
Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, and Yashar Mehdad. 2021 · 2021
Later among the works it cites.
Deep learning on a data diet: Finding important examples early in training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Competence-based curriculum learning for neural machine translation
Emmanouil Antonios Platanios, Otilia Stretcu, Graham Neubig, Barnabás Poczós, and Tom Mitchell. 2019 · 2019
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019 · 2019
Cited alongside, same era.
Slurp: A spoken language understanding resource package
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, and Verena Rieser. 2020 · 2020
Cited alongside, same era.
Selection via proxy: Efficient data selection for deep learning
C Coleman, C Yeh, S Mussmann, B Mirzasoleiman, P Bailis, P Liang, J Leskovec, and M Zaharia. 2020 · 2020
Cited alongside, same era.
Deep active learning for sequence labeling based on diversity and uncertainty in gradient
Yekyung Kim. 2020 · 2020
Cited alongside, same era.
Do we need to create big datasets to learn a task?
Swaroop Mishra and Bhavdeep Singh Sachdeva. 2020 · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite. 2021 · 2021
Later among the works it cites.
Accelerating deep learning with dynamic data pruning
Ravi S Raju, Kyle Daruwalla, and Mikko Lipasti. 2021 · 2021
Later among the works it cites.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. 2021 · 2021
Later among the works it cites.
Training dynamics for curriculum learning: A study on monolingual and cross-lingual NLU
Fenia Christopoulou, Gerasimos Lampouras, and Ignacio Iacobacci. 2022 · 2022
Later among the works it cites.
Bert on a data diet: Finding important examples by gradient-based pruning
Mohsen Fayyaz, Ehsan Aghazadeh, Ali Modarressi, Mohammad Taher Pilehvar, Yadollah Yaghoobzadeh, Samira Ebrahimi Kahou, and CIFAR Mila. 2022 · 2022
Later among the works it cites.
Prioritized training on points that are learnable, worth learning, and not yet learnt
Sören Mindermann, Jan M Brauner, Muhammed T Razzak, Mrinank Sharma, Andreas Kirsch, Winnie Xu, Benedikt Höltgen, Aidan N Gomez, Adrien Morisot, Sebastian Farquhar, et al. 2022 · 2022
Later among the works it cites.
Active learning is a strong baseline for data subset selection
Dongmin Park, Dimitris Papailiopoulos, and Kangwook Lee · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S Morcos. 2022 · 2022
Later among the works it cites.
Dataset pruning: Reducing training data by examining generalization influence
Shuo Yang, Zeke Xie, Hanyu Peng, Min Xu, Mingming Sun, and Ping Li. 2022 · 2022
Later among the works it cites.
Deep learning on a healthy data diet: Finding important examples for fairness
Abdelrahman Zayed, Prasanna Parthasarathi, Goncalo Mordido, Hamid Palangi, Samira Shabanian, and Sarath Chandar. 2022 · 2022
Later among the works it cites.