Fetching the paper…
Reading the bibliography…
Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data; however, using only scale to improve performance means that resource consumption also grows.
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
The State of Sparsity in Deep Neural Networks
Trevor Gale, Erich Elsen, and Sara Hooker. 2019 · 1902
Earlier work this paper cites.
Generating Long Sequences with Sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Discriminative Active Learning
Daniel Gissin and Shai Shalev-Shwartz. 2019 · 1907
Earlier work this paper cites.
Curriculum Learning for Domain Adaptation in Neural Machine Translation
Xuan Zhang, Pamela Shapiro, Gaurav Kumar, Paul McNamee, Marine Carpuat, and Kevin Duh. 2019 · 1915
Earlier work this paper cites.
Optimal Brain Damage
Yann LeCun, John Denker, and Sara Solla. 1989 · 1989
Earlier work this paper cites.
Adaptive Mixtures of Local Experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991 · 1991
Earlier work this paper cites.
Learning and development in neural networks: The importance of starting small
Jeffrey L. Elman. 1993 · 1993
Earlier work this paper cites.
A Sequential Algorithm for Training Text Classifiers
David D. Lewis and William A. Gale. 1994 · 1994
Earlier work this paper cites.
Multitask Learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2002
Earlier work this paper cites.
Active Learning for Statistical Natural Language Parsing
Min Tang, Xiaoqiang Luo, and Salim Roukos. 2002 · 2002
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Towards efficient supercomputing: A quest for the right metric
C-H Hsu, W-C Feng, and Jeremy S Archuleta. 2005 · 2005
Earlier work this paper cites.
Power Consumption Variation over Activation Functions
Leon Derczynski. 2020 · 2006
Earlier work this paper cites.
The Computational Limits of Deep Learning
Neil C. Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F. Manso. 2020 · 2007
Earlier work this paper cites.
Sample Selection Bias Correction Theory
Corinna Cortes, Mehryar Mohri, Michael Riley, and Afshin Rostamizadeh. 2008 · 2008
Earlier work this paper cites.
Active Learning with Real Annotation Costs
Burr Settles, Mark Craven, and Lewis Friedland. 2008 · 2008
Earlier work this paper cites.
Curriculum Learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Nur Ahmed and Muntasir Wahed. 2020 · 2010
Earlier work this paper cites.
Active Learning with Clustering
Zalán Bodó, Zsolt Minier, and Lehel Csató. 2011 · 2010
Earlier work this paper cites.
Characterising Bias in Compressed Models
Sara Hooker, Nyalleng Moorosi, Gregory Clark, Samy Bengio, and Emily Denton. 2020 · 2010
Earlier work this paper cites.
Self-Paced Learning for Latent Variable Models
M. Kumar, Benjamin Packer, and Daphne Koller. 2010 · 2010
Earlier work this paper cites.
Active Learning , volume 18 of Synthesis Lectures on Artificial Intelligence and Machine Learning
Burr Settles. 2012 · 2012
Earlier work this paper cites.
Practical Bayesian Optimization of Machine Learning Algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams. 2012 · 2012
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Efficient and Robust Automated Machine Learning
Matthias Feurer, Aaron Klein, Katharina Eggensperger, Jost Springenberg, Manuel Blum, and Frank Hutter. 2015 · 2015
Earlier work this paper cites.
Learning both Weights and Connections for Efficient Neural Networks
Song Han, Jeff Pool, John Tran, and William Dally. 2015 · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Non-stochastic Best Arm Identification and Hyperparameter Optimization
Kevin Jamieson and Ameet Talwalkar. 2016 · 2016
Earlier work this paper cites.
Sequence-Level Knowledge Distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Learning to Prune Deep Neural Networks via Layer-wise Optimal Brain Surgeon
Xin Dong, Shangyu Chen, and Sinno Pan. 2017 · 2017
Earlier work this paper cites.
Deep Bayesian Active Learning with Image Data
Yarin Gal, Riashat Islam, and Zoubin Ghahramani. 2017 · 2017
Earlier work this paper cites.
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017 · 2017
Earlier work this paper cites.
Learning to Generate Reviews and Discovering Sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever. 2017 · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017 · 2017
Earlier work this paper cites.
Reporting Score Distributions Makes a Difference: Performance Study of LSTM-networks for Sequence Tagging
Nils Reimers and Iryna Gurevych. 2017 · 2017
Earlier work this paper cites.
An Overview of Multi-Task Learning in Deep Neural Networks
Sebastian Ruder. 2017 · 2017
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
Search Engine Guided Non-Parametric Neural Machine Translation
Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor O. K. Li. 2018 · 2018
Earlier work this paper cites.
Measuring the Intrinsic Dimension of Objective Landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski. 2018 · 2018
Earlier work this paper cites.
Learning to Actively Learn Neural Machine Translation
Ming Liu, Wray Buntine, and Gholamreza Haffari. 2018 · 2018
Earlier work this paper cites.
Learning Sparse Neural Networks through L 0 Regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2018 · 2018
Earlier work this paper cites.
Ensuring More Sustainable Reporting in Europe Using Non-Financial Disclosure—De Facto and De Jure Evidence
Francesca Manes-Rossi, Adriana Tiron-Tudor, Giuseppe Nicolò, and Gianluca Zanellato. 2018 · 2018
Earlier work this paper cites.
Deep Contextualized Word Representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Pieces of Eight: 8-bit Neural Machine Translation
Jerry Quinn and Miguel Ballesteros. 2018 · 2018
Earlier work this paper cites.
Active Learning for Convolutional Neural Networks: A Core-Set Approach
Ozan Sener and Silvio Savarese. 2018 · 2018
Earlier work this paper cites.
Self-Attention with Relative Position Representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Earlier work this paper cites.
SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Earlier work this paper cites.
Simple, Scalable Adaptation for Neural Machine Translation
Ankur Bapna and Orhan Firat. 2019 · 2019
Earlier work this paper cites.
Efficient 8-Bit Quantization of Transformer Neural Machine Language Translation Model
Aishwarya Bhandare, Vamsi Sripathi, Deepthi Karkada, Vivek Menon, Sun Choi, Kushal Datta, and Vikram Saletore. 2019 · 2019
Earlier work this paper cites.
Adaptively Sparse Transformers
Gonçalo M. Correia, Vlad Niculae, and André F. T. Martins. 2019 · 2019
Earlier work this paper cites.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Earlier work this paper cites.
Universal Transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Show Your Work: Improved Reporting of Experimental Results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. 2019 · 2019
Earlier work this paper cites.
Rethinking ImageNet pre-training
Kaiming He, Ross Girshick, and Piotr Dollár. 2019 · 2019
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
From Research to Production and Back: Ludicrously Fast Neural Machine Translation
Young Jin Kim, Marcin Junczys-Dowmunt, Hany Hassan, Alham Fikri Aji, Kenneth Heafield, Roman Grundkiewicz, and Nikolay Bogoychev. 2019 · 2019
Earlier work this paper cites.
BatchBALD: Efficient and Diverse Batch Acquisition for Deep Bayesian Active Learning
Andreas Kirsch, Joost van Amersfoort, and Yarin Gal. 2019 · 2019
Earlier work this paper cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Earlier work this paper cites.
Practical Obstacles to Deploying Active Learning
David Lowell, Zachary C. Lipton, and Byron C. Wallace. 2019 · 2019
Earlier work this paper cites.
Quantifying the Carbon Emissions of Machine Learning
Sasha Luccioni, Victor Schmidt, Alexandre Lacoste, and Thomas Dandres. 2019 · 2019
Earlier work this paper cites.
Are Sixteen Heads Really Better than One?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Earlier work this paper cites.
Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization
Hesham Mostafa and Xin Wang. 2019 · 2019
Earlier work this paper cites.
Sparse Sequence-to-Sequence Models
Ben Peters, Vlad Niculae, and André F. T. Martins. 2019 · 2019
Earlier work this paper cites.
Language Models as Knowledge Bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Earlier work this paper cites.
Energy and Policy Considerations for Deep Learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 2019
Earlier work this paper cites.
Simple and Effective Curriculum Pointer-Generator Networks for Reading Comprehension over Long Narratives
Yi Tay, Shuohang Wang, Anh Tuan Luu, Jie Fu, Minh C. Phan, Xingdi Yuan, Jinfeng Rao, Siu Cheung Hui, and Aston Zhang. 2019 · 2019
Earlier work this paper cites.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Earlier work this paper cites.
Q8BERT: Quantized 8Bit BERT
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 2019
Earlier work this paper cites.
ETC: Encoding Long and Structured Inputs in Transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020 · 2020
Earlier work this paper cites.
CarbonTracker: Tracking and predicting the carbon footprint of training deep learning models
Lasse F Wolff Anthony, Benjamin Kanding, and Raghavendra Selvan. 2020 · 2020
Earlier work this paper cites.
Deep Batch Active Learning by Diverse, Uncertain Gradient Lower Bounds
Jordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal. 2020 · 2020
Earlier work this paper cites.
Proceedings of the Fourth Workshop on Neural Generation and Translation . Association for Computational Linguistics, Online
Alexandra Birch, Andrew Finch, Hiroaki Hayashi, Kenneth Heafield, Marcin Junczys-Dowmunt, Ioannis Konstas, Xian Li, Graham Neubig, and Yusuke Oda, editors. 2020 · 2020
Earlier work this paper cites.
What is the State of Neural Network Pruning?
Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag. 2020 · 2020
Earlier work this paper cites.
Edinburgh’s Submissions to the 2020 Machine Translation Efficiency Task
Nikolay Bogoychev, Roman Grundkiewicz, Alham Fikri Aji, Maximiliana Behnke, Kenneth Heafield, Sidharth Kashyap, Emmanouil-Ioannis Farsarakis, and Mateusz Chudyk. 2020 · 2020
Earlier work this paper cites.
Survey of machine-learning experimental methods at NeurIPS2019 and ICLR2020
Xavier Bouthillier and Gaël Varoquaux. 2020 · 2020
Earlier work this paper cites.
Towards Accurate and Reliable Energy Measurement of NLP Models
Qingqing Cao, Aruna Balasubramanian, and Niranjan Balasubramanian. 2020 · 2020
Earlier work this paper cites.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Earlier work this paper cites.
Balancing Cost and Benefit with Tied-Multi Transformers
Raj Dabre, Raphael Rubino, and Atsushi Fujita. 2020 · 2020
Earlier work this paper cites.
SMYRF - Efficient Attention using Asymmetric Clustering
Giannis Daras, Nikita Kitaev, Augustus Odena, and Alexandros G Dimakis. 2020 · 2020
Cited alongside, same era.
Location Attention for Extrapolation to Longer Sequences
Yann Dubois, Gautier Dagan, Dieuwke Hupkes, and Elia Bruni. 2020 · 2020
Cited alongside, same era.
Active Learning for BERT: An Empirical Study
Liat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch, Lena Dankin, Leshem Choshen, Marina Danilevsky, Ranit Aharonov, Yoav Katz, and Noam Slonim. 2020 · 2020
Cited alongside, same era.
Depth-Adaptive Transformer
Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Reducing Transformer Depth on Demand with Structured Dropout
Angela Fan, Edouard Grave, and Armand Joulin. 2020 · 2020
Cited alongside, same era.
Compressing BERT: Studying the Effects of Weight Pruning on Transfer Learning
Mitchell Gordon, Kevin Duh, and Nicholas Andrews. 2020 · 2020
AdapterDrop: On the Efficiency of Adapters in Transformers
Andreas Rücklé, Gregor Geigle, Max Glockner, Tilman Beck, Jonas Pfeiffer, Nils Reimers, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
ProFormer: Towards On-Device LSH Projection Based Transformers
Chinnadhurai Sankar, Sujith Ravi, and Zornitsa Kozareva. 2021 · 2021
Later among the works it cites.
It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
Timo Schick and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
Metadata Archaeology: Unearthing Data Subsets by Leveraging Training Dynamics
Shoaib Ahmed Siddiqui, Nitarshan Rajkumar, Tegan Maharaj, David Krueger, and Sara Hooker. 2021 · 2021
Later among the works it cites.
Towards a Comprehensive Understanding and Accurate Evaluation of Societal Biases in Pre-Trained Transformers
Andrew Silva, Pradyumna Tambwekar, and Matthew Gombolay. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A^3: Accelerating Attention Mechanisms in Neural Networks with Approximation
Tae Jun Ham, Sung Jun Jung, Seonghak Kim, Young H. Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee, Kyoung Park, Jae W. Lee, and Deog-Kyoon Jeong. 2020 · 2020
Cited alongside, same era.
Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning
Peter Henderson, Jieru Hu, Joshua Romoff, Emma Brunskill, Dan Jurafsky, and Joelle Pineau. 2020 · 2020
Cited alongside, same era.
SqueezeBERT: What can computer vision teach NLP about efficient neural networks?
Forrest Iandola, Albert Shaw, Ravi Krishna, and Kurt Keutzer. 2020 · 2020
Cited alongside, same era.
TinyBERT: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Cited alongside, same era.
Generalization through Memorization: Nearest Neighbor Language Models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Cited alongside, same era.
Does Knowledge Distillation Really Work?
Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A Alemi, and Andrew G Wilson. 2021 · 2021
Later among the works it cites.
Training with Quantization Noise for Extreme Model Compression
Pierre Stock, Angela Fan, Benjamin Graham, Edouard Grave, Rémi Gribonval, Herve Jegou, and Armand Joulin. 2021 · 2021
Later among the works it cites.
Training Neural Networks with Fixed Sparse Masks
Yi-Lin Sung, Varun Nair, and Colin A Raffel. 2021 · 2021
Later among the works it cites.
EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
Thierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia, En-Yu Yang, Marco Donato, Victor Sanh, Paul Whatmough, Alexander M. Rush, David Brooks, and Gu-Yeon Wei. 2021 · 2021
Later among the works it cites.
Long Range Arena : A Benchmark for Efficient Transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. 2021 · 2021
Later among the works it cites.
Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
Kale-ab Tessera, Sara Hooker, and Benjamin Rosman. 2021 · 2021
Later among the works it cites.
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
Hanrui Wang, Zhekai Zhang, and Song Han. 2021a · 2021
Later among the works it cites.
Meta-learning Hyperparameter Performance Prediction with Neural Processes
Ying Wei, Peilin Zhao, and Junzhou Huang. 2021 · 2021
Later among the works it cites.
Beyond Preserved Accuracy: Evaluating Loyalty and Robustness of BERT Compression
Canwen Xu, Wangchunshu Zhou, Tao Ge, Ke Xu, Julian McAuley, and Furu Wei. 2021 · 2021
Later among the works it cites.
Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Ge Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. 2021 · 2021
Later among the works it cites.
Adaptive Semiparametric Language Models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. 2021 · 2021
Later among the works it cites.
Prune Once for All: Sparse Pre-Trained Language Models
Ofir Zafrir, Ariel Larey, Guy Boudoukh, Haihao Shen, and Moshe Wasserblat. 2021 · 2021
Later among the works it cites.
Shuangfei Zhai, Walter Talbott, Nitish Srivastava, Chen Huang, Hanlin Goh, Ruixiang Zhang, and Josh Susskind. 2021 · 2021
Later among the works it cites.
Combining Curriculum Learning and Knowledge Distillation for Dialogue Generation
Qingqing Zhu, Xiuying Chen, Pengfei Wu, JunFei Liu, and Dongyan Zhao. 2021 · 2021
Later among the works it cites.
Auto-Pytorch: Multi-Fidelity MetaLearning for Efficient and Robust AutoDL
Lucas Zimmer, Marius Lindauer, and Frank Hutter. 2021 · 2021
Later among the works it cites.
Estimating Example Difficulty Using Variance of Gradients
Chirag Agarwal, Daniel D’souza, and Sara Hooker. 2022 · 2022
Closest in time.
How does the pre-training objective affect what large language models learn about linguistic properties?
Ahmed Alajrami and Nikolaos Aletras. 2022 · 2022
Closest in time.
Neuro-Symbolic Language Modeling with Automaton-augmented Retrieval
Uri Alon, Frank Xu, Junxian He, Sudipta Sengupta, Dan Roth, and Graham Neubig. 2022 · 2022
Closest in time.
ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning
Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q. Tran, Dara Bahri, Jianmo Ni, Jai Gupta, Kai Hui, Sebastian Ruder, and Donald Metzler. 2022 · 2022
Closest in time.
PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts
Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Fries, Maged Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Dragomir Radev, Mike Tian-Jian Jiang, and Alexander Rush. 2022 · 2022
Closest in time.
Pathways: Asynchronous Distributed Dataflow for ML
Paul Barham, Aakanksha Chowdhery, Jeff Dean, Sanjay Ghemawat, Steven Hand, Daniel Hurt, Michael Isard, Hyeontaek Lim, Ruoming Pang, Sudip Roy, Brennan Saeta, Parker Schuh, Ryan Sepassi, Laurent Shafey, Chandu Thekkath, and Yonghui Wu. 2022 · 2022
Closest in time.
Modeling the Machine Learning Multiverse
Samuel Bell, Onno Kampman, Jesse Dodge, and Neil D Lawrence. 2022 · 2022
Closest in time.
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. 2022 · 2022
Closest in time.
Improving Language Models by Retrieving from Trillions of Tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego De Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack Rae, Erich Elsen, and Laurent Sifre. 2022 · 2022
Closest in time.
Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
Beidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang, Zhao Song, Atri Rudra, and Christopher Re. 2022 · 2022
Closest in time.
PaLM: Scaling Language Modeling with Pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022 · 2022
Closest in time.
The Efficiency Misnomer
Mostafa Dehghani, Yi Tay, Anurag Arnab, Lucas Beyer, and Ashish Vaswani. 2022 · 2022
Closest in time.
Measuring the Carbon Intensity of AI in Cloud Instances
Jesse Dodge, Taylor Prewitt, Remi Tachet des Combes, Erika Odmark, Roy Schwartz, Emma Strubell, Alexandra Sasha Luccioni, Noah A. Smith, Nicole DeCario, and Will Buchanan. 2022 · 2022
Closest in time.
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, Liam Fedus, Maarten P Bosma, Zongwei Zhou, Tao Wang, Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathleen Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang, Quoc Le, Yonghui Wu, Zhifeng Chen, and Claire Cui. 2022 · 2022
Closest in time.
Understanding Dataset Difficulty with 𝒱 \mathcal{V} -Usable Information
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. 2022 · 2022
Closest in time.
Auto-Sklearn 2.0: Hands-free AutoML via Meta-Learning
Matthias Feurer, Katharina Eggensperger, Stefan Falkner, Marius Lindauer, and Frank Hutter. 2022 · 2022
Closest in time.
EdgeFormer: A Parameter-Efficient Transformer for On-Device Seq2seq Generation
Tao Ge, Si-Qing Chen, and Furu Wei. 2022 · 2022
Closest in time.
Sources of Irreproducibility in Machine Learning: A Review
Odd Erik Gundersen, Kevin Coakley, and Christine Kirkpatrick. 2022 · 2022
Closest in time.
Diagonal State Spaces are as Effective as Structured State Spaces
Ankit Gupta, Albert Gu, and Jonathan Berant. 2022 · 2022
Closest in time.
How Much Does Attention Actually Attend? Questioning the Importance of Attention in Pretrained Transformers
Michael Hassid, Hao Peng, Daniel Rotem, Jungo Kasai, Ivan Montero, Noah A. Smith, and Roy Schwartz. 2022 · 2022
Closest in time.
Towards Climate Awareness in NLP Research
Daniel Hershcovich, Nicolas Webersinke, Mathias Kraus, Julia Bingler, and Markus Leippold. 2022 · 2022
Closest in time.
Bridging Fairness and Environmental Sustainability in Natural Language Processing
Marius Hessenthaler, Emma Strubell, Dirk Hovy, and Anne Lauscher. 2022 · 2022
Closest in time.
The Forward-Forward Algorithm: Some Preliminary Investigations
Geoffrey Hinton. 2022 · 2022
Closest in time.
An Empirical Analysis of Compute-Optimal Large Language Model Training
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack William Rae, and Laurent Sifre. 2022 · 2022
Closest in time.
LoRA: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Closest in time.
How Well Do Sparse ImageNet Models Transfer?
Eugenia Iofinova, Alexandra Peste, Mark Kurtz, and Dan Alistarh. 2022 · 2022
Closest in time.
Prompt-free and Efficient Few-shot Learning with Language Models
Rabeeh Karimi Mahabadi, Luke Zettlemoyer, James Henderson, Lambert Mathias, Marzieh Saeidi, Veselin Stoyanov, and Majid Yazdani. 2022 · 2022
Closest in time.
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022 · 2022
Closest in time.
FP8 Quantization: The Power of the Exponent
Andrey Kuzmin, Mart Van Baalen, Yuwei Ren, Markus Nagel, Jorn Peters, and Tijmen Blankevoort. 2022 · 2022
Closest in time.
A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model
Imad Lakim, Ebtesam Almazrouei, Ibrahim Abualhaol, Merouane Debbah, and Julien Launay. 2022 · 2022
Closest in time.
FNet: Mixing tokens with Fourier transforms
James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon. 2022 · 2022
Closest in time.
SMAC3: A Versatile Bayesian Optimization Package for Hyperparameter Optimization
Marius Lindauer, Katharina Eggensperger, Matthias Feurer, André Biedenkapp, Difan Deng, Carolin Benjamins, Tim Ruhkopf, René Sass, and Frank Hutter. 2022 · 2022
Closest in time.
Towards Efficient NLP: A Standard Evaluation and A Strong Baseline
Xiangyang Liu, Tianxiang Sun, Junliang He, Jiawen Wu, Lingling Wu, Xinyu Zhang, Hao Jiang, Zhao Cao, Xuanjing Huang, and Xipeng Qiu. 2022b · 2022
Closest in time.
Chunk-based Nearest Neighbor Machine Translation
Pedro Henrique Martins, Zita Marinho, and André F. T. Martins. 2022c · 2022
Closest in time.
Fast Nearest Neighbor Machine Translation
Yuxian Meng, Xiaoya Li, Xiayu Zheng, Fei Wu, Xiaofei Sun, Tianwei Zhang, and Jiwei Li. 2022 · 2022
Closest in time.
What Do Compressed Multilingual Machine Translation Models Forget?
Alireza Mohammadshahi, Vassilina Nikoulina, Alexandre Berard, Caroline Brun, James Henderson, and Laurent Besacier. 2022 · 2022
Closest in time.
Adaptable Adapters
Nafise Moosavi, Quentin Delfosse, Kristian Kersting, and Iryna Gurevych. 2022 · 2022
Closest in time.
Multimodal Contrastive Learning with LIMoE: The Language-Image Mixture of Experts
Basil Mustafa, Carlos Riquelme Ruiz, Joan Puigcerver, Rodolphe Jenatton, and Neil Houlsby. 2022 · 2022
Closest in time.
8-bit Numerical Formats for Deep Neural Networks
Badreddine Noune, Philip Jones, Daniel Justus, Dominic Masters, and Carlo Luschi. 2022 · 2022
Closest in time.
Intriguing Properties of Compression on Multilingual Models
Kelechi Ogueji, Orevaoghene Ahia, Gbemileke Onilude, Sebastian Gehrmann, Sara Hooker, and Julia Kreutzer. 2022 · 2022
Closest in time.
Combining Modular Skills in Multitask Learning
Edoardo M Ponti, Alessandro Sordoni, and Siva Reddy. 2022 · 2022
Closest in time.
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Ofir Press, Noah Smith, and Mike Lewis. 2022 · 2022
Closest in time.
DOTA: detect and omit weak attentions for scalable transformer acceleration
Zheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen, Yufei Ding, and Yuan Xie. 2022 · 2022
Closest in time.
DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He. 2022 · 2022
Closest in time.
Multitask Prompted Training Enables Zero-Shot Task Generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. 2022 · 2022
Closest in time.
Efficient Transformers: A Survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2022 · 2022
Closest in time.
Predicting Attention Sparsity in Transformers
Marcos Treviso, António Góis, Patrick Fernandes, Erick Fonseca, and Andre Martins. 2022 · 2022
Closest in time.
DyLoRA: Parameter Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low Rank Adaptation
Mojtaba Valipour, Mehdi Rezagholizadeh, Ivan Kobyzev, and Ali Ghodsi. 2022 · 2022
Closest in time.
AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning
Yaqing Wang, Sahaj Agarwal, Subhabrata Mukherjee, Xiaodong Liu, Jing Gao, Ahmed Hassan Awadallah, and Jianfeng Gao. 2022b · 2022
Closest in time.
Should You Mask 15% in Masked Language Modeling?
Alexander Wettig, Tianyu Gao, Zexuan Zhong, and Danqi Chen. 2022 · 2022
Closest in time.
Structured Pruning Learns Compact and Accurate Models
Mengzhou Xia, Zexuan Zhong, and Danqi Chen. 2022 · 2022
Closest in time.
Can Model Compression Improve NLP Fairness
Guangxuan Xu and Qingyuan Hu. 2022 · 2022
Closest in time.
Adapting Coreference Resolution Models through Active Learning
Michelle Yuan, Patrick Xia, Chandler May, Benjamin Van Durme, and Jordan Boyd-Graber. 2022 · 2022
Closest in time.
Mokey: enabling narrow fixed-point inference for out-of-the-box floating-point transformer models
Ali Hadi Zadeh, Mostafa Mahmoud, Ameer Abdelhadi, and Andreas Moshovos. 2022 · 2022
Closest in time.
OPT: Open Pre-trained Transformer Language Models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022 · 2022
Closest in time.
Teach Less, Learn More: On the Undistillable Classes in Knowledge Distillation
Yichen Zhu, Ning Liu, Zhiyuan Xu, Xin Liu, Weibin Meng, Louis Wang, Zhicai Ou, and Jian Tang. 2022 · 2022
Closest in time.
Designing Effective Sparse Expert Models
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. 2022 · 2022
Closest in time.
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2023 · 2023
Closest in time.
Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Closest in time.
Mega: Moving Average Equipped Gated Attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer. 2023 · 2023
Closest in time.
Long Range Language Modeling via Gated State Spaces
Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur. 2023 · 2023
Closest in time.
On the effect of dropping layers of pre-trained transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2023 · 2023
Closest in time.
A Survey on Dynamic Neural Networks for Natural Language Processing
Canwen Xu and Julian McAuley. 2023 · 2023
Closest in time.