Fetching the paper…
Reading the bibliography…
Fine-tuning pre-trained language models has become the prevalent paradigm for building downstream NLP models.
Evidence for universality and cultural variation of differential emotion response patterning
Klaus R Scherer and Harald G Wallbott · 1994
Earlier work this paper cites.
Popular ensemble methods: An empirical study
David Opitz and Richard Maclin · 1999
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik Tjong Kim Sang and Fien De Meulder · 2003
Earlier work this paper cites.
Emotions from text: machine learning for text-based emotion prediction
Cecilia Ovesdotter Alm, Dan Roth, and Richard Sproat · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Bill Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Ontonotes: the 90% solution
Eduard Hovy, Mitch Marcus, Martha Palmer, Lance Ramshaw, and Ralph Weischedel · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and William B Dolan · 2007
Earlier work this paper cites.
Semeval-2007 task 14: Affective text
Carlo Strapparava and Rada Mihalcea · 2007
Earlier work this paper cites.
Ensemble-based classifiers
Lior Rokach · 2010
Earlier work this paper cites.
# emotional tweets
Saif Mohammad · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Sentiment, emotion, purpose, and style in electoral tweets
Saif M Mohammad, Xiaodan Zhu, Svetlana Kiritchenko, and Joel Martin · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia · 2017
Earlier work this paper cites.
Dailydialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu · 2017
Earlier work this paper cites.
Grounded emotions
Vicki Liu, Carmen Banea, and Rada Mihalcea · 2017
Earlier work this paper cites.
Data-free knowledge distillation for deep neural networks
Raphael Gontijo Lopes, Stefano Fenu, and Thad Starner · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
Wassa-2017 shared task on emotion intensity
Saif Mohammad and Felipe Bravo-Marquez · 2017
Earlier work this paper cites.
Annotation, modelling and analysis of fine-grained emotions on a stance and sentiment detection corpus
Hendrik Schuff, Jeremy Barnes, Julian Mohme, Sebastian Padó, and Roman Klinger · 2017
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
An analysis of annotated corpora for emotion classification in text
Laura Ana Maria Oberländer and Roman Klinger · 2018
Cited alongside, same era.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman · 2018
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Later among the works it cites.
Temporally-informed analysis of named entity recognition
Shruti Rijhwani and Daniel Preotiuc-Pietro · 2020
Later among the works it cites.
Model fusion via optimal transport
Sidak Pal Singh and Martin Jaggi · 2020
Later among the works it cites.
Ensemble of averages: Improving model selection and boosting performance in domain generalization
Devansh Arpit, Huan Wang, Yingbo Zhou, and Caiming Xiong · 2021
Later among the works it cites.
Swad: Domain generalization by seeking flat minima
Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
On the convergence of fedavg on non-iid data
Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Zero-shot knowledge distillation in deep networks
Gaurav Kumar Nayak, Konda Reddy Mopuri, Vaisakh Shaj, Venkatesh Babu Radhakrishnan, and Anirban Chakraborty · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing, 2021
Pengcheng He, Jianfeng Gao, and Weizhu Chen · 2021
Later among the works it cites.
Mergedistill: Merging language models using pre-trained distillation
Simran Khanuja, Melvin Johnson, and Partha Talukdar · 2021
Later among the works it cites.
Merging models with fisher-weighted averaging
Michael Matena and Colin Raffel · 2021
Later among the works it cites.
Model fusion of heterogeneous neural networks via cross-layer alignment
Dang Nguyen, Khai Nguyen, Dinh Phung, Hung Bui, and Nhat Ho · 2021
Later among the works it cites.
What to pre-train on? Efficient intermediate task selection
Clifton Poth, Jonas Pfeiffer, Andreas Rücklé, and Iryna Gurevych · 2021
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Samuel K Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa · 2022
Closest in time.
Fusing finetuned models for better pretraining
Leshem Choshen, Elad Venezian, Noam Slonim, and Yoav Katz · 2022
Closest in time.
Branch-train-merge: Embarrassingly parallel training of expert language models
Margaret Li, Suchin Gururangan, Tim Dettmers, Mike Lewis, Tim Althoff, Noah A Smith, and Luke Zettlemoyer · 2022
Closest in time.
FedNLP: Benchmarking federated learning methods for natural language processing tasks
Bill Yuchen Lin, Chaoyang He, Zihang Ze, Hulin Wang, Yufen Hua, Christophe Dupuy, Rahul Gupta, Mahdi Soltanolkotabi, Xiang Ren, and Salman Avestimehr · 2022
Closest in time.
Diverse weight averaging for out-of-distribution generalization
Alexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy, Patrick Gallinari, and Matthieu Cord · 2022
Closest in time.
Adamix: Mixture-of-adapter for parameter-efficient tuning of large language models
Yaqing Wang, Subhabrata Mukherjee, Xiaodong Liu, Jing Gao, Ahmed Hassan Awadallah, and Jianfeng Gao · 2022
Closest in time.
When to use multi-task learning vs intermediate fine-tuning for pre-trained encoder transfer learning
Orion Weller, Kevin Seppi, and Matt Gardner · 2022
Closest in time.
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al · 2022
Closest in time.