Fetching the paper…
Reading the bibliography…
Fine-tuning pre-trained language models, particularly large language models, demands extensive computing resources and can result in varying performance outcomes across different domains and datasets.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Training feedforward neural networks using genetic algorithms
David J Montana, Lawrence Davis, et al. 1989 · 1989
Earlier work this paper cites.
Evidence for universality and cultural variation of differential emotion response patterning
Klaus R Scherer and Harald G Wallbott. 1994 · 1994
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2002
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
Kenneth O Stanley and Risto Miikkulainen. 2002 · 2002
Earlier work this paper cites.
Federated learning with matched averaging
Hongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris Papailiopoulos, and Yasaman Khazaeni. 2020a · 2002
Earlier work this paper cites.
Emotions from text: machine learning for text-based emotion prediction
Cecilia Ovesdotter Alm, Dan Roth, and Richard Sproat. 2005 · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Bill Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Self-adaptive differential evolution algorithm for numerical optimization
A Kai Qin and Ponnuthurai N Suganthan. 2005 · 2005
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and William B Dolan. 2007 · 2007
Earlier work this paper cites.
Semeval-2007 task 14: Affective text
Carlo Strapparava and Rada Mihalcea. 2007 · 2007
Earlier work this paper cites.
Differential evolution for neural architecture search
Noor Awad, Neeratyoy Mallik, and Frank Hutter. 2020 · 2012
Earlier work this paper cites.
# emotional tweets
Saif Mohammad. 2012 · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Success-history based parameter adaptation for differential evolution
Ryoji Tanabe and Alex Fukunaga. 2013 · 2013
Earlier work this paper cites.
Differential evolution algorithms applied to neural network training suffer from stagnation
Adam P Piotrowski. 2014 · 2014
Earlier work this paper cites.
Sentiment, emotion, purpose, and style in electoral tweets
Saif M Mohammad, Xiaodan Zhu, Svetlana Kiritchenko, and Joel Martin. 2015 · 2015
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
Dailydialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Cited alongside, same era.
Grounded emotions
Vicki Liu, Carmen Banea, and Rada Mihalcea. 2017 · 2017
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. 2020 · 2020
Later among the works it cites.
Differential evolution: A review of more than two decades of research
Millie Pant, Hira Zaheer, Laura Garcia-Hernandez, Ajith Abraham, et al. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Later among the works it cites.
Transformer-based language model fine-tuning methods for covid-19 fake news detection
Ben Chen, Bin Chen, Dehong Gao, Qijin Chen, Chengfu Huo, Xiaonan Meng, Weijun Ren, and Yang Zhou. 2021 · 2021
Later among the works it cites.
Mergedistill: Merging language models using pre-trained distillation
Simran Khanuja, Melvin Johnson, and Partha Talukdar. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. 2017 · 2017
Cited alongside, same era.
Wassa-2017 shared task on emotion intensity
Saif Mohammad and Felipe Bravo-Marquez. 2017 · 2017
Cited alongside, same era.
An overview of multi-task learning in deep neural networks
Sebastian Ruder. 2017 · 2017
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. 2017 · 2017
Cited alongside, same era.
Annotation, modelling and analysis of fine-grained emotions on a stance and sentiment detection corpus
Hendrik Schuff, Jeremy Barnes, Julian Mohme, Sebastian Padó, and Roman Klinger. 2017 · 2017
Cited alongside, same era.
A genetic programming approach to designing convolutional neural network architectures
Masanori Suganuma, Shinichi Shirakawa, and Tomoharu Nagao. 2017 · 2017
Cited alongside, same era.
Deep learning with darwin: Evolutionary synthesis of deep neural networks
Mohammad Javad Shafiee, Akshaya Mishra, and Alexander Wong. 2018 · 2018
Cited alongside, same era.
What to pre-train on? efficient intermediate task selection
Clifton Poth, Jonas Pfeiffer, Andreas Rücklé, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2022 · 2022
Later among the works it cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2022 · 2022
Later among the works it cites.
Merging models with fisher-weighted averaging
Michael S Matena and Colin A Raffel. 2022 · 2022
Later among the works it cites.
Dataless knowledge fusion by merging weights of language models
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng. 2023 · 2023
Later among the works it cites.
Ties-merging: Resolving interference when merging models
Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal. 2023 · 2023
Later among the works it cites.
Language models are super mario: Absorbing abilities from homologous models as a free lunch
Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li. 2023 · 2023
Later among the works it cites.
Minicpm: Unveiling the potential of small language models with scalable training strategies
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, et al. 2024 · 2024
Closest in time.
Cade: Cosine annealing differential evolution for spiking neural network
Runhua Jiang, Guodong Du, Shuyang Yu, Yifei Guo, Sim Kuan Goh, and Ho-Kin Tang. 2024 · 2024
Closest in time.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Guillermo Ortiz-Jimenez, Alessandro Favero, and Pascal Frossard. 2024 · 2024
Closest in time.
Ties-merging: Resolving interference when merging models
Prateek Yadav, Derek Tam, Leshem Choshen, Colin A Raffel, and Mohit Bansal. 2024 · 2024
Closest in time.