Fetching the paper…
Reading the bibliography…
Large language models (LLMs) encounter significant adaptation challenges in diverse multitask finetuning.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, et al · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Michael I. Jordan and Robert A. Jacobs · 1994
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gut feelings: The intelligence of the unconscious
Gerd Gigerenzer · 2007
Earlier work this paper cites.
How intuition contributes to high performance: An educational perspective
Christian Harteis, Tina Koch, and Barbara Morgenthaler · 2008
Earlier work this paper cites.
Mixture of experts: a literature survey
Saeed Masoudnia and Reza Ebrahimpour · 2014
Earlier work this paper cites.
Cross-stitch networks for multi-task learning
Ishan Misra, Abhinav Shrivastava, Abhinav Gupta, and Martial Hebert · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, et al · 2017
Earlier work this paper cites.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Latent multi-task architecture learning
Sebastian Ruder, Joachim Bingel, Isabelle Augenstein, and Anders Søgaard · 2019
Earlier work this paper cites.
Socialiqa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, et al · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
SuperGLUE: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, et al · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, et al · 2020
Cited alongside, same era.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie and Adina others Williams · 2020
Cited alongside, same era.
In Communications , volume 64, pages 99–106, 2021
Winogrande: An adversarial winograd schema challenge at scale · 2021
Cited alongside, same era.
Repvgg: Making vgg-style convnets great again
Xiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han, Guiguang Ding, and Jian Sun · 2021
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Designing effective sparse expert models
Barret Zoph · 2022
Later among the works it cites.
Opencompass: A universal evaluation platform for foundation models
OpenCompass Contributors · 2023
Later among the works it cites.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ning Ding, Yujia Qin, et al · 2023
Later among the works it cites.
On the effectiveness of parameter-efficient fine-tuning
Zihao Fu, Haoran Yang, et al · 2023
Later among the works it cites.
Lorahub: Efficient cross-task generalization via dynamic lora composition
Chengsong Huang, Qian Liu, et al · 2023
Later among the works it cites.
Phi-2: The surprising power of small language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, et al · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
J. Edward Hu, Yelong Shen, et al · 2021
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, et al · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Cited alongside, same era.
Hash layers for large sparse models
Stephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, and Jason Weston · 2021
Cited alongside, same era.
Javaheripi, Mojan, Bubeck, et al · 2023
Later among the works it cites.
Albert Qiaochu Jiang, Alexandre, et al · 2023
Later among the works it cites.
Mowe: Mixture of weather experts for multiple adverse weather removal
Yulin Luo, Rui Zhao, et al · 2023
Later among the works it cites.
From sparse to soft mixtures of experts
Joan Puigcerver, Carlos Riquelme, Basil Mustafa, and Neil Houlsby · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, et al · 2023
Later among the works it cites.
Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning
Ted Zadouri, Ahmet Üstün, et al · 2023
Later among the works it cites.
Sira: Sparse mixture of low rank adaptation
Yun Zhu, Nevan Wichers, et al · 2023
Later among the works it cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Closest in time.
Nomic embed: Training a reproducible long context text embedder
Zach Nussbaum, John X Morris, Brandon Duderstadt, and Andriy Mulyar · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, et al · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Alex Young, Bei Chen, et al · 2024
Closest in time.