Fetching the paper…
Reading the bibliography…
Sparsely activated neural networks with conditional computation learn to route their inputs through different "expert" subnetworks, providing a form of modularity that densely activated models lack.
Fastbert: a self-distilling bert with adaptive inference time
Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Haotang Deng, and Qi Ju · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Bill Dolan and Chris Brockett · 2005
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo · 2009
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D. Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Earlier work this paper cites.
Gradient estimation using stochastic computation graphs
John Schulman, Nicolas Heess, Theophane Weber, and Pieter Abbeel · 2015
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia · 2017
Earlier work this paper cites.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Will Grathwohl, Dami Choi, Yuhuai Wu, Geoffrey Roeder, and David Duvenaud · 2017
Earlier work this paper cites.
Learning to reason: End-to-end module networks for visual question answering
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi · 2017
Earlier work this paper cites.
First quora dataset release: question pairs (2017)
Iyer Shankar, Dandekar Nikhil, and Csernai Kornel · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models
George Tucker, Andriy Mnih, Chris J Maddison, John Lawson, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Modular networks: Learning to decompose neural computation
Louis Kirsch, Julius Kunze, and David Barber · 2018
Earlier work this paper cites.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Arm: Augment-reinforce-merge gradient for stochastic binary networks
Mingzhang Yin and Mingyuan Zhou · 2018
Earlier work this paper cites.
Taskonomy: Disentangling task transfer learning
Amir R Zamir, Alexander Sax, William Shen, Leonidas J Guibas, Jitendra Malik, and Silvio Savarese · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
Self-assembling modular networks for interpretable multi-hop reasoning
Yichen Jiang and Mohit Bansal · 2019
Earlier work this paper cites.
Buy 4 reinforce samples, get a baseline for free!
Wouter Kool, Herke van Hoof, and Max Welling · 2019
Earlier work this paper cites.
Snr: Sub-network routing for flexible parameter sharing in multi-task learning
Jiaqi Ma, Zhe Zhao, Jilin Chen, Ang Li, Lichan Hong, and Ed H Chi · 2019
Cited alongside, same era.
Flexible multi-task networks by learning parameter allocation
Krzysztof Maziarz, Efi Kokiopoulou, Andrea Gesmundo, Luciano Sbaiz, Gabor Bartok, and Jesse Berent · 2019
Cited alongside, same era.
Moment matching for multi-source domain adaptation
Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang · 2019
Cited alongside, same era.
How multilingual is multilingual bert?
Telmo Pires, Eva Schlinger, and Dan Garrette · 2019
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2019
Cited alongside, same era.
Hash layers for large sparse models
Stephen Roller, Sainbayar Sukhbaatar, Jason Weston, et al · 2021
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al · 2021
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Later among the works it cites.
Basisnet: Two-stage model synthesis for efficient inference
Mingda Zhang, Chun-Te Chu, Andrey Zhmoginov, Andrew Howard, Brendan Jou, Yukun Zhu, Li Zhang, Rebecca Hwa, and Adriana Kovashka · 2021
Later among the works it cites.
Taming sparsely activated transformer with stochastic experts
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Condconv: Conditionally parameterized convolutions for efficient inference
Brandon Yang, Gabriel Bender, Quoc V Le, and Jiquan Ngiam · 2019
Cited alongside, same era.
Understanding the role of individual units in a deep neural network
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba · 2020
Cited alongside, same era.
Latent domain learning with dynamic residual adapters
Lucas Deecke, Timothy Hospedales, and Hakan Bilen · 2020
Cited alongside, same era.
Disarm: An antithetic gradient estimator for binary latent variables
Zhe Dong, Andriy Mnih, and George Tucker · 2020
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2020
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Cited alongside, same era.
Mad-x: An adapter-based framework for multi-task cross-lingual transfer
Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder · 2020
Cited alongside, same era.
Simiao Zuo, Xiaodong Liu, Jian Jiao, Young Jin Kim, Hany Hassan, Ruofei Zhang, Tuo Zhao, and Jianfeng Gao · 2021
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Samuel K Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa · 2022
Later among the works it cites.
Promptsource: An integrated development environment and repository for natural language prompts
Stephen H Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, et al · 2022
Later among the works it cites.
On the representation collapse of sparse mixture of experts
Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, and Furu Wei · 2022
Later among the works it cites.
Unified scaling laws for routed language models
Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, et al · 2022
Later among the works it cites.
Stablemoe: Stable routing strategy for mixture of experts
Damai Dai, Li Dong, Shuming Ma, Bo Zheng, Zhifang Sui, Baobao Chang, and Furu Wei · 2022
Later among the works it cites.
Cold fusion: Collaborative descent for distributed multitask finetuning
Shachar Don-Yehiya, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen · 2022
Later among the works it cites.
Glam: Efficient scaling of language models with mixture-of-experts
Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al · 2022
Later among the works it cites.
A review of sparse expert models in deep learning
William Fedus, Jeff Dean, and Barret Zoph · 2022
Later among the works it cites.
Sparsely activated mixture-of-experts are robust multi-task learners
Shashank Gupta, Subhabrata Mukherjee, Krishan Subudhi, Eduardo Gonzalez, Damien Jose, Ahmed Hassan Awadallah, and Jianfeng Gao · 2022
Later among the works it cites.
Patching open-vocabulary models by interpolating weights
Gabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song, Hannaneh Hajishirzi, Simon Kornblith, Ali Farhadi, and Ludwig Schmidt · 2022
Later among the works it cites.
Dataless knowledge fusion by merging weights of language models
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng · 2022
Later among the works it cites.
Linear connectivity reveals generalization strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra · 2022
Later among the works it cites.
Sparse upcycling: Training mixture-of-experts from dense checkpoints
Aran Komatsuzaki, Joan Puigcerver, James Lee-Thorp, Carlos Riquelme Ruiz, Basil Mustafa, Joshua Ainslie, Yi Tay, Mostafa Dehghani, and Neil Houlsby · 2022
Later among the works it cites.
Sparse mixers: Combining moe and mixing to build a more efficient bert
James Lee-Thorp and Joshua Ainslie · 2022
Later among the works it cites.
Is a modular architecture enough?
Sarthak Mittal, Yoshua Bengio, and Guillaume Lajoie · 2022
Later among the works it cites.
Lifting the curse of multilinguality by pre-training modular transformers
Jonas Pfeiffer, Naman Goyal, Xi Victoria Lin, Xian Li, James Cross, Sebastian Riedel, and Mikel Artetxe · 2022
Later among the works it cites.
Combining modular skills in multitask learning
Edoardo M Ponti, Alessandro Sordoni, and Siva Reddy · 2022
Later among the works it cites.
Exploring mode connectivity for pre-trained language models
Yujia Qin, Cheng Qian, Jing Yi, Weize Chen, Yankai Lin, Xu Han, Zhiyuan Liu, Maosong Sun, and Jie Zhou · 2022
Later among the works it cites.
One model, multiple tasks: Pathways for natural language understanding
Duyu Tang, Fan Zhang, Yong Dai, Cong Zhou, Shuangzhi Wu, and Shuming Shi · 2022
Later among the works it cites.
Eliciting transferability in multi-task learning with task-level mixture-of-experts
Qinyuan Ye, Juan Zha, and Xiang Ren · 2022
Later among the works it cites.
Efficient language modeling with sparse all-mlp
Ping Yu, Mikel Artetxe, Myle Ott, Sam Shleifer, Hongyu Gong, Ves Stoyanov, and Xian Li · 2022
Later among the works it cites.
Moefication: Transformer feed-forward layers are mixtures of experts
Zhengyan Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew Dai, Zhifeng Chen, Quoc Le, and James Laudon · 2022
Later among the works it cites.
Designing effective sparse expert models
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus · 2022
Later among the works it cites.
Jonas Pfeiffer, Sebastian Ruder, Ivan Vulić, and Edoardo Maria Ponti · 2023
Closest in time.
From sparse to soft mixtures of experts
Joan Puigcerver, Carlos Riquelme, Basil Mustafa, and Neil Houlsby · 2023
Closest in time.
p i pi -tuning: Transferring multimodal foundation models with optimal multi-task interpolation
Chengyue Wu, Teng Wang, Yixiao Ge, Zeyu Lu, Ruisong Zhou, Ying Shan, and Ping Luo · 2023
Closest in time.
Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning
Ted Zadouri, Ahmet Üstün, Arash Ahmadian, Beyza Ermiş, Acyr Locatelli, and Sara Hooker · 2023
Closest in time.