Fetching the paper…
Reading the bibliography…
Multi-task learning (MTL) encapsulates multiple learned tasks in a single model and often lets those tasks learn better jointly.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Michael I Jordan and Robert A Jacobs · 1994
Earlier work this paper cites.
Improved learning algorithms for mixture of experts in multiclass classification
Ke Chen, Lei Xu, and Huisheng Chi · 1999
Earlier work this paper cites.
Combining predictors: comparison of five meta machine learning methods
Jakob Vogdrup Hansen · 1999
Earlier work this paper cites.
Learning to detect natural image boundaries using local brightness, color, and texture cues
David R Martin, Charless C Fowlkes, and Jitendra Malik · 2004
Earlier work this paper cites.
Learning gaussian processes from multiple tasks
Kai Yu, Volker Tresp, and Anton Schwaighofer · 2005
Earlier work this paper cites.
Multi-task learning for classification with dirichlet process priors
Ya Xue, Xuejun Liao, Lawrence Carin, and Balaji Krishnapuram · 2007
Earlier work this paper cites.
Learning a meta-level prior for feature relevance from multiple related tasks
Su-In Lee, Vassil Chatalbashev, David Vickrey, and Daphne Koller · 2007
Earlier work this paper cites.
Clustered multi-task learning: A convex formulation
Laurent Jacob, Jean-philippe Vert, and Francis Bach · 2008
Earlier work this paper cites.
Bayesian multitask learning with latent hierarchies
Hal Daumé III · 2009
Earlier work this paper cites.
Clustered multi-task learning via alternating structure optimization
Jiayu Zhou, Jianhui Chen, and Jieping Ye · 2011
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Learning task grouping and overlap in multi-task learning
Abhishek Kumar and Hal Daume III · 2012
Earlier work this paper cites.
Twenty years of mixture of experts
Seniha Esen Yuksel, Joseph N. Wilson, and Paul D. Gader · 2012
Earlier work this paper cites.
Multi-task bayesian optimization
Kevin Swersky, Jasper Snoek, and Ryan P Adams · 2013
Earlier work this paper cites.
Deep learning of representations: Looking forward
Yoshua Bengio · 2013
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2013
Earlier work this paper cites.
The role of context for object detection and semantic segmentation in the wild
Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille · 2014
Earlier work this paper cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
Conditional convolutional neural network for modality-aware face recognition
Chao Xiong, Xiaowei Zhao, Danhang Tang, Karlekar Jayashree, Shuicheng Yan, and Tae-Kyun Kim · 2015
Earlier work this paper cites.
A joint many-task model: Growing a neural network for multiple nlp tasks
Kazuma Hashimoto, Caiming Xiong, Yoshimasa Tsuruoka, and Richard Socher · 2016
Earlier work this paper cites.
Cross-stitch networks for multi-task learning
Ishan Misra, Abhinav Shrivastava, Abhinav Gupta, and Martial Hebert · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Network of experts for large-scale image categorization
Karim Ahmed, Mohammad Haris Baig, and Lorenzo Torresani · 2016
Earlier work this paper cites.
Neural relation extraction with selective attention over instances
Yankai Lin, Shiqi Shen, Zhiyuan Liu, Huanbo Luan, and Maosong Sun · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Ankur P Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
An overview of multi-task learning in deep neural networks
Sebastian Ruder · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Lukasz Kaiser, Aidan N Gomez, Noam Shazeer, Ashish Vaswani, Niki Parmar, Llion Jones, and Jakob Uszkoreit · 2017
Earlier work this paper cites.
Hard mixtures of experts for large scale weakly supervised vision
Sam Gross, Marc’Aurelio Ranzato, and Arthur Szlam · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille · 2017
Earlier work this paper cites.
Indoor semantic segmentation for robot navigating on mobile
Wonsuk Kim and Junhee Seok · 2018
Earlier work this paper cites.
Taskonomy: Disentangling task transfer learning
Amir R Zamir, Alexander Sax, William Shen, Leonidas J Guibas, Jitendra Malik, and Silvio Savarese · 2018
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi · 2018
Earlier work this paper cites.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich · 2018
Cited alongside, same era.
Dynamic task prioritization for multitask learning
Michelle Guo, Albert Haque, De-An Huang, Serena Yeung, and Li Fei-Fei · 2018
Cited alongside, same era.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2018
Cited alongside, same era.
Pad-net: Multi-tasks guided prediction-and-distillation network for simultaneous depth estimation and scene parsing
Dan Xu, Wanli Ouyang, Xiaogang Wang, and Nicu Sebe · 2018
Cited alongside, same era.
Joint task-recursive learning for semantic segmentation and depth estimation
Zhenyu Zhang, Zhen Cui, Chunyan Xu, Zequn Jie, Xiang Li, and Jian Yang · 2018
Cited alongside, same era.
Segformer: Simple and efficient design for semantic segmentation with transformers
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Later among the works it cites.
Fast drivable areas estimation with multi-task learning for real-time autonomous driving assistant
Dong-Gyu Lee · 2021
Later among the works it cites.
Multi-task learning for dense prediction tasks: A survey
Simon Vandenhende, Stamatios Georgoulis, Wouter Van Gansbeke, Marc Proesmans, Dengxin Dai, and Luc Van Gool · 2021
Later among the works it cites.
Sparse is enough in scaling transformers
Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, and Jonni Kanerva · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ozan Sener and Vladlen Koltun · 2018
Cited alongside, same era.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Cited alongside, same era.
Nddr-cnn: Layerwise feature fusing in multi-task cnns by neural discriminative dimensionality reduction
Yuan Gao, Jiayi Ma, Mingbo Zhao, Wei Liu, and Alan L Yuille · 2019
Cited alongside, same era.
End-to-end multi-task learning with attention
Shikun Liu, Edward Johns, and Andrew J Davison · 2019
Cited alongside, same era.
Latent multi-task architecture learning
Sebastian Ruder, Joachim Bingel, Isabelle Augenstein, and Anders Søgaard · 2019
Cited alongside, same era.
Joint cs-mri reconstruction and segmentation with a unified deep network
Liyan Sun, Zhiwen Fan, Xinghao Ding, Yue Huang, and John Paisley · 2019
Cited alongside, same era.
Pattern-affinitive propagation across depth, surface normal and semantic segmentation
Zhenyu Zhang, Zhen Cui, Chunyan Xu, Yan Yan, Nicu Sebe, and Jian Yang · 2019
Cited alongside, same era.
Unit: Multimodal multitask learning with a unified transformer
Ronghang Hu and Amanpreet Singh · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al · 2021
Later among the works it cites.
Hash layers for large sparse models
Stephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, and Jason E Weston · 2021
Later among the works it cites.
Tricks for training sparse translation models
Dheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross, Mike Lewis, and Angela Fan · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Later among the works it cites.
Base layers: Simplifying training of large, sparse models
Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer · 2021
Later among the works it cites.
Heterogeneous multi-task learning with expert diversity
Raquel Aoki, Frederick Tung, and Gabriel L Oliveira · 2021
Later among the works it cites.
Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning
Hussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy, Yihua Chen, Rahul Mazumder, Lichan Hong, and Ed Chi · 2021
Later among the works it cites.
Transgan: Two pure transformers can make one strong gan, and that can scale up
Yifan Jiang, Shiyu Chang, and Zhangyang Wang · 2021
Later among the works it cites.
Vitgan: Training gans with vision transformers
Kwonjoon Lee, Huiwen Chang, Lu Jiang, Han Zhang, Zhuowen Tu, and Ce Liu · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
Chasing sparsity in vision transformers: An end-to-end exploration
Tianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan, Lei Zhang, and Zhangyang Wang · 2021
Later among the works it cites.
Do we really need explicit position encodings for vision transformers?
Xiangxiang Chu, Bo Zhang, Zhi Tian, Xiaolin Wei, and Huaxia Xia · 2021
Later among the works it cites.
Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang · 2021
Later among the works it cites.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al · 2021
Later among the works it cites.
End-to-end human pose and mesh reconstruction with transformers
Kevin Lin, Lijuan Wang, and Zicheng Liu · 2021
Later among the works it cites.
Point transformer
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun · 2021
Later among the works it cites.
Pct: Point cloud transformer
Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R Martin, and Shi-Min Hu · 2021
Later among the works it cites.
Delayed propagation transformer: A universal computation engine towards practical control in cyber-physical systems
Wenqing Zheng, Qiangqiang Guo, Hao Yang, Peihao Wang, and Zhangyang Wang · 2021
Later among the works it cites.
Accommodating transformer onto fpga: Coupling the balanced model compression and fpga-implementation optimization
Panjie Qi, Yuhong Song, Hongwu Peng, Shaoyi Huang, Qingfeng Zhuge, and Edwin Hsing-Mean Sha · 2021
Later among the works it cites.
Accelerating transformer-based deep learning models on fpgas using column balanced block pruning
Hongwu Peng, Shaoyi Huang, Tong Geng, Ang Li, Weiwen Jiang, Hang Liu, Shusen Wang, and Caiwen Ding · 2021
Later among the works it cites.
Scaling vision with sparse mixture of experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby · 2021
Later among the works it cites.
Task switching network for multi-task learning
Guolei Sun, Thomas Probst, Danda Pani Paudel, Nikola Popović, Menelaos Kanakis, Jagruti Patel, Dengxin Dai, and Luc Van Gool · 2021
Later among the works it cites.
Task adaptive parameter sharing for multi-task learning
Matthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran, Charless Fowlkes, Rahul Bhotika, and Stefano Soatto · 2022
Closest in time.
Unified scaling laws for routed language models
Aidan Clark, Diego de las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake Hechtman, Trevor Cai, Sebastian Borgeaud, et al · 2022
Closest in time.
Anti-oversmoothing in deep vision transformers via the fourier domain analysis: From theory to practice
Peihao Wang, Wenqing Zheng, Tianlong Chen, and Zhangyang Wang · 2022
Closest in time.
Vision transformer for nerf-based view synthesis from a single input image
Kai-En Lin, Lin Yen-Chen, Wei-Sheng Lai, Tsung-Yi Lin, Yi-Chang Shih, and Ravi Ramamoorthi · 2022
Closest in time.
Is attention all nerf needs?
Mukund Varma T, Peihao Wang, Xuxi Chen, Tianlong Chen, Subhashini Venugopalan, and Zhangyang Wang · 2022
Closest in time.
Dall-e for detection: Language-driven context image synthesis for object detection
Yunhao Ge, Jiashu Xu, Brian Nlong Zhao, Laurent Itti, and Vibhav Vineet · 2022
Closest in time.
Vaqf: Fully automatic software-hardware co-design framework for low-bit vision transformer
Mengshu Sun, Haoyu Ma, Guoliang Kang, Yifan Jiang, Tianlong Chen, Xiaolong Ma, Zhangyang Wang, and Yanzhi Wang · 2022
Closest in time.