Fetching the paper…
Reading the bibliography…
Multi-task learning (MTL) has been widely used in recommender systems, wherein predicting each type of user feedback on items (e.g, click, purchase) are treated as individual tasks and jointly trained with a unified model.
Preparing lessons: Improve knowledge distillation with better supervision
Wen, T.; Lai, S.; and Qian, X. 2019 · 1911
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Ma, J.; Zhao, Z.; Yi, X.; Chen, J.; Hong, L.; and Chi, E. H. 2018a · 1939
Earlier work this paper cites.
Multitask learning
Caruana, R. 1997 · 1997
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
Platt, J.; et al. 1999 · 1999
Earlier work this paper cites.
A simple generalisation of the area under the ROC curve for multiple class classification problems
Hand, D. J.; and Till, R. J. 2001 · 2001
Earlier work this paper cites.
Understanding and improving knowledge distillation
Tang, J.; Shivanna, R.; Zhao, Z.; Lin, D.; Singh, A.; Chi, E. H.; and Jain, S. 2020b · 2002
Earlier work this paper cites.
Predicting good probabilities with supervised learning
Niculescu-Mizil, A.; and Caruana, R. 2005 · 2005
Earlier work this paper cites.
Self-distillation as instance-specific label smoothing
Zhang, Z.; and Sabuncu, M. R. 2020 · 2006
Earlier work this paper cites.
An experimental comparison of performance measures for classification
Ferri, C.; Hernández-Orallo, J.; and Modroiu, R. 2009 · 2009
Earlier work this paper cites.
BPR: Bayesian personalized ranking from implicit feedback
Rendle, S.; Freudenthaler, C.; Gantner, Z.; and Schmidt-Thieme, L. 2012 · 2012
Earlier work this paper cites.
Low resource dependency parsing: Cross-lingual parameter sharing in a neural network parser
Duong, L.; Cohn, T.; Bird, S.; and Cook, P. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Cited alongside, same era.
Cross-stitch networks for multi-task learning
Misra, I.; Shrivastava, A.; Gupta, A.; and Hebert, M. 2016 · 2016
Cited alongside, same era.
Deep multi-task representation learning: A tensor factorisation approach
Yang, Y.; and Hospedales, T. 2016 · 2016
Cited alongside, same era.
An overview of multi-task learning in deep neural networks
Ruder, S. 2017 · 2017
Cited alongside, same era.
Optimizing ranking for response prediction via triplet-wise learning from historical feedback
Shan, L.; Lin, L.; Sun, C.; Wang, X.; and Liu, B. 2017 · 2017
Cited alongside, same era.
Snr: Sub-network routing for flexible parameter sharing in multi-task learning
Ma, J.; Zhao, Z.; Chen, J.; Li, A.; Hong, L.; and Chi, E. H. 2019 · 2019
Later among the works it cites.
Predicting different types of conversions with multi-task learning in online advertising
Pan, J.; Mao, Y.; Ruiz, A. L.; Sun, Y.; and Flores, A. 2019 · 2019
Later among the works it cites.
Towards understanding knowledge distillation
Phuong, M.; and Lampert, C. 2019 · 2019
Later among the works it cites.
Privileged features distillation at Taobao recommendations
Xu, C.; Li, Q.; Ge, J.; Gao, J.; Yang, X.; Pei, C.; Sun, F.; Wu, J.; Sun, H.; and Ou, W. 2020 · 2020
Later among the works it cites.
Gradient surgery for multi-task learning
Yu, T.; Kumar, S.; Gupta, A.; Levine, S.; Hausman, K.; and Finn, C. 2020 · 2020
Later among the works it cites.
Revisiting knowledge distillation via label smoothing regularization
Yuan, L.; Tay, F. E.; Li, G.; Wang, T.; and Feng, J. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Large scale distributed neural network training through online distillation
Anil, R.; Pereyra, G.; Passos, A.; Ormandi, R.; Dahl, G. E.; and Hinton, G. E. 2018 · 2018
Cited alongside, same era.
Why I like it: multi-task learning for recommendation and explanation
Lu, Y.; Dong, R.; and Smyth, B. 2018 · 2018
Cited alongside, same era.
Combined regression and tripletwise learning for conversion rate prediction in real-time bidding advertising
Shan, L.; Lin, L.; and Sun, C. 2018 · 2018
Cited alongside, same era.
Ranking distillation: Learning compact ranking models with high performance for recommender system
Tang, J.; and Wang, K. 2018 · 2018
Cited alongside, same era.
Explainable recommendation via multi-task learning in opinionated text data
Wang, N.; Wang, H.; Jia, Y.; and Yin, Y. 2018 · 2018
Cited alongside, same era.
Rocket launching: A universal and efficient framework for training well-performing light net
Zhou, G.; Fan, Y.; Cui, R.; Bian, W.; Zhu, X.; and Gai, K. 2018 · 2018
Cited alongside, same era.
Entire space multi-task model: An effective approach for estimating post-click conversion rate
Ma, X.; Zhao, L.; Huang, G.; Wang, Z.; Hu, Z.; Zhu, X.; and Gai, K. 2018b
Cited in the paper.
Later among the works it cites.
Ensembled CTR prediction via knowledge distillation
Zhu, J.; Liu, J.; Li, W.; Lai, J.; He, X.; Chen, L.; and Zheng, Z. 2020 · 2020
Later among the works it cites.
Boosting Multi-task Learning Through Combination of Task Labels-with Applications in ECG Phenotyping
Hsieh, M.-E.; and Tseng, V. 2021 · 2021
Later among the works it cites.
Self-knowledge distillation with progressive refinement of targets
Kim, K.; Ji, B.; Yoon, D.; and Hwang, S. 2021 · 2021
Later among the works it cites.
A survey on multi-task learning
Zhang, Y.; and Yang, Q. 2021 · 2021
Later among the works it cites.
Rethinking soft labels for knowledge distillation: A bias-variance tradeoff perspective
Zhou, H.; Song, L.; Chen, J.; Zhou, Y.; Wang, G.; Yuan, J.; and Zhang, Q. 2021 · 2021
Later among the works it cites.