Fetching the paper…
Reading the bibliography…
Multi-Task Learning (MTL) is a powerful learning paradigm to improve generalization performance via knowledge sharing.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D · 1907
Earlier work this paper cites.
Masked gradient-based causal structure learning
Ng, I., Fang, Z., Zhu, S., and Chen, Z · 1910
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
A model of inductive bias learning
Baxter, J · 2000
Earlier work this paper cites.
Invariant risk minimization games
Ahuja, K., Shanmugam, K., Varshney, K. R., and Dhurandhar, A · 2002
Earlier work this paper cites.
Particular formulae for the moore–penrose inverse of a columnwise partitioned matrix
Baksalary, J. K. and Baksalary, O. M · 2006
Earlier work this paper cites.
A notion of task relatedness yielding provable multiple-task learning guarantees
Ben-David, S. and Borbely, R. S · 2008
Earlier work this paper cites.
When is invariance useful in an out-of-distribution generalization problem?
Koyama, M. and Yamaguchi, S · 2008
Earlier work this paper cites.
The effectiveness of memory replay in large scale continual learning
Balaji, Y., Farajtabar, M., Yin, D., Mott, A., and Li, A · 2010
Earlier work this paper cites.
Learning causal semantic representation for out-of-distribution prediction
Liu, C., Sun, X., Wang, J., Li, T., Qin, T., Chen, W., and Liu, T · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Torralba, A. and Efros, A. A · 2011
Earlier work this paper cites.
Entropy balancing for causal effects: A multivariate reweighting method to produce balanced samples in observational studies
Hainmueller, J · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from RGBD images
Silberman, N., Hoiem, D., Kohli, P., and Fergus, R · 2012
Earlier work this paper cites.
Learning fair representations
Zemel, R. S., Wu, Y., Swersky, K., Pitassi, T., and Dwork, C · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2014
Earlier work this paper cites.
Segnet: A deep convolutional encoder-decoder architecture for image segmentation
Badrinarayanan, V., Kendall, A., and Cipolla, R · 2015
Earlier work this paper cites.
Discovering hidden factors of variation in deep networks
Cheung, B., Livezey, J. A., Bansal, A. K., and Olshausen, B. A · 2015
Earlier work this paper cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Reducing overfitting in deep networks by decorrelating representations
Cogswell, M., Ahmed, F., Girshick, R. B., Zitnick, L., and Batra, D · 2016
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., Marchand, M., and Lempitsky, V. S · 2016
Earlier work this paper cites.
The movielens datasets: History and context
Harper, F. M. and Konstan, J. A · 2016
Earlier work this paper cites.
From dependence to causation
Lopez-Paz, D · 2016
Earlier work this paper cites.
The benefit of multitask representation learning
Maurer, A., Pontil, M., and Romera-Paredes, B · 2016
Earlier work this paper cites.
Cross-stitch networks for multi-task learning
Misra, I., Shrivastava, A., Gupta, A., and Hebert, M · 2016
Earlier work this paper cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
Parisotto, E., Ba, L. J., and Salakhutdinov, R · 2016
Cited alongside, same era.
Causal inference by using invariant prediction: identification and confidence intervals
Peters, J., Bühlmann, P., and Meinshausen, N · 2016
Cited alongside, same era.
Annotation artifacts in natural language inference data
Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R., and Smith, N. A · 2017
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2017
Cited alongside, same era.
Fully-adaptive feature sharing in multi-task networks with applications in person attribute classification
Lu, Y., Kumar, A., Zhai, S., Cheng, Y., Javidi, T., and Feris, R. S · 2017
Cited alongside, same era.
Robustly disentangled causal mechanisms: Validating deep representations for interventional robustness
Suter, R., Miladinovic, D., Schölkopf, B., and Bauer, S · 2019
Later among the works it cites.
Balanced datasets are not enough: Estimating and mitigating gender bias in deep image representations
Wang, T., Zhao, J., Yatskar, M., Chang, K., and Ordonez, V · 2019
Later among the works it cites.
A meta-transfer objective for learning to disentangle causal mechanisms
Bengio, Y., Deleu, T., Rahaman, N., Ke, N. R., Lachapelle, S., Bilaniuk, O., Goyal, A., and Pal, C. J · 2020
Later among the works it cites.
Just pick a sign: Optimizing deep multitask models with gradient sign dropout
Chen, Z., Ngiam, J., Huang, Y., Luong, T., Kretzschmar, H., Chai, Y., and Anguelov, D · 2020
Later among the works it cites.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J., Michaelis, C., Zemel, R. S., Brendel, W., Bethge, M., and Wichmann, F. A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Cited alongside, same era.
Recognition in terra incognita
Beery, S., Horn, G. V., and Perona, P · 2018
Cited alongside, same era.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Chen, Z., Badrinarayanan, V., Lee, C., and Rabinovich, A · 2018
Cited alongside, same era.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Kendall, A., Gal, Y., and Cipolla, R · 2018
Cited alongside, same era.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Ma, J., Zhao, Z., Yi, X., Chen, J., Hong, L., and Chi, E. H · 2018
Cited alongside, same era.
Learning independent causal mechanisms
Parascandolo, G., Kilbertus, N., Rojas-Carulla, M., and Schölkopf, B · 2018
Cited alongside, same era.
Routing networks: Adaptive selection of non-linear functions for multi-task learning
Rosenbaum, C., Klinger, T., and Riemer, M · 2018
Cited alongside, same era.
Learning to branch for multi-task learning
Guo, P., Lee, C., and Ulbricht, D · 2020
Later among the works it cites.
Gradient-based neural DAG learning
Lachapelle, S., Brouillard, P., Deleu, T., and Lacoste-Julien, S · 2020
Later among the works it cites.
Distributionally robust neural networks
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P · 2020
Later among the works it cites.
Which tasks should be learned together in multi-task learning?
Standley, T., Zamir, A. R., Chen, D., Guibas, L. J., Malik, J., and Savarese, S · 2020
Later among the works it cites.
On the theory of transfer learning: The importance of task diversity
Tripuraneni, N., Jordan, M. I., and Jin, C · 2020
Later among the works it cites.
Understanding and improving information transfer in multi-task learning
Wu, S., Zhang, H. R., and Ré, C · 2020
Later among the works it cites.
Gradient surgery for multi-task learning
Yu, T., Kumar, S., Gupta, A., Levine, S., Hausman, K., and Finn, C · 2020
Later among the works it cites.
Systematic generalisation with group invariant predictions
Ahmed, F., Bengio, Y., van Seijen, H., and Courville, A. C · 2021
Later among the works it cites.
Few-shot learning via learning the representation, provably
Du, S. S., Hu, W., Kakade, S. M., Lee, J. D., and Lei, Q · 2021
Later among the works it cites.
Efficiently identifying task groupings for multi-task learning
Fifty, C., Amid, E., Zhao, Z., Yu, T., Anil, R., and Finn, C · 2021
Later among the works it cites.
Removing spurious features can hurt accuracy and affect groups disproportionately
Khani, F. and Liang, P · 2021
Later among the works it cites.
Out-of-distribution generalization via risk extrapolation (rex)
Krueger, D., Caballero, E., Jacobsen, J., Zhang, A., Binas, J., Zhang, D., Priol, R. L., and Courville, A. C · 2021
Later among the works it cites.
Disentanglement via mechanism sparsity regularization: A new principle for nonlinear ica
Lachapelle, S., López, P. R., Sharma, Y., Everett, K., Priol, R. L., Lacoste, A., and Lacoste-Julien, S · 2021
Later among the works it cites.
Representation learning via invariant causal mechanisms
Mitrovic, J., McWilliams, B., Walker, J. C., Buesing, L. H., and Blundell, C · 2021
Later among the works it cites.
Understanding the failure modes of out-of-distribution generalization
Nagarajan, V., Andreassen, A., and Neyshabur, B · 2021
Later among the works it cites.
The risks of invariant risk minimization
Rosenfeld, E., Ravikumar, P. K., and Risteski, A · 2021
Later among the works it cites.
Toward causal representation learning
Schölkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y · 2021
Later among the works it cites.
Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models
Wang, Z., Tsvetkov, Y., Firat, O., and Cao, Y · 2021
Later among the works it cites.
A survey on negative transfer
Zhang, W., Deng, L., Zhang, L., and Wu, D · 2021
Later among the works it cites.