Fetching the paper…
Reading the bibliography…
When training and evaluating machine learning models on a large number of tasks, it is important to not only look at average task accuracy -- which may be biased by easy or redundant tasks -- but also worst-case accuracy (i.e.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 1910
Earlier work this paper cites.
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang · 1911
Earlier work this paper cites.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
On information and sufficiency
Solomon Kullback and Richard A Leibler · 1951
Earlier work this paper cites.
Information theory and statistical mechanics
Edwin T Jaynes · 1957
Earlier work this paper cites.
Bimatrix equilibrium points and mathematical programming
Carlton E Lemke · 1965
Earlier work this paper cites.
Perplexity—a measure of the difficulty of speech recognition tasks
Fred Jelinek, Robert L Mercer, Lalit R Bahl, and James K Baker · 1977
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadij Semenoviˇc Nemirovskij and David Borisovich Yudin · 1983
Earlier work this paper cites.
Relative information—what for?
Guy Jumarie · 1990
Earlier work this paper cites.
Entropy optimization principles and their applications
Jagat Narain Kapur and Hiremaglur K Kesavan · 1992
Earlier work this paper cites.
Exponentiated gradient versus gradient descent for linear predictors
Jyrki Kivinen and Manfred K Warmuth · 1997
Earlier work this paper cites.
Robust optimization , volume 28
Aharon Ben-Tal, Laurent El Ghaoui, and Arkadi Nemirovski · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky et al · 2009
Earlier work this paper cites.
Kullback-leibler divergence constrained distributionally robust optimization
Zhaolin Hu and L Jeff Hong · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Distributionally robust stochastic optimization with wasserstein distance
Rui Gao and Anton J Kleywegt · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
An Overview of Multi-Task Learning in Deep Neural Networks
Sebastian Ruder · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Cited alongside, same era.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich · 2018
Cited alongside, same era.
Dynamic task prioritization for multitask learning
Michelle Guo, Albert Haque, De-An Huang, Serena Yeung, and Li Fei-Fei · 2018
Cited alongside, same era.
What kind of language is hard to language-model?
Sabrina J. Mielke, Ryan Cotterell, Kyle Gorman, Brian Roark, and Jason Eisner · 2019
Later among the works it cites.
Distributionally robust language modeling
Yonatan Oren, Shiori Sagawa, Tatsunori Hashimoto, and Percy Liang · 2019
Later among the works it cites.
Distributionally Robust Optimization: A Review
Hamed Rahimian and Sanjay Mehrotra · 2019
Later among the works it cites.
Neural transfer learning for natural language processing
Sebastian Ruder · 2019
Later among the works it cites.
A hierarchical multi-task approach for learning embeddings from semantic tasks
Victor Sanh, Thomas Wolf, and Sebastian Ruder · 2019
Later among the works it cites.
Pyramidal person re-identification via multi-loss dynamic training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2018
Cited alongside, same era.
Scheduled multi-task learning: From syntax to translation
Eliyahu Kiperwasser and Miguel Ballesteros · 2018
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo · 2018
Cited alongside, same era.
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2018
Cited alongside, same era.
The natural language decathlon: Multitask learning as question answering
Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Cited alongside, same era.
Routing networks: Adaptive selection of non-linear functions for multi-task learning
Clemens Rosenbaum, Tim Klinger, and Matthew Riemer · 2018
Cited alongside, same era.
Multi-Task Learning as Multi-Objective Optimization
Ozan Sener and Vladlen Koltun · 2018
Cited alongside, same era.
Feng Zheng, Cheng Deng, Xing Sun, Xinyang Jiang, Xiaowei Guo, Zongqiao Yu, Feiyue Huang, and Rongrong Ji · 2019
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Later among the works it cites.
Multi-task learning with deep neural networks: A survey
Michael Crawshaw · 2020
Later among the works it cites.
XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson · 2020
Later among the works it cites.
Large-scale methods for distributionally robust optimization
Daniel Levy, Yair Carmon, John C Duchi, and Aaron Sidford · 2020
Later among the works it cites.
Effcient Continuous Pareto Exploration in Multi-Task Learning
Pingchuan Ma, Tao Du, and Wojciech Matusik · 2020
Later among the works it cites.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Later among the works it cites.
Glue benchmark leaderboard
Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2021
Closest in time.
Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual Models
Zirui Wang, Yulia Tsvetkov, Orhan Firat, and Yuan Cao · 2021
Closest in time.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel · 2021
Closest in time.
Examining and combating spurious features under distribution shift
Chunting Zhou, Xuezhe Ma, Paul Michel, and Graham Neubig · 2021
Closest in time.