Fetching the paper…
Reading the bibliography…
In most settings of practical concern, machine learning practitioners know in advance what end-task they wish to boost with auxiliary tasks.
Learning many related tasks at the same time with backpropagation
Rich Caruana · 1995
Earlier work this paper cites.
On learning how to learn learning strategies
Jürgen Schmidhuber · 1995
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Lifelong learning algorithms
Sebastian Thrun · 1998
Earlier work this paper cites.
Permutation, parametric and bootstrap tests of hypotheses: a practical guide to resampling methods for testing hypotheses
Phillip I Good · 2005
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Ronan Collobert and Jason Weston · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Word representations: A simple and general method for semi-supervised learning
Joseph Turian, Lev-Arie Ratinov, and Yoshua Bengio · 2010
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le · 2015
Earlier work this paper cites.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Russ R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Chemprot-3.0: a global chemical biology diseases mapping
Jens Kringelum, Sonny Kim Kjaerulff, Søren Brunak, Ole Lund, Tudor I Oprea, and Olivier Taboureau · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Semi-supervised sequence tagging with bidirectional language models
Matthew E Peters, Waleed Ammar, Chandra Bhagavatula, and Russell Power · 2017
Earlier work this paper cites.
An overview of multi-task learning in deep neural networks
Sebastian Ruder · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Measuring the evolution of a scientific field through citation frames
David Jurgens, Srijan Kumar, Raine Hoover, Dan McFarland, and Dan Jurafsky · 2018
Cited alongside, same era.
Reviving and improving recurrent back-propagation
Renjie Liao, Yuwen Xiong, Ethan Fetaya, Lisa Zhang, KiJung Yoon, Xaq Pitkow, Raquel Urtasun, and Richard Zemel · 2018
Cited alongside, same era.
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2018
Cited alongside, same era.
Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi · 2018
Cited alongside, same era.
The natural language decathlon: Multitask learning as question answering
Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith · 2020
Later among the works it cites.
Lessons from archives: Strategies for collecting sociocultural data in machine learning
Eun Seo Jo and Timnit Gebru · 2020
Later among the works it cites.
Weight poisoning attacks on pre-trained models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Later among the works it cites.
Biobert: a pre-trained biomedical language representation model for biomedical text mining
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Meta-learning update rules for unsupervised representation learning
Luke Metz, Niru Maheswaranathan, Brian Cheung, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Understanding short-horizon bias in stochastic meta-optimization, 2018
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger Grosse · 2018
Cited alongside, same era.
Scibert: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan · 2019
Cited alongside, same era.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 2019
Cited alongside, same era.
Adaptive auxiliary task weighting for reinforcement learning
Xingyu Lin, Harjatin Singh Baweja, George Kantor, and David Held · 2019
Cited alongside, same era.
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang · 2020
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
Jonathan Lorraine, Paul Vicol, and David Duvenaud · 2020
Later among the works it cites.
Auxiliary learning by implicit differentiation
Aviv Navon, Idan Achituve, Haggai Maron, Gal Chechik, and Ethan Fetaya · 2020
Later among the works it cites.
Optimizing data usage via differentiable rewards
Xinyi Wang, Hieu Pham, Paul Michel, Antonios Anastasopoulos, Jaime Carbonell, and Graham Neubig · 2020
Later among the works it cites.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Later among the works it cites.
When do you need billions of words of pretraining data?
Yian Zhang, Alex Warstadt, Haau-Sing Li, and Samuel R Bowman · 2020
Later among the works it cites.
Muppet: Massive multi-task representations with pre-finetuning
Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta · 2021
Closest in time.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Closest in time.
Extracting training data from large language models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel · 2021
Closest in time.
Auxiliary task update decomposition : the good, the bad and the neutral
Lucio M. Dery, Yann Dauphin, and David Grangier · 2021
Closest in time.
Towards understanding and mitigating social biases in language models
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2021
Closest in time.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Timo Schick, Sahana Udupa, and Hinrich Schütze · 2021
Closest in time.
Nlp from scratch without large-scale pretraining: A simple and efficient framework
Xingcheng Yao, Yanan Zheng, Xiaocong Yang, and Zhilin Yang · 2021
Closest in time.
A survey on multi-task learning
Yu Zhang and Qiang Yang · 2021
Closest in time.