Fetching the paper…
Reading the bibliography…
Transfer learning has recently become the dominant paradigm of machine learning.
Routing networks and the challenges of modular and compositional computation
Clemens Rosenbaum, Ignacio Cases, Matthew Riemer, and Tim Klinger · 1904
Earlier work this paper cites.
Three scenarios for continual learning
Gido M. van de Ven and Andreas S. Tolias · 1904
Earlier work this paper cites.
Sparse transfer learning via winning lottery tickets
Rahul Mehta · 1905
Earlier work this paper cites.
Feature partitioning for efficient multi-task architectures
Alejandro Newell, Lu Jiang, Chong Wang, Li-Jia Li, and Jia Deng · 1908
Earlier work this paper cites.
Learning neural causal models from unknown interventions
Nan Rosemary Ke, Olexa Bilaniuk, Anirudh Goyal, Stefan Bauer, Hugo Larochelle, Chris Pal, and Yoshua Bengio · 1910
Earlier work this paper cites.
First draft of a report on the EDVAC
John von Neumann · 1945
Earlier work this paper cites.
The modularity of Mind
Jerry A. Fodor · 1983
Earlier work this paper cites.
Optimization by simulated annealing
Scott Kirkpatrick, C. Daniel Gelatt Jr, and Mario P. Vecchi · 1983
Earlier work this paper cites.
Cortical connections and parallel processing: Structure and function
Dana H Ballard · 1986
Earlier work this paper cites.
Two problems with back propagation and other steepest descent learning procedures for networks
Richard S. Sutton · 1986
Earlier work this paper cites.
Toward a theory of reinforcement-learning connectionist systems
Ronald J. Williams · 1988
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J. Cohen · 1989
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E. Hinton · 1992
Earlier work this paper cites.
The meta-pi network: Building distributed knowledge representations for robust multisource pattern recognition
John B. Hampshire and Alex Waibel · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Michael I. Jordan and Robert A. Jacobs · 1994
Earlier work this paper cites.
On the computational power of neural nets
Hava T. Siegelmann and Eduardo D. Sontag · 1995
Earlier work this paper cites.
The role of product architecture in the manufacturing firm
Karl Ulrich · 1995
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M. French · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Design rules: The power of modularity
Carliss Young Baldwin and Kim B. Clark · 2000
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Thomas G. Dietterich · 2000
Earlier work this paper cites.
Temporal Abstraction in Reinforcement Learning
Doina Precup · 2000
Earlier work this paper cites.
Stochastic weight averaging in parallel: Large-batch training that generalizes well
Vipul Gupta, Santiago Akle Serrano, and Dennis DeCoste · 2001
Earlier work this paper cites.
The unified medical language system (UMLS): integrating biomedical terminology
Olivier Bodenreider · 2004
Earlier work this paper cites.
Richard Meyes, Constantin Waubert de Puiseau, Andres Posada-Moreno, and Tobias Meisen · 2004
Earlier work this paper cites.
Global workspace theory of consciousness: toward a cognitive neuroscience of human experience
Bernard J. Baars · 2005
Earlier work this paper cites.
Spontaneous evolution of modularity and network motifs
Nadav Kashtan and Uri Alon · 2005
Earlier work this paper cites.
VerbNet: A broad-coverage, comprehensive verb lexicon
Karin Kipper Schuler · 2005
Earlier work this paper cites.
Natural selection and the origin of modules
Günter P. Wagner, Jason Mezey, and Raffaele Calabretta · 2005
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Object-oriented analysis and design with applications, third edition
Grady Booch, Robert A. Maksimchuk, Michael W. Engle, Bobbi J. Young, Jim Conallen, and Kelli A. Houston · 2008
Earlier work this paper cites.
Multi-task learning with deep neural networks: A survey
Michael Crawshaw · 2009
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
The indian buffet process: An introduction and review
Thomas L Griffiths and Zoubin Ghahramani · 2011
Earlier work this paper cites.
On the binding problem in artificial neural networks
Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber · 2012
Earlier work this paper cites.
On causal and anticausal learning
Bernhard Schölkopf, Dominik Janzing, Jonas Peters, Eleni Sgouritsa, Kun Zhang, and Joris M. Mooij · 2012
Earlier work this paper cites.
Unsupervised domain adaptation by domain invariant projection
Mahsa Baktashmotlagh, Mehrtash Tafazzoli Harandi, Brian C. Lovell, and Mathieu Salzmann · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomás Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell · 2014
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Fastfood: Approximate kernel expansions in loglinear time
Quoc Viet Le, Tamás Sarlós, and Alexander Johannes Smola · 2014
Earlier work this paper cites.
Facial landmark detection by deep multi-task learning
Zhanpeng Zhang, Ping Luo, Chen Change Loy, and Xiaoou Tang · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William J. Dally · 2015
Earlier work this paper cites.
Learning to compose neural networks for question answering
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Learning feed-forward one-shot learners
Luca Bertinetto, João F. Henriques, Jack Valmadre, Philip H. S. Torr, and Andrea Vedaldi · 2016
Earlier work this paper cites.
Domain separation networks
Konstantinos Bousmalis, George Trigeorgis, Nathan Silberman, Dilip Krishnan, and Dumitru Erhan · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor S. Lempitsky · 2016
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwinska, Sergio Gomez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John P. Agapiou, Adrià Puigdomènech Badia, Karl Moritz Hermann, Yori Zwols, Georg Ostrovski, Adam Cain, Helen King, Christopher Summerfield, Phil Blunsom, Koray Kavukcuoglu, and Demis Hassabis · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning and transfer of modulated locomotor controllers
Nicolas Heess, Gregory Wayne, Yuval Tassa, Timothy P. Lillicrap, Martin A. Riedmiller, and David Silver · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D. Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Earlier work this paper cites.
Cross-stitch networks for multi-task learning
Ishan Misra, Abhinav Shrivastava, Abhinav Gupta, and Martial Hebert · 2016
Earlier work this paper cites.
Neural programmer-interpreters
Scott E. Reed and Nando de Freitas · 2016
Earlier work this paper cites.
Andrei A. Rusu, Neil C. Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Earlier work this paper cites.
Learning hidden unit contributions for unsupervised acoustic model adaptation
Pawel Swietojanski, Jinyu Li, and Steve Renals · 2016
Earlier work this paper cites.
Expert gate: Lifelong learning with a network of experts
Rahaf Aljundi, Punarjay Chakravarty, and Tinne Tuytelaars · 2017
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Earlier work this paper cites.
Structured pruning of deep convolutional neural networks
Sajid Anwar, Kyuyeon Hwang, and Wonyong Sung · 2017
Earlier work this paper cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Earlier work this paper cites.
Universal representations: The missing link between faces, text, planktons, and cat breeds
Hakan Bilen and Andrea Vedaldi · 2017
Earlier work this paper cites.
Modulating early visual processing by language
Harm de Vries, Florian Strub, Jérémie Mary, Hugo Larochelle, Olivier Pietquin, and Aaron C. Courville · 2017
Earlier work this paper cites.
Learning modular neural network policies for multi-task and multi-robot transfer
Coline Devin, Abhishek Gupta, Trevor Darrell, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Pathnet: Evolution channels gradient descent in super neural networks
Chrisantha Fernando, Dylan Banarse, Charles Blundell, Yori Zwols, David Ha, Andrei A. Rusu, Alexander Pritzel, and Daan Wierstra · 2017
Earlier work this paper cites.
Stochastic neural networks for hierarchical reinforcement learning
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C. Daniel Freeman and Joan Bruna · 2017
Earlier work this paper cites.
Hypernetworks
David Ha, Andrew M. Dai, and Quoc V. Le · 2017
Earlier work this paper cites.
DSD: dense-sparse-dense training for deep neural networks
Song Han, Jeff Pool, Sharan Narang, Huizi Mao, Enhao Gong, Shijian Tang, Erich Elsen, Peter Vajda, Manohar Paluri, John Tran, Bryan Catanzaro, and William J. Dally · 2017
Earlier work this paper cites.
Learning to reason: End-to-end module networks for visual question answering
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Earlier work this paper cites.
Adversarial adaptation of synthetic or stale data
Young-Bum Kim, Karl Stratos, and Dongchan Kim · 2017
Earlier work this paper cites.
Fully-adaptive feature sharing in multi-task networks with applications in person attribute classification
Yongxi Lu, Abhishek Kumar, Shuangfei Zhai, Yu Cheng, Tara Javidi, and Rogério Schmidt Feris · 2017
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh · 2017
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz · 2017
Earlier work this paper cites.
Attend, adapt and transfer: Attentive deep architecture for adaptive transfer from multiple sources in the same domain
Janarthanan Rajendran, Aravind S. Lakshminarayanan, Mitesh M. Khapra, P. Prasanna, and Balaraman Ravindran · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi · 2017
Earlier work this paper cites.
An overview of multi-task learning in deep neural networks
Sebastian Ruder · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep multi-task representation learning: A tensor factorisation approach
Yongxin Yang and Timothy M. Hospedales · 2017
Earlier work this paper cites.
Modular meta-learning
Ferran Alet, Tomás Lozano-Pérez, and Leslie Pack Kaelbling · 2018
Earlier work this paper cites.
Multinomial adversarial networks for multi-domain text classification
Xilun Chen and Claire Cardie · 2018
Earlier work this paper cites.
Neural modular control for embodied question answering
Abhishek Das, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred A. Hamprecht · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P. Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Multi-source domain adaptation with mixture of experts
Jiang Guo, Darsh Shah, and Regina Barzilay · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Modular networks: Learning to decompose neural computation
Louis Kirsch, Julius Kunze, and David Barber · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden M. Lake and Marco Baroni · 2018
Earlier work this paper cites.
Measuring the intrinsic dimension of objective landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi · 2018
Earlier work this paper cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Arun Mallya and Svetlana Lazebnik · 2018
Earlier work this paper cites.
Piggyback: Adapting a single network to multiple tasks by learning to mask weights
Arun Mallya, Dillon Davis, and Svetlana Lazebnik · 2018
Earlier work this paper cites.
Beyond Shared Hierarchies: Deep Multitask Learning through Soft Layer Ordering
Elliot Meyerson and Risto Miikkulainen · 2018
Earlier work this paper cites.
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine · 2018
Earlier work this paper cites.
Learning independent causal mechanisms
Giambattista Parascandolo, Niki Kilbertus, Mateo Rojas-Carulla, and Bernhard Schölkopf · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville · 2018
Earlier work this paper cites.
Contextual parameter generation for universal neural machine translation
Emmanouil Antonios Platanios, Mrinmaya Sachan, Graham Neubig, and Tom Mitchell · 2018
Earlier work this paper cites.
Efficient parametrization of multi-domain deep neural networks
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi · 2018
Earlier work this paper cites.
Routing networks: Adaptive selection of non-linear functions for multi-task learning
Clemens Rosenbaum, Tim Klinger, and Matthew Riemer · 2018
Earlier work this paper cites.
Freezing subnetworks to analyze domain adaptation in neural machine translation
Brian Thompson, Huda Khayrallah, Antonios Anastasopoulos, Arya D. McCarthy, Kevin Duh, Rebecca Marvin, Paul McNamee, Jeremy Gwinnup, Tim Anderson, and Philipp Koehn · 2018
Earlier work this paper cites.
Neural arithmetic logic units
Andrew Trask, Felix Hill, Scott E Reed, Jack Rae, Chris Dyer, and Phil Blunsom · 2018
Earlier work this paper cites.
Deep elastic networks with model selection for multi-task learning
Chanho Ahn, Eunwoo Kim, and Songhwai Oh · 2019
Earlier work this paper cites.
Giving BERT a calculator: Finding operations and arguments with reading comprehension
Daniel Andor, Luheng He, Kenton Lee, and Emily Pitler · 2019
Earlier work this paper cites.
Simple, scalable adaptation for neural machine translation
Ankur Bapna and Orhan Firat · 2019
Earlier work this paper cites.
Budget-aware adapters for multi-domain learning
Rodrigo Ferreira Berriel, Stéphane Lathuilière, Moin Nabi, Tassilo Klein, Thiago Oliveira-Santos, Nicu Sebe, and Elisa Ricci · 2019
Earlier work this paper cites.
Stochastic filter groups for multi-task cnns: Learning specialist and generalist convolution kernels
Felix J. S. Bragman, Ryutaro Tanno, Sébastien Ourselin, Daniel C. Alexander, and Manuel Jorge Cardoso · 2019
Cited alongside, same era.
Recursive routing networks: Learning to compose modules for language understanding
Ignacio Cases, Clemens Rosenbaum, Matthew Riemer, Atticus Geiger, Tim Klinger, Alex Tamkin, Olivia Li, Sandhini Agarwal, Joshua D. Greene, Dan Jurafsky, Christopher Potts, and Lauri Karttunen · 2019
Cited alongside, same era.
Automatically composing representation transformations as a means for generalization
Michael Chang, Abhishek Gupta, Sergey Levine, and Thomas L. Griffiths · 2019
Cited alongside, same era.
On self modulation for generative adversarial networks
Ting Chen, Mario Lučić, Neil Houlsby, and Sylvain Gelly · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Toward causal representation learning
Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio · 2021
Later among the works it cites.
Multilingual domain adaptation for NMT: decoupling language and domain information with adapters
Asa Cooper Stickland, Alexandre Berard, and Vassilina Nikoulina · 2021
Later among the works it cites.
Training neural networks with fixed sparse masks
Yi-Lin Sung, Varun Nair, and Colin Raffel · 2021
Later among the works it cites.
Residual adapters for parameter-efficient ASR adaptation to atypical and accented speech
Katrin Tomanek, Vicky Zayats, Dirk Padfield, Kara Vaillancourt, and Fadi Biadsy · 2021
Later among the works it cites.
Multilingual unsupervised neural machine translation with denoising adapters
Ahmet Üstün, Alexandre Berard, Laurent Besacier, and Matthias Gallé · 2021
Later among the works it cites.
Neural algorithmic reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Cited alongside, same era.
NDDR-CNN: layerwise feature fusing in multi-task cnns by neural discriminative dimensionality reduction
Yuan Gao, Jiayi Ma, Mingbo Zhao, Wei Liu, and Alan L. Yuille · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Cited alongside, same era.
Language as an abstraction for hierarchical deep reinforcement learning
Yiding Jiang, Shixiang Gu, Kevin Murphy, and Chelsea Finn · 2019
Cited alongside, same era.
Large-scale multilingual speech recognition with a streaming end-to-end model
Anjuli Kannan, Arindrima Datta, Tara N. Sainath, Eugene Weinstein, Bhuvana Ramabhadran, Yonghui Wu, Ankur Bapna, Zhifeng Chen, and Seungji Lee · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey E. Hinton · 2019
Cited alongside, same era.
Multi-task deep neural networks for natural language understanding
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao · 2019
Cited alongside, same era.
Petar Veličković and Charles Blundell · 2021
Later among the works it cites.
Adaptable and interpretable neural MemoryOver symbolic knowledge
Pat Verga, Haitian Sun, Livio Baldini Soares, and William Cohen · 2021
Later among the works it cites.
K-Adapter: Infusing Knowledge into Pre-Trained Models with Adapters
Ruize Wang, Duyu Tang, Nan Duan, Zhongyu Wei, Xuanjing Huang, Jianshu Ji, Guihong Cao, Daxin Jiang, and Ming Zhou · 2021
Later among the works it cites.
Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models
Zirui Wang, Yulia Tsvetkov, Orhan Firat, and Yuan Cao · 2021
Later among the works it cites.
Exploring sparse expert models and beyond
An Yang, Junyang Lin, Rui Men, Chang Zhou, Le Jiang, Xianyan Jia, Ang Wang, Jie Zhang, Jiamang Wang, Yong Li, Di Zhang, Wei Lin, Lin Qu, Jingren Zhou, and Hongxia Yang · 2021
Later among the works it cites.
CrossFit: A few-shot learning challenge for cross-task generalization in NLP
Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren · 2021
Later among the works it cites.
Unsupervised domain adaptation with adapter
Rongsheng Zhang, Yinhe Zheng, Xiaoxi Mao, and Minlie Huang · 2021
Later among the works it cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Later among the works it cites.
Factual probing is [MASK]: Learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen · 2021
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Samuel K. Ainsworth, Jonathan Hayase, and Siddhartha S. Srinivasa · 2022
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan · 2022
Later among the works it cites.
Composable sparse fine-tuning for cross-lingual transfer
Alan Ansell, Edoardo Ponti, Anna Korhonen, and Ivan Vulić · 2022
Later among the works it cites.
Efficient large scale language modeling with mixtures of experts
Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giridharan Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O’Horo, Jeffrey Wang, Luke Zettlemoyer, Mona Diab, Zornitsa Kozareva, and Veselin Stoyanov · 2022
Later among the works it cites.
ATTEMPT: Parameter-efficient multi-task tuning via attentional mixtures of soft prompts
Akari Asai, Mohammadreza Salehi, Matthew Peters, and Hannaneh Hajishirzi · 2022
Later among the works it cites.
XLS-R: self-supervised cross-lingual speech representation learning at scale
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli · 2022
Later among the works it cites.
Parameter-efficient conformers via sharing sparsely-gated experts for end-to-end speech recognition
Ye Bai, Jie Li, Wenjing Han, Hao Ni, Kaituo Xu, Zhuo Zhang, Cheng Yi, and Xiaorui Wang · 2022
Later among the works it cites.
Multilingual machine translation with hyper-adapters
Christos Baziotis, Mikel Artetxe, James Cross, and Shruti Bhosale · 2022
Later among the works it cites.
BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel · 2022
Later among the works it cites.
A scalable model specialization framework for training and inference using submodels and its application to speech model personalization
Fadi Biadsy, Youzheng Chen, Xia Zhang, Oleg Rybakov, Andrew Rosenberg, and Pedro J. Moreno · 2022
Later among the works it cites.
IGLUE: A benchmark for transfer learning across modalities, tasks, and languages
Emanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, and Ivan Vulić · 2022
Later among the works it cites.
Multi-head adapter routing for data-efficient fine-tuning, 2022
Lucas Caccia, Edoardo Ponti, Lucas Liu, Matheus Pereira, Nicolas Le Roux, and Alessandro Sordoni · 2022
Later among the works it cites.
Graphical clusterability and local specialization in deep neural networks
Stephen Casper, Shlomi Hod, Daniel Filan, Cody Wild, Andrew Critch, and Stuart Russell · 2022
Later among the works it cites.
On the representation collapse of sparse mixture of experts, 2022
Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, and Furu Wei · 2022
Later among the works it cites.
Data-efficient cross-lingual transfer with language-specific subnetworks
Rochelle Choenni, Dan Garrette, and Ekaterina Shutova · 2022
Later among the works it cites.
Fusing finetuned models for better pretraining, 2022
Leshem Choshen, Elad Venezian, Noam Slonim, and Yoav Katz · 2022
Later among the works it cites.
Efficient hierarchical domain adaptation for pretrained language models
Alexandra Chronopoulou, Matthew Peters, and Jesse Dodge · 2022
Later among the works it cites.
Unified scaling laws for routed language models
Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Rutherford, Tom Hennigan, Matthew J. Johnson, Albin Cassirer, Chris Jones, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Marc’Aurelio Ranzato, Jack W. Rae, Erich Elsen, Koray Kavukcuoglu, and Karen Simonyan · 2022
Later among the works it cites.
No language left behind: Scaling human-centered machine translation
Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loïc Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang · 2022
Later among the works it cites.
Brain-like functional specialization emerges spontaneously in deep neural networks
Katharina Dobs, Julio Martinez, Alexander JE Kell, and Nancy Kanwisher · 2022
Later among the works it cites.
Cold fusion: Collaborative descent for distributed multitask finetuning
Shachar Don-Yehiya, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen · 2022
Later among the works it cites.
Glam: Efficient scaling of language models with mixture-of-experts
Nan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, Liam Fedus, Maarten P. Bosma, Zongwei Zhou, Tao Wang, Yu Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathleen S. Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang, Quoc V. Le, Yonghui Wu, Zhifeng Chen, and Claire Cui · 2022
Later among the works it cites.
Tricks for training sparse translation models
Dheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross, Mike Lewis, and Angela Fan · 2022
Later among the works it cites.
Using adapters to overcome catastrophic forgetting in end-to-end automatic speech recognition
Steven Vander Eeckt and Hugo Van hamme · 2022
Later among the works it cites.
Phylogeny-inspired adaptation of multilingual models to new languages
Fahim Faisal and Antonios Anastasopoulos · 2022
Later among the works it cites.
DRAFT: A novel framework to reduce domain shifting in self-supervised learning and its application to children’s ASR
Ruchao Fan and Abeer Alwan · 2022
Later among the works it cites.
A review of sparse expert models in deep learning
William Fedus, Jeff Dean, and Barret Zoph · 2022
Later among the works it cites.
Discovering language-neutral sub-networks in multilingual language models
Negar Foroutan, Mohammadreza Banaei, Rémi Lebret, Antoine Bosselut, and Karl Aberer · 2022
Later among the works it cites.
Deep end-to-end causal inference
Tomas Geffner, Javier Antoran, Adam Foster, Wenbo Gong, Chao Ma, Emre Kiciman, Amit Sharma, Angus Lamb, Martin Kukla, Nick Pawlowski, Miltiadis Allamanis, and Cheng Zhang · 2022
Later among the works it cites.
A multi-agent framework for the asynchronous and collaborative extension of multitask ML systems
Andrea Gesmundo · 2022
Later among the works it cites.
Sparsely activated mixture-of-experts are robust multi-task learners
Shashank Gupta, Subhabrata Mukherjee, Krishan Subudhi, Eduardo Gonzalez, Damien Jose, Ahmed Hassan Awadallah, and Jianfeng Gao · 2022
Later among the works it cites.
DEMix layers: Disentangling domains for modular language modeling
Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, and Luke Zettlemoyer · 2022
Later among the works it cites.
Towards a unified view of parameter-efficient transfer learning
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig · 2022
Later among the works it cites.
Hyperprompt: Prompt-based task-conditioning of transformers
Yun He, Huaixiu Steven Zheng, Yi Tay, Jai Prakash Gupta, Yu Du, Vamsi Aribandi, Zhe Zhao, YaGuang Li, Zhao Chen, Donald Metzler, Heng-Tze Cheng, and Ed H. Chi · 2022
Later among the works it cites.
Adapter-based extension of multi-speaker text-to-speech model for new speakers
Cheng-Ping Hsieh, Subhankar Ghosh, and Boris Ginsburg · 2022
Later among the works it cites.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch · 2022
Later among the works it cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Túlio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2022
Later among the works it cites.
Rotograd: Gradient homogenization in multitask learning
Adrián Javaloy and Isabel Valera · 2022
Later among the works it cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Later among the works it cites.
Clustering units in neural networks: upstream vs downstream information
Richard D. Lange, David S. Rolnick, and Konrad P. Kording · 2022
Later among the works it cites.
Quantifying adaptability in pre-trained language models with 500 tasks
Belinda Li, Jane Yu, Madian Khabsa, Luke Zettlemoyer, Alon Halevy, and Jacob Andreas · 2022
Later among the works it cites.
Parameter-efficient neural reranking for cross-lingual and multilingual retrieval
Robert Litschko, Ivan Vulić, and Goran Glavaš · 2022
Later among the works it cites.
P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang · 2022
Later among the works it cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2022
Later among the works it cites.
Storydall-e: Adapting pretrained text-to-image transformers for story continuation
Adyasha Maharana, Darryl Hannan, and Mohit Bansal · 2022
Later among the works it cites.
Cross-task generalization via natural language crowdsourcing instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi · 2022
Later among the works it cites.
Is a modular architecture enough?
Sarthak Mittal, Yoshua Bengio, and Guillaume Lajoie · 2022
Later among the works it cites.
Residual adapters for few-shot text-to-speech speaker adaptation
Nobuyuki Morioka, Heiga Zen, Nanxin Chen, Yu Zhang, and Yifan Ding · 2022
Later among the works it cites.
Models with Conditional Computation Learn Suboptimal Solutions
Mohammed Muqeeth, Haokun Liu, and Colin Raffel · 2022
Later among the works it cites.
Learning to compose soft prompts for compositional zero-shot learning
Nihal V Nayak, Peilin Yu, and Stephen H Bach · 2022
Later among the works it cites.
ST-Adapter: Parameter-efficient image-to-video transfer learning for action recognition
Junting Pan, Ziyi Lin, Xiatian Zhu, Jing Shao, and Hongsheng Li · 2022
Later among the works it cites.
Hierarchical3d adapters for long video-to-text summarization
Pinelopi Papalampidi and Mirella Lapata · 2022
Later among the works it cites.
BAD-X: Bilingual adapters improve zero-shot cross-lingual transfer
Marinela Parović, Goran Glavaš, Ivan Vulić, and Anna Korhonen · 2022
Later among the works it cites.
xGQA: Cross-lingual visual question answering
Jonas Pfeiffer, Gregor Geigle, Aishwarya Kamath, Jan-Martin Steitz, Stefan Roth, Ivan Vulić, and Iryna Gurevych · 2022
Later among the works it cites.
Lifting the curse of multilinguality by pre-training modular transformers
Jonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, and Mikel Artetxe · 2022
Later among the works it cites.
Combining modular skills in multitask learning
Edoardo M. Ponti, Alessandro Sordoni, and Siva Reddy · 2022
Later among the works it cites.
Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation AI scale
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He · 2022
Later among the works it cites.
Scott E. Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Nando de Freitas · 2022
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M. Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush · 2022
Later among the works it cites.
Contextual adapters for personalized speech recognition in neural transducers
Kanthashree Mysore Sathyendra, Thejaswi Muniyappa, Feng-Ju Chang, Jing Liu, Jinru Su, Grant P. Strimel, Athanasios Mouchtaris, and Siegfried Kunzmann · 2022
Later among the works it cites.
Domain adaptation and multi-domain adaptation for neural machine translation: A survey
Danielle Saunders · 2022
Later among the works it cites.
Same neurons, different languages: Probing morphosyntax in multilingual pre-trained models
Karolina Stanczak, Edoardo Ponti, Lucas Torroba Hennigen, Ryan Cotterell, and Isabelle Augenstein · 2022
Later among the works it cites.
BBTv2: pure black-box optimization can be comparable to gradient descent for few-shot learning
Tianxiang Sun, Zhengfu He, Hong Qian, Xuanjing Huang, and Xipeng Qiu · 2022
Later among the works it cites.
VL-ADAPTER: parameter-efficient transfer learning for vision-and-language tasks
Yi-Lin Sung, Jaemin Cho, and Mohit Bansal · 2022
Later among the works it cites.
Efficient adapter transfer of self-supervised speech models for automatic speech recognition
Bethan Thomas, Samuel Kessler, and Salah Karout · 2022
Later among the works it cites.
Hyper-X: A unified hypernetwork for multi-task multilingual transfer
Ahmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van Noord, and Sebastian Ruder · 2022
Later among the works it cites.
Overcoming catastrophic forgetting in zero-shot cross-lingual generation
Tu Vu, Aditya Barua, Brian Lester, Daniel Cer, Mohit Iyyer, and Noah Constant · 2022
Later among the works it cites.
Overcoming catastrophic forgetting in zero-shot cross-lingual generation
Tu Vu, Aditya Barua, Brian Lester, Daniel Cer, Mohit Iyyer, and Noah Constant · 2022
Later among the works it cites.
SPoT: Better frozen model adaptation through soft prompt transfer
Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou’, and Daniel Cer · 2022
Later among the works it cites.
Bottleneck low-rank transformers for low-resource spoken language understanding
Pu Wang and Hugo Van hamme · 2022
Later among the works it cites.
AdaMix: Mixture-of-adaptations for parameter-efficient model tuning
Yaqing Wang, Sahaj Agarwal, Subhabrata Mukherjee, Xiaodong Liu, Jing Gao, Ahmed Hassan Awadallah, and Jianfeng Gao · 2022
Later among the works it cites.
Do prompt-based models really understand the meaning of their prompts?
Albert Webson and Ellie Pavlick · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt · 2022
Later among the works it cites.
UFO: unified feature optimization
Teng Xi, Yifan Sun, Deli Yu, Bi Li, Nan Peng, Gang Zhang, Xinyu Zhang, Zhigang Wang, Jinwen Chen, Jian Wang, Lufei Liu, Haocheng Feng, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang · 2022
Later among the works it cites.
Speechmoe2: Mixture-of-experts model with improved routing
Zhao You, Shulin Feng, Dan Su, and Dong Yu · 2022
Later among the works it cites.
Meta-dmoe: Adapting to domain shift by meta-distillation from mixture-of-experts
Tao Zhong, Zhixiang Chi, Li Gu, Yang Wang, Yuanhao Yu, and Jin Tang · 2022
Later among the works it cites.
Conditional prompt learning for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.
Towards better meta-initialization with task augmentation for kindergarten-aged speech recognition
Yunzheng Zhu, Ruchao Fan, and Abeer Alwan · 2022
Later among the works it cites.
ST-MoE: designing stable and transferable sparse expert models
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus · 2022
Later among the works it cites.
Taming sparsely activated transformer with stochastic experts
Simiao Zuo, Xiaodong Liu, Jian Jiao, Young Jin Kim, Hany Hassan, Ruofei Zhang, Jianfeng Gao, and Tuo Zhao · 2022
Later among the works it cites.
Exploring efficient-tuning methods in self-supervised speech models
Zih-Ching Chen, Chin-Lun Fu, Chih-Ying Liu, Shang-Wen (Daniel) Li, and Hung-yi Lee · 2023
Closest in time.
Decouple knowledge from paramters for plug-and-play language modeling
Xin Cheng, Yankai Lin, Xiuying Chen, Dongyan Zhao, and Rui Yan · 2023
Closest in time.
Knowledge is a region in weight space for fine-tuned language models
Almog Gueta, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen · 2023
Closest in time.
Exploring the benefits of training expert language models over instruction tuning, 2023
Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, and Minjoon Seo · 2023
Closest in time.
Dataless knowledge fusion by merging weights of language models, 2023
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng · 2023
Closest in time.
Git-theta: A git extension for collaborative development of machine learning models, 2023
Nikhil Kandpal, Brian Lester, Mohammed Muqeeth, Anisha Mascarenhas, Monty Evans, Vishal Baskaran, Tenghao Huang, Haokun Liu, and Colin Raffel · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2023
Closest in time.
Augmented language models: A survey
Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ramakanth Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, Edouard Grave, Yann LeCun, and Thomas Scialom · 2023
Closest in time.
Soft merging of experts with adaptive routing, 2023
Mohammed Muqeeth, Haokun Liu, and Colin Raffel · 2023
Closest in time.
From sparse to soft mixtures of experts, 2023
Joan Puigcerver, Carlos Riquelme, Basil Mustafa, and Neil Houlsby · 2023
Closest in time.
Language Models are Multilingual Chain-of-Thought Reasoners
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei · 2023
Closest in time.
Resolving interference when merging models, 2023
Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal · 2023
Closest in time.
Retrieval-augmented multimodal language modeling
Michihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Richard James, Jure Leskovec, Percy Liang, Mike Lewis, Luke Zettlemoyer, and Wen-Tau Yih · 2023
Closest in time.
Augmentation-adapted retriever improves generalization of language models as generic plug-in
Zichun Yu, Chenyan Xiong, Shi Yu, and Zhiyuan Liu · 2023
Closest in time.
AutoPEFT: automatic configuration search for parameter-efficient fine-tuning
Han Zhou, Xingchen Wan, Ivan Vulić, and Anna Korhonen · 2023
Closest in time.
Deep multi-task learning with low level tasks supervised at lower layers
Anders Søgaard and Yoav Goldberg · 2038
Closest in time.
Learning hidden unit contribution for adapting neural machine translation models
David Vilar · 2080
Closest in time.