Pay less attention with lightweight and dynamic convolutions
Original
Felix Wu, Angela Fan, Alexei Baevski, Yann N Dauphin, and Michael Auli. 2019 · 1901
Earlier work this paper cites.
Fixup initialization: Residual learning without normalization
Original
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma. 2019 · 1901
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Original
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 1902
Earlier work this paper cites.
Lingvo: a modular and scalable framework for sequence-to-sequence modeling
Original
Jonathan Shen, Patrick Nguyen, Yonghui Wu, et al. 2019 · 1902
Earlier work this paper cites.
Multilingual neural machine translation with knowledge distillation
Original
Xu Tan, Yi Ren, Di He, Tao Qin, Zhou Zhao, and Tie-Yan Liu. 2019 · 1902
Earlier work this paper cites.
Massively multilingual neural machine translation
Original
Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019 · 1903
Earlier work this paper cites.
The missing ingredient in zero-shot neural machine translation
Original
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Roee Aharoni, Melvin Johnson, and Wolfgang Macherey. 2019 · 1903
Earlier work this paper cites.
Continual learning via neural pruning
Original
Siavash Golkar, Michael Kagan, and Kyunghyun Cho. 2019 · 1903
Earlier work this paper cites.
Reinforcement learning based curriculum optimization for neural machine translation
Original
Gaurav Kumar, George Foster, Colin Cherry, and Maxim Krikun. 2019 · 1903
Earlier work this paper cites.
Competence-based curriculum learning for neural machine translation
Original
Emmanouil Antonios Platanios, Otilia Stretcu, Graham Neubig, Barnabás Póczos, and Tom M. Mitchell. 2019 · 1903
Earlier work this paper cites.
Consistency by agreement in zero-shot neural machine translation
Original
Maruan Al-Shedivat and Ankur P Parikh. 2019 · 1904
Earlier work this paper cites.
Text repair model for neural machine translation
Original
Markus Freitag, Isaac Caswell, and Scott Roy. 2019 · 1904
Earlier work this paper cites.
Towards interlingua neural machine translation
Original
Carlos Escolano, Marta R Costa-jussà, and José AR Fonollosa. 2019 · 1905
Earlier work this paper cites.
Mass: Masked sequence to sequence pre-training for language generation
Original
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019 · 1905
Earlier work this paper cites.
Improved zero-shot neural machine translation via ignoring spurious correlations
Original
Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor OK Li. 2019 · 1906
Earlier work this paper cites.
Evaluating the supervised and zero-shot performance of multi-lingual translation models
Original
Chris Hokamp, John Glover, and Demian Gholipour. 2019 · 1906
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi. 2018a · 1939
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Art 2: Self-organization of stable category recognition codes for analog input patterns
Gail A Carpenter and Stephen Grossberg. 1987 · 1987
Earlier work this paper cites.
Neurocomputing: Foundations of research
Stephen Grossberg. 1988 · 1988
Earlier work this paper cites.
Lifelong robot learning
Sebastian Thrun and Tom M. Mitchell. 1995 · 1995
Earlier work this paper cites.
Multitask learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
Machine learning
Tom M. Mitchell. 1997 · 1997
Earlier work this paper cites.
Learning to Learn
Sebastian Thrun and Lorien Pratt, editors. 1998 · 1998
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
A perspective view and survey of meta-learning
Ricardo Vilalta and Youssef Drissi. 2002 · 2002
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Europarl: A Parallel Corpus for Statistical Machine Translation
Philipp Koehn. 2005 · 2005
Earlier work this paper cites.
To transfer or not to transfer
Michael T Rosenstein, Zvika Marx, Leslie Pack Kaelbling, and Thomas G Dietterich. 2005 · 2005
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Ronan Collobert and Jason Weston. 2008 · 2008
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Tags for Identifying Languages
A. Phillips and M Davis. 2009 · 2009
Earlier work this paper cites.
Multi-domain learning by confidence-weighted parameter combination
Mark Dredze, Alex Kulesza, and Koby Crammer. 2010 · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
A survey on transfer learning
S. J. Pan and Q. Yang. 2010 · 2010
Earlier work this paper cites.
Large scale parallel document mining for machine translation
Jakob Uszkoreit, Jay M Ponte, Ashok C Popat, and Moshe Dubiner. 2010 · 2010
Earlier work this paper cites.
Learning with whom to share in multi-task feature learning
Zhuoliang Kang, Kristen Grauman, and Fei Sha. 2011 · 2011
Earlier work this paper cites.
One shot learning of simple visual concepts
Brenden M. Lake, Ruslan R. Salakhutdinov, Jason Gross, and Joshua B. Tenenbaum. 2011 · 2011
Earlier work this paper cites.
Wit3: Web inventory of transcribed and translated talks
Mauro Cettolo, C Girardi, and M Federico. 2012 · 2012
Earlier work this paper cites.
Multi-domain learning: When do domains matter?
Mahesh Joshi, Mark Dredze, William W. Cohen, and Carolyn Rose. 2012 · 2012
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
New types of deep neural network learning for speech recognition and related applications: an overview
L. Deng, G. Hinton, and B. Kingsbury. 2013 · 2013
Earlier work this paper cites.
Recurrent continuous translation models
Nal Kalchbrenner and Phil Blunsom. 2013 · 2013
Earlier work this paper cites.
Lifelong machine learning systems: Beyond learning algorithms
Daniel L. Silver, Qiang Yang, and Lianghao Li. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Original
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Training deep neural networks with low precision multiplications
Original
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. 2014 · 2014
Earlier work this paper cites.
BEER: BEtter evaluation as ranking
Milos Stanojevic and Khalil Sima’an. 2014 · 2014
Earlier work this paper cites.