Fetching the paper…
Reading the bibliography…
Transfer and multi-task learning have traditionally focused on either a single source-target pair or very few, similar tasks.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Efficient Normal-Form Parsing for Combinatory Categorial Grammar
Jason Eisner. 1996 · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jurgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Gated Word-Character Recurrent Language Model
Yasumasa Miyamoto and Kyunghyun Cho. 2016 · 1997
Earlier work this paper cites.
Chunking with Support Vector Machines
Taku Kudo and Yuji Matsumoto. 2001 · 2001
Earlier work this paper cites.
Feature-Rich Part-of-Speech Tagging with a Cyclic Dependency Network
Kristina Toutanova, Dan Klein, Christopher D Manning, and Yoram Singer. 2003 · 2003
Earlier work this paper cites.
Framewise Phoneme Classification with Bidirectional LSTM and Other Neural Network Architectures
Alex Graves and Jurgen Schmidhuber. 2005 · 2005
Earlier work this paper cites.
Chunking and Dependency Parsing
Giuseppe Attardi and Felice Dell’Orletta. 2008 · 2008
Earlier work this paper cites.
Semi-Supervised Sequential Labeling and Segmentation Using Giga-Word Scale Unlabeled Data
Jun Suzuki and Hideki Isozaki. 2008 · 2008
Earlier work this paper cites.
Top Accuracy and Fast Dependency Parsing is not a Contradiction
Bernd Bohnet. 2010 · 2010
Earlier work this paper cites.
Natural Language Processing (Almost) from Scratch
Ronan Collobert, Jason Weston, Leon Bottou, Michael Karlen nad Koray Kavukcuoglu, and Pavel Kuksa. 2011 · 2011
Earlier work this paper cites.
Semi-supervised condensed nearest neighbor for part-of-speech tagging
Anders Søgaard. 2011 · 2011
Earlier work this paper cites.
Learning with Lookahead: Can History-Based Models Rival Globally Optimized Models?
Yoshimasa Tsuruoka, Yusuke Miyao, and Jun’ichi Kazama. 2011 · 2011
Earlier work this paper cites.
Maxout Networks
Ian J. Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio. 2013 · 2013
Earlier work this paper cites.
Distributed Representations of Words and Phrases and their Compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Illinois-LH: A Denotational and Distributional Approach to Semantics
Alice Lai and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
SemEval-2014 Task 1: Evaluation of Compositional Distributional Semantic Models on Full Sentences through Semantic Relatedness and Textual Entailment
Marco Marelli, Luisa Bentivogli, Marco Baroni, Raffaella Bernardi, Stefano Menini, and Roberto Zamparelli. 2014 · 2014
Earlier work this paper cites.
Dropout improves Recurrent Neural Networks for Handwriting Recognition
Vu Pham, Theodore Bluche, Christopher Kermorvant, and Jerome Louradour. 2014 · 2014
Cited alongside, same era.
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Cited alongside, same era.
Improved Transition-Based Parsing and Tagging with Neural Networks
Chris Alberti, David Weiss, Greg Coppola, and Slav Petrov. 2015 · 2015
Cited alongside, same era.
Transition-Based Dependency Parsing with Stack Long Short-Term Memory
Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, and Noah A. Smith. 2015 · 2015
Cited alongside, same era.
Recurrent Convolutional Neural Networks for Text Classification
Siwei Lai, Liheng Xu, Kang Liu, and Jun Zhao. 2015 · 2015
Cited alongside, same era.
Zhizhong Li and Derek Hoiem. 2016 · 2016
Closest in time.
Multi-task Sequence to Sequence Learning
Minh-Thang Luong, Ilya Sutskever, Quoc V. Le, Oriol Vinyals, and Lukasz Kaiser. 2016 · 2016
Closest in time.
End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF
Xuezhe Ma and Eduard Hovy. 2016 · 2016
Closest in time.
Cross-stitch Networks for Multi-task Learning
Ishan Misra, Abhinav Shrivastava, Abhinav Gupta, and Martial Hebert. 2016 · 2016
Closest in time.
End-to-End Relation Extraction using LSTMs on Sequences and Tree Structures
Makoto Miwa and Mohit Bansal. 2016 · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramon Fermandez, Silvio Amir, Luis Marujo, and Tiago Luis. 2015 · 2015
Cited alongside, same era.
Word Embedding-based Antonym Detection using Thesauri and Distributional Information
Masataka Ono, Makoto Miwa, and Yutaka Sasaki. 2015 · 2015
Cited alongside, same era.
Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks
Kai Sheng Tai, Richard Socher, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Structured Training for Neural Network Transition-Based Parsing
David Weiss, Chris Alberti, Michael Collins, and Slav Petrov. 2015 · 2015
Cited alongside, same era.
Globally Normalized Transition-Based Neural Networks
Daniel Andor, Chris Alberti, David Weiss, Aliaksei Severyn, Alessandro Presta, Kuzman Ganchev, Slav Petrov, and Michael Collins. 2016 · 2016
Cited alongside, same era.
Enhancing and Combining Sequential and Tree LSTM for Natural Language Inference
Qian Chen, Xiaodan Zhu, Zhenhua Ling, Si Wei, and Hui Jiang. 2016 · 2016
Cited alongside, same era.
Parsing as Language Modeling
Do Kook Choe and Eugene Charniak. 2016 · 2016
Cited alongside, same era.
Andrei A. Rusu, Neil C. Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. 2016 · 2016
Closest in time.
Deep multi-task learning with low level tasks supervised at lower layers
Anders Søgaard and Yoav Goldberg. 2016 · 2016
Closest in time.
CHARAGRAM: Embedding Words and Sentences via Character n-grams
John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2016 · 2016
Closest in time.
ABCNN: Attention-Based Convolutional Neural Network for Modeling Sentence Pairs
Wenpeng Yin, Hinrich Schütze, Bing Xiang, and Bowen Zhou. 2016 · 2016
Closest in time.
Stack-propagation: Improved Representation Learning for Syntax
Yuan Zhang and David Weiss. 2016 · 2016
Closest in time.
Modelling Sentence Pairs with Tree-structured Attentive Encoder
Yao Zhou, Cong Liu, and Yan Pan. 2016 · 2016
Closest in time.
Enriching Word Vectors with Subword Information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Closest in time.
Deep Biaffine Attention for Neural Dependency Parsing
Timothy Dozat and Christopher D. Manning. 2017 · 2017
Closest in time.
Neural Machine Translation with Source-Side Latent Graph Parsing
Kazuma Hashimoto and Yoshimasa Tsuruoka. 2017 · 2017
Closest in time.
What Do Recurrent Neural Network Grammars Learn About Syntax?
Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, and Noah A. Smith. 2017 · 2017
Closest in time.
Dependency Parsing as Head Selection
Xingxing Zhang, Jianpeng Cheng, and Mirella Lapata. 2017 · 2017
Closest in time.