Fetching the paper…
Reading the bibliography…
Recent work using auxiliary prediction task classifiers to investigate the properties of LSTM representations has begun to shed light on why pretrained representations, like ELMo (Peters et al., 2018) and CoVe (McCann et al., 2017), are so beneficial for neural language understanding models.
Building a Large Annotated Corpus of English: The Penn Treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jüergen Schmidhuber · 1997
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Reservoir-based techniques for speech recognition
David Verstraeten, Benjamin Schrauwen, and Dirk Stroobandt · 2006
Earlier work this paper cites.
CCGbank: A Corpus of CCG Derivations and Dependency Structures Extracted from the Penn Treebank
Julia Hockenmaier and Mark Steedman · 2007
Earlier work this paper cites.
Moses: Open Source Toolkit for Statistical Machine Translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst · 2007
Earlier work this paper cites.
Learning grammatical structure with echo state networks
Matthew H. Tong, Adam D. Bickett, Eric M. Christiansen, and Garrison W. Cottrell · 2007
Earlier work this paper cites.
Extracting and Composing Robust Features with Denoising Autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Reservoir computing approaches to recurrent neural network training
Mantas Lukoševičius and Herbert Jaeger · 2009
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Semi-supervised Sequence Learning
Andrew M. Dai and Quoc V. Le · 2015
Cited alongside, same era.
Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2015
Cited alongside, same era.
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg · 2016
Cited alongside, same era.
Findings of the 2016 Conference on Machine Translation (WMT16)
Ondrej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurelie Neveol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri · 2016
Six Challenges for Neural Machine Translation
Philip Koehn and Rebecca Knowles · 2017
Later among the works it cites.
Learned in Translation: Contextualized Word Vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher · 2017
Later among the works it cites.
Deep belief echo-state network and its application to time series prediction
Xiaochuan Sun, Tao Li, Qun Li, Yue Huang, and Yingqi Li · 2017
Later among the works it cites.
Deep Image Prior
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2017
Later among the works it cites.
Deep RNNs Learn Hierarchical Syntax
Terra Blevins, Omer Levy, and Luke Zettlemoyer · 2018
Closest in time.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, Germàn Kruszewski, Guillaume Lample, Loï Barrault, and Marco Baroni · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Does String-Based Neural MT Learn Source Syntax?
Xing Shi, Inkit Padhi, and Kevin Knight · 2016
Cited alongside, same era.
Supervised Learning of Universal Sentence Representations from Natural Language Inference Data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loic Barrault, and Antoine Bordes · 2017
Cited alongside, same era.
Deep Learning Scaling is Predictable, Empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory F. Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Cited alongside, same era.
OpenNMT: Open-Source Toolkit for Neural Machine Translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander M. Rush · 2017
Cited alongside, same era.
Yonatan Belinkov, Lluís Màrquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass
Cited in the paper.
What do Neural Machine Translation Models Learn about Morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James R. Glass
Cited in the paper.
Universal Language Model Fine-tuning for Text Classification
Jeremy Howard and Sebastian Ruder · 2018
Closest in time.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Closest in time.