Fetching the paper…
Reading the bibliography…
Recently, pre-trained language models have achieved remarkable success in a broad range of natural language processing tasks.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. 2001 · 2001
Earlier work this paper cites.
Towards computational guessing of unknown word meanings: The ontological semantic approach
Julia M Taylor, Victor Raskin, and Christian F Hempelmann. 2011 · 2011
Earlier work this paper cites.
Robust learning in random subspaces: Equipping nlp for oov effects
Anders Søgaard and Anders Johannsen. 2012 · 2012
Earlier work this paper cites.
Deep canonical correlation analysis
Galen Andrew, Raman Arora, Jeff Bilmes, and Karen Livescu. 2013 · 2013
Earlier work this paper cites.
Universal dependency annotation for multilingual parsing
Ryan McDonald, Joakim Nivre, Yvonne Quirmbach-Brundage, Yoav Goldberg, Dipanjan Das, Kuzman Ganchev, Keith Hall, Slav Petrov, Hao Zhang, Oscar Täckström, et al. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Graph propagation for paraphrasing out-of-vocabulary words in statistical machine translation
Majid Razmara, Maryam Siahbani, Reza Haffari, and Anoop Sarkar. 2013 · 2013
Earlier work this paper cites.
Bilingual word embeddings for phrase-based machine translation
Will Y Zou, Richard Socher, Daniel Cer, and Christopher D Manning. 2013 · 2013
Earlier work this paper cites.
Improving vector space word representations using multilingual correlation
Manaal Faruqui and Chris Dyer. 2014 · 2014
Earlier work this paper cites.
Sentiment lexicon interpolation and polarity estimation of objective and out-of-vocabulary words to improve sentiment classification on microblogging
Yongyos Kaewpitakkun, Kiyoaki Shirai, and Masnizah Mohd. 2014 · 2014
Earlier work this paper cites.
Learning character-level representations for part-of-speech tagging
Cicero D Santos and Bianca Zadrozny. 2014 · 2014
Earlier work this paper cites.
On using very large target vocabulary for neural machine translation
Sébastien Jean Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Foreebank: Syntactic analysis of customer support forums
Rasoul Kaljahi, Jennifer Foster, Johann Roturier, Corentin Ribeyre, Teresa Lynn, and Joseph Le Roux. 2015 · 2015
Earlier work this paper cites.
Finding function in form: Compositional character models for open vocabulary word representation
Wang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramon Fermandez, Silvio Amir, Luis Marujo, and Tiago Luis. 2015a · 2015
Earlier work this paper cites.
Deep multilingual correlation for improved word embeddings
Ang Lu, Weiran Wang, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2015 · 2015
Earlier work this paper cites.
Addressing the rare word problem in neural machine translation
Thang Luong, Ilya Sutskever, Quoc Le, Oriol Vinyals, and Wojciech Zaremba. 2015 · 2015
Earlier work this paper cites.
Named entity recognition for chinese social media with jointly trained embeddings
Nanyun Peng and Mark Dredze. 2015 · 2015
Earlier work this paper cites.
Boosting named entity recognition with neural character embeddings
Cıcero dos Santos, Victor Guimaraes, RJ Niterói, and Rio de Janeiro. 2015 · 2015
Earlier work this paper cites.
Adapting lexical representation and oov handling from written to spoken language with word embedding
Jeremie Tafforeau, Thierry Artieres, Benoit Favre, and Frederic Bechet. 2015 · 2015
Earlier work this paper cites.
Normalized word embedding and orthogonal transform for bilingual word translation
Chao Xing, Dong Wang, Chao Liu, and Yiye Lin. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
A character-level decoder without explicit segmentation for neural machine translation
Junyoung Chung, Kyunghyun Cho, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Multilingual language processing from bytes
Dan Gillick, Cliff Brunk, Oriol Vinyals, and Amarnag Subramanya. 2016 · 2016
Earlier work this paper cites.
Charagram: Embedding words and sentences via character n-grams
John Wieting Mohit Bansal Kevin Gimpel and Karen Livescu. 2016 · 2016
Cited alongside, same era.
Tracking the world state with recurrent entity networks
Mikael Henaff, Jason Weston, Arthur Szlam, Antoine Bordes, and Yann LeCun. 2016 · 2016
Cited alongside, same era.
Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush. 2016 · 2016
Cited alongside, same era.
Towards zero unknown word in neural machine translation
Xiaoqing Li, Jiajun Zhang, and Chengqing Zong. 2016 · 2016
Cited alongside, same era.
Leveraging lexical resources for learning entity embeddings in multi-relational data
Teng Long, Ryan Lowe, Jackie Chi Kit Cheung, and Doina Precup. 2016 · 2016
Cited alongside, same era.
A sub-character architecture for korean language processing
Karl Stratos. 2017 · 2017
Later among the works it cites.
Breaking the softmax bottleneck: A high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen. 2017 · 2017
Later among the works it cites.
Joint embeddings of chinese words, characters, and fine-grained subcharacter components
Jinxing Yu, Xun Jian, Hao Xin, and Yangqiu Song. 2017 · 2017
Later among the works it cites.
Overview of the CALCS 2018 Shared Task: Named Entity Recognition on Code-switched Data
Gustavo Aguilar, Fahad AlGhamdi, Victor Soto, Mona Diab, Julia Hirschberg, and Thamar Solorio. 2018 · 2018
Later among the works it cites.
Gromov-wasserstein alignment of word embedding spaces
David Alvarez-Melis and Tommi Jaakkola. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minh-Thang Luong and Christopher D Manning. 2016 · 2016
Cited alongside, same era.
Mapping unseen words to task-trained embedding spaces
Pranava Swaroop Madhyastha, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2016 · 2016
Cited alongside, same era.
Multilingual part-of-speech tagging with bidirectional long short-term memory models and auxiliary loss
Barbara Plank, Anders Søgaard, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
Ultradense word embeddings by orthogonal transformation
Sascha Rothe, Sebastian Ebert, and Hinrich Schütze. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Learning to compute word embeddings on the fly
Dzmitry Bahdanau, Tom Bosc, Stanisław Jastrzebski, Edward Grefenstette, Pascal Vincent, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Compositional representation of morphologically-rich input for neural machine translation
Duygu Ataman and Marcello Federico. 2018 · 2018
Later among the works it cites.
Findings of the 2018 conference on machine translation (wmt18)
Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, and Christof Monz. 2018 · 2018
Later among the works it cites.
Combining character and word information in neural machine translation using a multi-level attention
Huadong Chen, Shujian Huang, David Chiang, Xinyu Dai, and Jiajun Chen. 2018 · 2018
Later among the works it cites.
Word translation without parallel data
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
” bilingual expert” can find translation errors
Kai Fan, Bo Li, Fengming Zhou, and Jiayi Wang. 2018 · 2018
Later among the works it cites.
Universal neural machine translation for extremely low resource languages
Jiatao Gu, Hany Hassan, Jacob Devlin, and Victor OK Li. 2018 · 2018
Later among the works it cites.
Learning to generate word representations using subword information
Yeachan Kim, Kang-Min Kim, Ji-Min Lee, and SangKeun Lee. 2018 · 2018
Later among the works it cites.
Phrase-based & neural unsupervised machine translation
Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018 · 2018
Later among the works it cites.
Subword-level composition functions for learning word embeddings
Bofang Li, Aleksandr Drozd, Tao Liu, and Xiaoyong Du. 2018 · 2018
Later among the works it cites.
Generating Wikipedia by summarizing long sequences
Peter J Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018 · 2018
Later among the works it cites.
Using morphological knowledge in open-vocabulary neural language models
Austin Matthews, Graham Neubig, and Chris Dyer. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Contextual parameter generation for universal neural machine translation
Emmanouil Antonios Platanios, Mrinmaya Sachan, Graham Neubig, and Tom Mitchell. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
On the limitations of unsupervised bilingual dictionary induction
Anders Søgaard, Sebastian Ruder, and Ivan Vulić. 2018 · 2018
Later among the works it cites.
Findings of the wmt 2018 shared task on quality estimation
Lucia Specia, Frédéric Blain, Varvara Logacheva, Ramón Astudillo, and André FT Martins. 2018 · 2018
Later among the works it cites.
Code-switched named entity recognition with embedding attention
Changhan Wang, Kyunghyun Cho, and Douwe Kiela. 2018 · 2018
Later among the works it cites.
Subword-augmented embedding for cloze reading comprehension
Zhuosheng Zhang, Yafang Huang, and Hai Zhao. 2018 · 2018
Later among the works it cites.
Addressing troublesome words in neural machine translation
Yang Zhao, Jiajun Zhang, Zhongjun He, Chengqing Zong, and Hua Wu. 2018 · 2018
Later among the works it cites.