Fetching the paper…
Reading the bibliography…
The pre-training of text encoders normally processes text as a sequence of tokens corresponding to small text units, such as word pieces in English and characters in Chinese.
ERNIE: Enhanced Representation through Knowledge Integration
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu. 2019a · 1904
Earlier work this paper cites.
What Does BERT Look At? An Analysis of BERT’s Attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. 2019 · 1906
Earlier work this paper cites.
Pre-Training with Whole Word Masking for Chinese BERT
Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Ziqing Yang, Shijin Wang, and Guoping Hu. 2019 · 1906
Earlier work this paper cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 1906
Earlier work this paper cites.
ERNIE 2.0: A Continual Pre-training Framework for Language Understanding
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2019b · 1907
Earlier work this paper cites.
K-BERT: Enabling Language Representation with Knowledge Graph
Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Qi Ju, Haotang Deng, and Ping Wang. 2019b · 1909
Earlier work this paper cites.
NEZHA: Neural Contextualized Representation for Chinese Language Understanding
Junqiu Wei, Xiaozhe Ren, Xiaoguang Li, Wenyong Huang, Yi Liao, Yasheng Wang, Jiashu Lin, Xin Jiang, Xiao Chen, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
LCQMC: A Large-scale Chinese Question Matching Corpus
Xin Liu, Qingcai Chen, Chong Deng, Huajun Zeng, Jing Chen, Dongfang Li, and Buzhou Tang. 2018 · 1962
Earlier work this paper cites.
The Second International Chinese Word Segmentation Bakeoff
Thomas Emerson. 2005 · 2005
Earlier work this paper cites.
The Penn Chinese TreeBank: Phrase Structure Annotation of a Large Corpus
Naiwen Xue, Fei Xia, Fu Dong Chiou, and Martha Palmer. 2005 · 2005
Earlier work this paper cites.
Transliteration of Name Entity via Improved Statistical Translation on Character Sequences
Yan Song, Chunyu Kit, and Xiao Chen. 2009 · 2009
Earlier work this paper cites.
Natural Language Processing (Almost) from Scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011 · 2011
Earlier work this paper cites.
Using a Goodness Measurement for Domain Adaptation: A Case Study on Chinese Word Segmentation
Yan Song and Fei Xia. 2012 · 2012
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Cited alongside, same era.
Deep Convolutional Neural Networks for Sentiment Analysis of Short Texts
Cicero Dos Santos and Maira Gatti. 2014 · 2014
Cited alongside, same era.
Glove: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Two/Too Simple Adaptations of Word2Vec for Syntax Problems
Wang Ling, Chris Dyer, Alan W Black, and Isabel Trancoso. 2015 · 2015
Cited alongside, same era.
Named Entity Recognition in Chinese Clinical Text Using Deep Neural Network
Yonghui Wu, Min Jiang, Jianbo Lei, and Hua Xu. 2015 · 2015
Cited alongside, same era.
A Hybrid Word-Character Model for Abstractive Summarization
Chieh-Teng Chang, Chi-Chia Huang, and Jane Yung-jen Hsu. 2018 · 2018
Later among the works it cites.
XNLI: Evaluating Cross-lingual Sentence Representations
Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Deep Contextualized Word Representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018a · 2018
Later among the works it cites.
Improving Language Understanding by Generative Pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016 · 2016
Cited alongside, same era.
context2vec: Learning Generic Context Embedding with Bidirectional LSTM
Oren Melamud, Jacob Goldberger, and Ido Dagan. 2016 · 2016
Cited alongside, same era.
THUCTC: An Efficient Chinese Text Classifier
M Sun, J Li, Z Guo, Z Yu, Y Zheng, X Si, and Z Liu. 2016 · 2016
Cited alongside, same era.
Enriching Word Vectors with Subword Information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. 2017 · 2017
Cited alongside, same era.
Learning Word Representations with Regularization from Prior Knowledge
Yan Song, Chia-Jung Lee, and Fei Xia. 2017 · 2017
Cited alongside, same era.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, L · 2017
Cited alongside, same era.
Later among the works it cites.
Complementary Learning of Word Embeddings
Yan Song and Shuming Shi. 2018 · 2018
Later among the works it cites.
Directional Skip-Gram: Explicitly Distinguishing Left and Right Context for Word Embeddings
Yan Song, Shuming Shi, Jing Li, and Haisong Zhang. 2018 · 2018
Later among the works it cites.
Incorporating Word Attention into Character-Based Word Segmentation
Shohei Higashiyama, Masao Utiyama, Eiichiro Sumita, Masao Ideuchi, Yoshiaki Oida, Yohei Sakamoto, and Isaac Okada. 2019 · 2019
Closest in time.
What Does BERT Learn about the Structure of Language?
Ganesh Jawahar, Benoît Sagot, Djamé Seddah, Samuel Unicomb, Gerardo Iñiguez, Márton Karsai, Yannick Léo, Márton Karsai, Carlos Sarraute, Éric Fleury, et al. 2019 · 2019
Closest in time.
An Encoding Strategy Based Word-Character LSTM for Chinese NER
Wei Liu, Tongge Xu, Qinghua Xu, Jiayu Song, and Yueran Zu. 2019a · 2019
Closest in time.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
Adapting bert for target-oriented multimodal sentiment classification
Jianfei Yu and Jing Jiang. 2019 · 2019
Closest in time.